From 8cef7a9c00ef034ddafb87b374be5e87b6d4eb4e Mon Sep 17 00:00:00 2001 From: Lilleman auf Larv Date: Fri, 18 Sep 2026 01:10:44 +0200 Subject: [PATCH] Spell the address number column alike in the geo and locale records, and record what ports across countries --- DATA-LICENSES.md | 8 ++------ README.md | 26 +++++++++++++++++++------- data/en_US/address.json | 2 +- data/geo/SE/address.json | 4 ++-- data/geo/US/address.json | 4 ++-- data/sv_SE/address.json | 2 +- todo.md | 3 +++ 7 files changed, 30 insertions(+), 19 deletions(-) diff --git a/DATA-LICENSES.md b/DATA-LICENSES.md index 9cf38a5..3b63be0 100644 --- a/DATA-LICENSES.md +++ b/DATA-LICENSES.md @@ -1,12 +1,8 @@ # Data licenses Every shipped dataset, its source, its licence and the attribution it asks for. A -`data-import/` script rebuilds each sourced table; a curated one is hand-written. -The scripts run through `docker compose run --rm data-import data-import/.py` -and cache their downloads under `data-import/cache/`; `geo-se.py` needs a -Trafikverket API key, free at [data.trafikverket.se](https://data.trafikverket.se/), -in `TRAFIKVERKET_API_KEY` or a `--key-file`; `geo-us.py` fetches two TIGER/Line -files per county it ships, a few hundred megabytes in all. +`data-import/` script rebuilds each sourced table, run as the README's +[Development](README.md#development) section says; a curated one is hand-written. | Table | Source | Licence | Attribution | Rebuild | |-------|--------|---------|-------------|---------| diff --git a/README.md b/README.md index d5db683..beaed69 100644 --- a/README.md +++ b/README.md @@ -229,22 +229,25 @@ Each locale carries `address`, `color`, `company`, `date`, `email`, `ip`, [`DATA-LICENSES.md`](DATA-LICENSES.md) names each table's source and licence. A `geo` folder holds one tree per country under its alpha-2 code: five -[linked tables](#linked-tables) named alike, so a template ports across countries, -and an `address` record over one consistent draw of them, which the locale's -`address` reads. +[linked tables](#linked-tables) named alike, and an `address` record over one +consistent draw of them, which the locale's `address` reads. | Table | `geo.SE` | `geo.US` | Weight | |-------|----------|----------|--------| | `region` | län, by code or name | state, by USPS abbreviation or name; `code` is the FIPS code | population | | `municipality` | kommun, by code or name | county, by FIPS code or name | population | -| `locality` | postort, by name | incorporated place of 25,000 people or more, by GEOID or name; Hawaii has none | population | +| `locality` | postort, by name | incorporated place of 25,000 people or more, by GEOID or name; Hawaii has none | tätort population, the kommun's where the postort names it, else 200; place population | | `postal-code` | postnummer with street delivery, by code | ZCTA, by code | one; address ranges | | `street` | gatunamn, the ten with most road segments per postort | street name, the ten with most address ranges per place | segments; address ranges | `geo.SE.region[Skåne län].municipality` draws a kommun in Skåne, `geo.SE.locality[Lund].street` a street in Lund, and `geo.US.region[IL].locality[Springfield]` settles which Springfield. A region row -carries its `timezone`, a locality its `lat` and `lon`. +carries its `timezone`, the state's predominant zone, and a locality its `lat` and +`lon`. What ports across countries is the five table names, the `name` column, +selection by name, and the `address` record's columns `street`, `street-number`, +`postal-code` and `locality`; every other column is the country's own, `code` on a +Swedish region but `abbr` on a US one. ## Data format @@ -1017,6 +1020,11 @@ renamed or retyped line is a major. two countries match in size, and `--min-population` and `--streets-per-locality` on the import scripts build a fuller set. The two trees add about 20 ms to `New`, which loads the shipped set in about 45 ms. +- **A locale's `address` restates its country record's format.** A record cannot + read another whole and keep its columns, so `sv_SE.address` names the same four + columns as `geo.SE.address`, each a reference into it, and the format appears + twice; a column is spelled the same in both, `street-number`, so the two never + disagree on a name. - **A postort's kommun comes from its name, its tätort or its codes, never from distance.** GeoNames leaves a fifth of Sweden's codes without a kommun and carries stale spellings; the nearest code across a border named the wrong kommun @@ -1071,7 +1079,11 @@ REPIN=1 docker compose run --rm --user "$(id -u):$(id -g)" test A shipped table built from a source is rebuilt by its script under [`data-import/`](data-import), one command per dataset, fetching the source named in -[`DATA-LICENSES.md`](DATA-LICENSES.md): +[`DATA-LICENSES.md`](DATA-LICENSES.md). Downloads are cached under +`data-import/cache/`, so delete it to fetch afresh; `geo-us.py` fetches two +TIGER/Line files per county it ships, a few hundred megabytes, and `geo-se.py` needs +a Trafikverket API key, free at [data.trafikverket.se](https://data.trafikverket.se/), +in `TRAFIKVERKET_API_KEY` or a `--key-file`: ```sh docker compose run --rm --user "$(id -u):$(id -g)" data-import data-import/country.py @@ -1108,7 +1120,7 @@ datatype.go column datatypes: DataType, where datatype and null may sit, a c value.go the value proof: what a typed column or calc operand holds, checked at load data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge cmd/fejkdata/ the fejkdata CLI -data/ shipped data (JSON, and a TSV per table), embedded at build: locale folders + a misc folder +data/ shipped data (JSON, and a TSV per table), embedded at build: locale folders, geo, misc data-import/ the scripts that rebuild each sourced table (see DATA-LICENSES.md) release-tooling/ the release CI publishes from the changelog heading testdata/ the pinned shipped shape (see Versioning) diff --git a/data/en_US/address.json b/data/en_US/address.json index cde0c13..4bc72a7 100644 --- a/data/en_US/address.json +++ b/data/en_US/address.json @@ -1,7 +1,7 @@ { "format": "{street-number} {street}\n{locality}, {region} {postal-code}", "street": "{/geo.US.address.street}", - "street-number": "{/geo.US.address.number}", + "street-number": "{/geo.US.address.street-number}", "locality": "{/geo.US.address.locality}", "region": "{/geo.US.address.region}", "postal-code": "{/geo.US.address.postal-code}" diff --git a/data/geo/SE/address.json b/data/geo/SE/address.json index ca927bd..8c14e63 100644 --- a/data/geo/SE/address.json +++ b/data/geo/SE/address.json @@ -1,7 +1,7 @@ { - "format": "{street} {number}\n{postal-code} {locality}", + "format": "{street} {street-number}\n{postal-code} {locality}", "street": "{.street.name}", - "number": [ + "street-number": [ "{int(1,9)}", "{int(10,99)}", { "format": "{int(100,999)}", "weight": 0.2 }, diff --git a/data/geo/US/address.json b/data/geo/US/address.json index 8f7c0ef..d99c8d8 100644 --- a/data/geo/US/address.json +++ b/data/geo/US/address.json @@ -1,7 +1,7 @@ { - "format": "{number} {street}\n{locality}, {region} {postal-code}", + "format": "{street-number} {street}\n{locality}, {region} {postal-code}", "street": "{.street.name}", - "number": [ + "street-number": [ "{int(10,99)}", { "format": "{int(100,999)}", "weight": 2 }, { "format": "{int(1000,9999)}", "weight": 0.5 }, diff --git a/data/sv_SE/address.json b/data/sv_SE/address.json index c8df275..67461da 100644 --- a/data/sv_SE/address.json +++ b/data/sv_SE/address.json @@ -1,7 +1,7 @@ { "format": "{street} {street-number}\n{postal-code} {locality}", "street": "{/geo.SE.address.street}", - "street-number": "{/geo.SE.address.number}", + "street-number": "{/geo.SE.address.street-number}", "postal-code": "{/geo.SE.address.postal-code}", "locality": "{/geo.SE.address.locality}" } diff --git a/todo.md b/todo.md index 598b907..2cfd8e6 100644 --- a/todo.md +++ b/todo.md @@ -66,6 +66,9 @@ countries; the README maps each to the native term. outer selector's pins seed every draw group of the render. - Ship the fuller sets, every US place of 10,000 and more streets per locality, as packs; `--min-population` and `--streets-per-locality` on the scripts build them. +- Fill the 398 Swedish localities weighted 200 from SCB småorter before v0.1.0. +- Give the address records one column set across countries: `region` and + `municipality` as building-block columns on `geo.SE.address` too, in step 6. - v0.1.0 ships SE and US; then NO, DK, FI, NL, FR, AU, CA, ES, GB, DE. - Revisit an application to Lantmäteriet for the exact street to postnummer pairing after v0.1.0; today a street goes to the nearest postal code centroid.