diff --git a/.gitignore b/.gitignore index 5995de4..fc04196 100644 --- a/.gitignore +++ b/.gitignore @@ -1,3 +1,4 @@ .claude *.out -__pycache__/ \ No newline at end of file +__pycache__/ +data-import/cache/ diff --git a/CHANGELOG.md b/CHANGELOG.md index 075ad90..b83d133 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -26,3 +26,9 @@ replacement, and each removed path, column or flag. `currency`, `flag`, `languages`, `numeric` and `tld` added; `misc.currency` the current ISO 4217 currencies with a minor unit, with `decimals` and `numeric` added, and its symbols from CLDR. `DATA-LICENSES.md` lists each source. +- `geo.SE` and `geo.US`: five linked tables per country, `region`, `municipality`, + `locality`, `postal-code` and `street`, weighted by population and address counts + and built from SCB, GeoNames, Trafikverket NVDB and the US Census Bureau, and an + `address` record over one consistent draw of them. `sv_SE.address` and + `en_US.address` read those records, so `en_US.address.street` no longer carries + `name` and `suffix`, and a locale folder loads only beside `geo`. diff --git a/DATA-LICENSES.md b/DATA-LICENSES.md index 137c296..3b63be0 100644 --- a/DATA-LICENSES.md +++ b/DATA-LICENSES.md @@ -1,10 +1,16 @@ # Data licenses Every shipped dataset, its source, its licence and the attribution it asks for. A -`data-import/` script rebuilds each sourced table; a curated one is hand-written. +`data-import/` script rebuilds each sourced table, run as the README's +[Development](README.md#development) section says; a curated one is hand-written. | Table | Source | Licence | Attribution | Rebuild | |-------|--------|---------|-------------|---------| +| `geo/SE/locality.tsv`, `postal-code.tsv` | [GeoNames](https://www.geonames.org/) postal codes for SE; populations from SCB tätorter 2023 | [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/); CC0 1.0 | "Postal codes from GeoNames, www.geonames.org" | `data-import/geo-se.py` | +| `geo/SE/region.tsv`, `municipality.tsv` | [SCB](https://www.scb.se/) län and kommun codes 2026 and population 2024 | [CC0 1.0](https://creativecommons.org/publicdomain/zero/1.0/) | none required | `data-import/geo-se.py` | +| `geo/SE/street.tsv` | [Trafikverket NVDB](https://www.trafikverket.se/) Gatunamn, through the open API | CC0 1.0 | none required | `data-import/geo-se.py` | +| `geo/US/region.tsv`, `municipality.tsv`, `locality.tsv` | [Census Bureau](https://www.census.gov/) Gazetteer 2026 and population estimates 2025 | [public domain](https://www.usa.gov/government-works) | none required | `data-import/geo-us.py` | +| `geo/US/postal-code.tsv`, `street.tsv` | Census Bureau ZCTA to place relationships 2020 and TIGER/Line 2025 address ranges and feature names | public domain | none required | `data-import/geo-us.py` | | `misc/country.tsv` | [datasets/country-codes](https://github.com/datasets/country-codes) | [PDDL 1.0](https://opendatacommons.org/licenses/pddl/1-0/) | none required | `data-import/country.py` | | `misc/currency.tsv` | [datasets/currency-codes](https://github.com/datasets/currency-codes); symbols from [Unicode CLDR](https://github.com/unicode-org/cldr) `en.xml` and `root.xml` | PDDL 1.0; [Unicode License v3](https://www.unicode.org/license.txt) | CLDR: "Copyright © 1991-2025 Unicode, Inc. Unicode and the Unicode Logo are registered trademarks of Unicode, Inc. in the United States and other countries." | `data-import/currency.py` | | `misc/httpstatus.tsv` | curated (IANA HTTP status codes are facts) | — | — | — | diff --git a/README.md b/README.md index 0fab5ea..bfd609c 100644 --- a/README.md +++ b/README.md @@ -18,6 +18,7 @@ fejkdata --seed 42 sv_SE.address # the same address every run fejkdata -n 3 --separator ', ' sv_SE.word # nät, barn, sol fejkdata --list # every path the data offers fejkdata 'misc.country[SE].capital' # Stockholm — a table's row, selected by key or name +fejkdata 'geo.SE.locality[Lund].street' # Fjelievägen — a linked table, drawn inside the row fejkdata --data-path ./mydata sv_SE.word # layer a directory over the shipped data fejkdata --no-shipped-data -d ./mydata --list # only your data fejkdata 'name: {/sv_SE.person.last}' # name: — an inline template @@ -227,6 +228,27 @@ Each locale carries `address`, `color`, `company`, `date`, `email`, `ip`, `misc.country[SE].capital` and `misc.currency[Euro].symbol` select a row; [`DATA-LICENSES.md`](DATA-LICENSES.md) names each table's source and licence. +A `geo` folder holds one tree per country under its alpha-2 code: five +[linked tables](#linked-tables) named alike, and an `address` record over one +consistent draw of them, which the locale's `address` reads. + +| Table | `geo.SE` | `geo.US` | Weight | +|-------|----------|----------|--------| +| `region` | län, by code or name | state, by USPS abbreviation or name; `code` is the FIPS code | population | +| `municipality` | kommun, by code or name | county, by FIPS code or name | population | +| `locality` | postort, by name | incorporated place of 25,000 people or more with a postal code of its own, by GEOID or name; Hawaii has none | tätort population, the kommun's where the postort names it, else 200; place population | +| `postal-code` | postnummer with street delivery, by code | ZCTA, by code | one; address ranges | +| `street` | gatunamn, the ten with most road segments per postort | street name, the ten with most address ranges per place | segments; address ranges | + +`geo.SE.region[Skåne län].municipality` draws a kommun in Skåne, +`geo.SE.locality[Lund].street` a street in Lund, and +`geo.US.region[IL].locality[Springfield]` settles which Springfield. A region row +carries its `timezone`, the state's predominant zone, and a locality its `lat` and +`lon`. What ports across countries is the five table names, the `name` column, +selection by name, and the `address` record's columns `street`, `street-number`, +`postal-code` and `locality`; every other column is the country's own, `code` on a +Swedish region but `abbr` on a US one. + ## Data format Every value is a **node**, nestable without limit: @@ -983,6 +1005,41 @@ renamed or retyped line is a major. - **A path is walked once without drawing before it is walked for real.** A path that fails below its first level then moves no seeded stream, at the cost of one draw-free walk per call, which allocates nothing. +- **A country's postal codes and streets are siblings under its locality.** No open + source pairs a Swedish street with its postnummer, and pairing the US through its + ZIPs would shape the two trees differently, so both draw inside the pinned + locality and an address agrees at that level. A street's own code is the exact + pairing to add when a source carries it. +- **A locale's `address` reads its country's `geo` tree, so the shipped set loads + whole.** `data/sv_SE` alone no longer loads: a test loads `data` and prefixes + the locale, and `--no-shipped-data -d` takes the whole `data` folder or a set of + one's own. +- **The default embed holds every Swedish postort the import can place and give a + street-delivery code and a street, and the US places of 25,000 or more.** Sweden + fits whole in 700 KB; every US place of 10,000 would pass a + megabyte and fetch 1,200 counties of TIGER files, so the threshold sits where the + two countries match in size, and `--min-population` and + `--streets-per-locality` on the import scripts build a fuller set. The two trees + add about 20 ms to `New`, which loads the shipped set in about 45 ms. +- **A locale's `address` restates its country record's format.** A record cannot + read another whole and keep its columns, so `sv_SE.address` names the same four + columns as `geo.SE.address`, each a reference into it, and the format appears + twice; a column is spelled the same in both, `street-number`, so the two never + disagree on a name. +- **A postort's kommun comes from its name, its tätort or its codes, never from + distance.** GeoNames leaves a fifth of Sweden's codes without a kommun and + carries stale spellings; the nearest code across a border named the wrong kommun + half the time it was tried, so a postort none of the three rules place is + dropped, as is one not cased like a place name. +- **A highway designation is not a street, and a US postal code belongs to the place + holding most of its land inside places.** `I- 55 Bus` and `US Hwy 1` carry the + most address ranges in many places and would head every address, so the import + drops names spelled as a route. A ZCTA goes to the place its largest in-place part + lies in, census-designated places left out since they never ship, and ships only + when that place does, so a few dozen places whose every code lies mostly in a + bigger neighbour ship no address; counting the land outside every place too would + drop a quarter of the places, whose codes straddle unincorporated land, for a + postal city the USPS mostly names the same way. - **`List` advertises direct descents only.** `region.municipality.locality` is listed, and `region.locality` resolves too but is not: the set of every descent through a chain of five tables is every subsequence of it, and the direct chain is @@ -1032,11 +1089,17 @@ REPIN=1 docker compose run --rm --user "$(id -u):$(id -g)" test A shipped table built from a source is rebuilt by its script under [`data-import/`](data-import), one command per dataset, fetching the source named in -[`DATA-LICENSES.md`](DATA-LICENSES.md): +[`DATA-LICENSES.md`](DATA-LICENSES.md). Downloads are cached under +`data-import/cache/`, so delete it to fetch afresh; `geo-us.py` fetches two +TIGER/Line files per county it ships, a few hundred megabytes, and `geo-se.py` needs +a Trafikverket API key, free at [data.trafikverket.se](https://data.trafikverket.se/), +in `TRAFIKVERKET_API_KEY` or a `--key-file`: ```sh docker compose run --rm --user "$(id -u):$(id -g)" data-import data-import/country.py docker compose run --rm --user "$(id -u):$(id -g)" data-import data-import/currency.py +docker compose run --rm --user "$(id -u):$(id -g)" data-import data-import/geo-us.py +docker compose run --rm --user "$(id -u):$(id -g)" -e TRAFIKVERKET_API_KEY data-import data-import/geo-se.py ``` To release, head `CHANGELOG.md` with the version's section in place of `Unreleased` @@ -1067,7 +1130,7 @@ datatype.go column datatypes: DataType, where datatype and null may sit, a c value.go the value proof: what a typed column or calc operand holds, checked at load data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge cmd/fejkdata/ the fejkdata CLI -data/ shipped data (JSON, and a TSV per table), embedded at build: locale folders + a misc folder +data/ shipped data (JSON, and a TSV per table), embedded at build: locale folders, geo, misc data-import/ the scripts that rebuild each sourced table (see DATA-LICENSES.md) release-tooling/ the release CI publishes from the changelog heading testdata/ the pinned shipped shape (see Versioning) diff --git a/bench_test.go b/bench_test.go index d3bc76f..e016d58 100644 --- a/bench_test.go +++ b/bench_test.go @@ -26,12 +26,12 @@ func benchPath(b *testing.B, dir, path string) { } } -func BenchmarkPerson(b *testing.B) { benchPath(b, "data/sv_SE", "person") } -func BenchmarkAddress(b *testing.B) { benchPath(b, "data/sv_SE", "address") } -func BenchmarkWord(b *testing.B) { benchPath(b, "data/sv_SE", "word") } -func BenchmarkCreditcard(b *testing.B) { benchPath(b, "data/misc", "creditcard") } -func BenchmarkSSN(b *testing.B) { benchPath(b, "data/sv_SE", "ssn") } -func BenchmarkUUIDv7(b *testing.B) { benchPath(b, "data/misc", "uuid") } +func BenchmarkPerson(b *testing.B) { benchPath(b, "data", "sv_SE.person") } +func BenchmarkAddress(b *testing.B) { benchPath(b, "data", "sv_SE.address") } +func BenchmarkWord(b *testing.B) { benchPath(b, "data", "sv_SE.word") } +func BenchmarkCreditcard(b *testing.B) { benchPath(b, "data", "misc.creditcard") } +func BenchmarkSSN(b *testing.B) { benchPath(b, "data", "sv_SE.ssn") } +func BenchmarkUUIDv7(b *testing.B) { benchPath(b, "data", "misc.uuid") } func tmpData(b *testing.B, name, body string) string { b.Helper() diff --git a/cmd/fejkdata/main_test.go b/cmd/fejkdata/main_test.go index d591992..626234a 100644 --- a/cmd/fejkdata/main_test.go +++ b/cmd/fejkdata/main_test.go @@ -14,6 +14,7 @@ import ( const ( svSE = "../../data/sv_SE" enUS = "../../data/en_US" + misc = "../../data/misc" ) func runOut(args ...string) (int, string, string) { @@ -405,14 +406,14 @@ func TestRunTemplateMisuse(t *testing.T) { } func TestRunNoShippedData(t *testing.T) { - code, list, errb := runOut("--no-shipped-data", "-d", svSE, "--list") + code, list, errb := runOut("--no-shipped-data", "-d", misc, "--list") if code != 0 { t.Fatalf("run = %d, stderr=%q", code, errb) } - if strings.Contains(list, "en_US") || !strings.Contains(list, "person\n") { + if strings.Contains(list, "sv_SE") || !strings.Contains(list, "uuid\n") { t.Errorf("--no-shipped-data --list = %q, want only the given dir", list) } - code, out, _ := runOut("--no-shipped-data", "-d", svSE, "-s", "3", "person") + code, out, _ := runOut("--no-shipped-data", "-d", misc, "-s", "3", "uuid") if code != 0 || strings.TrimSpace(out) == "" { t.Errorf("run = %d, out=%q", code, out) } diff --git a/data-import/country.py b/data-import/country.py index 326ab57..9c67cff 100644 --- a/data-import/country.py +++ b/data-import/country.py @@ -1,30 +1,24 @@ #!/usr/bin/env python3 """Rebuild data/misc/country.tsv from datasets/country-codes (PDDL). - data-import/country.py [--source URL_OR_FILE] [--out FILE] + data-import/country.py [--source URL_OR_FILE] [--cache DIR] [--out FILE] """ import argparse import csv import io import re -import sys -import urllib.request from pathlib import Path +import tsv + SOURCE = "https://raw.githubusercontent.com/datasets/country-codes/main/data/country-codes.csv" OUT = Path(__file__).resolve().parent.parent / "data" / "misc" / "country.tsv" +CACHE = Path(__file__).resolve().parent / "cache" COLUMNS = ["alpha2", "alpha3", "calling-code", "capital", "currency", "flag", "languages", "name", "numeric", "tld"] # Gaps in the source, keyed by alpha2. FIXUPS = {"TR": {"currency": "TRY"}} -def read(source): - if re.match(r"^https?://", source): - with urllib.request.urlopen(source, timeout=60) as r: - return r.read().decode("utf-8") - return Path(source).read_text(encoding="utf-8") - - def flag(alpha2): return "".join(chr(0x1F1E6 + ord(c) - ord("A")) for c in alpha2) @@ -58,23 +52,14 @@ def rows(text): yield row -def write(out, table): - lines = ["\t".join(COLUMNS)] - for row in sorted(table, key=lambda r: r["alpha2"]): - cells = [row[c] for c in COLUMNS] - assert not any("\t" in c or "\n" in c for c in cells), row - lines.append("\t".join(cells)) - Path(out).write_text("\n".join(lines) + "\n", encoding="utf-8") - return len(lines) - 1 - - def main(): p = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + p.add_argument("--cache", default=str(CACHE)) p.add_argument("--source", default=SOURCE) p.add_argument("--out", default=str(OUT)) a = p.parse_args() - n = write(a.out, rows(read(a.source))) - print(f"{a.out}: {n} rows", file=sys.stderr) + table = rows(tsv.fetch(a.source, a.cache, "country-codes.csv").decode("utf-8")) + tsv.write(a.out, COLUMNS, sorted(table, key=lambda r: r["alpha2"])) if __name__ == "__main__": diff --git a/data-import/currency.py b/data-import/currency.py index e25caf6..6f843aa 100644 --- a/data-import/currency.py +++ b/data-import/currency.py @@ -1,35 +1,28 @@ #!/usr/bin/env python3 """Rebuild data/misc/currency.tsv from datasets/currency-codes (PDDL) and CLDR's symbols (Unicode). - data-import/currency.py [--source URL_OR_FILE] [--symbols URL_OR_FILE ...] [--out FILE] + data-import/currency.py [--source URL_OR_FILE] [--symbols URL_OR_FILE ...] [--cache DIR] [--out FILE] The symbols come from the first locale file that has one, narrow symbols before wide. """ import argparse import csv import io -import re -import sys -import urllib.request import xml.etree.ElementTree as ET from pathlib import Path +import tsv + SOURCE = "https://raw.githubusercontent.com/datasets/currency-codes/main/data/codes-all.csv" SYMBOLS = [ "https://raw.githubusercontent.com/unicode-org/cldr/main/common/main/en.xml", "https://raw.githubusercontent.com/unicode-org/cldr/main/common/main/root.xml", ] OUT = Path(__file__).resolve().parent.parent / "data" / "misc" / "currency.tsv" +CACHE = Path(__file__).resolve().parent / "cache" COLUMNS = ["code", "decimals", "name", "numeric", "symbol"] -def read(source): - if re.match(r"^https?://", source): - with urllib.request.urlopen(source, timeout=60) as r: - return r.read().decode("utf-8") - return Path(source).read_text(encoding="utf-8") - - def symbols(xml_texts): """CLDR's symbol per code: the first locale's narrow symbol, else the first locale's wide one.""" narrow, wide = {}, {} @@ -58,24 +51,16 @@ def rows(text, symbol): } -def write(out, table): - lines = ["\t".join(COLUMNS)] - for row in sorted(table, key=lambda r: r["code"]): - cells = [row[c] for c in COLUMNS] - assert all(cells) and not any("\t" in c or "\n" in c for c in cells), row - lines.append("\t".join(cells)) - Path(out).write_text("\n".join(lines) + "\n", encoding="utf-8") - return len(lines) - 1 - - def main(): p = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + p.add_argument("--cache", default=str(CACHE)) p.add_argument("--source", default=SOURCE) p.add_argument("--symbols", nargs="+", default=SYMBOLS) p.add_argument("--out", default=str(OUT)) a = p.parse_args() - n = write(a.out, rows(read(a.source), symbols(read(s) for s in a.symbols))) - print(f"{a.out}: {n} rows", file=sys.stderr) + symbol = symbols(tsv.fetch(s, a.cache, Path(s).name).decode("utf-8") for s in a.symbols) + table = rows(tsv.fetch(a.source, a.cache, "codes-all.csv").decode("utf-8"), symbol) + tsv.write(a.out, COLUMNS, sorted(table, key=lambda r: r["code"])) if __name__ == "__main__": diff --git a/data-import/geo-se.py b/data-import/geo-se.py new file mode 100644 index 0000000..bb4a738 --- /dev/null +++ b/data-import/geo-se.py @@ -0,0 +1,246 @@ +#!/usr/bin/env python3 +"""Rebuild data/geo/SE/*.tsv from SCB (CC0), GeoNames (CC BY 4.0) and Trafikverket NVDB (CC0). + + TRAFIKVERKET_API_KEY=… data-import/geo-se.py [--key-file FILE] [--cache DIR] [--streets-per-locality N] [--out DIR] +""" +import argparse +import collections +import csv +import io +import json +import math +import os +import re +import sys +import urllib.request +import xml.etree.ElementTree as ET +import zipfile +from pathlib import Path + +import tsv + +CODES = "https://www.scb.se/contentassets/7a89e48960f741e08918e489ea36354a/kommunlankod-2026.xlsx" +POPULATION = "https://api.scb.se/OV0104/v1/doris/sv/ssd/START/BE/BE0101/BE0101A/BefolkningNy" +POPULATION_QUERY = { + "query": [ + {"code": "Region", "selection": {"filter": "all", "values": ["*"]}}, + {"code": "ContentsCode", "selection": {"filter": "item", "values": ["BE0101N1"]}}, + {"code": "Tid", "selection": {"filter": "top", "values": ["1"]}}, + ], + "response": {"format": "json"}, +} +TATORTER = "https://geodata.scb.se/geoserver/stat/wfs?service=WFS&version=2.0.0&request=GetFeature&typeNames=stat:Tatorter_2023&outputFormat=csv&propertyName=tatort,kommun,bef" +POSTAL_CODES = "https://download.geonames.org/export/zip/SE.zip" +NVDB = "https://api.trafikinfo.trafikverket.se/v2/data.json" +NVDB_PAGE = 50000 +OUT = Path(__file__).resolve().parent.parent / "data" / "geo" / "SE" +CACHE = Path(__file__).resolve().parent / "cache" +TIMEZONE = "Europe/Stockholm" +ONE_POSITION = {"Stockholm", "Göteborg", "Malmö"} +UNMATCHED_POPULATION = 200 +XLSX_NS = {"m": "http://schemas.openxmlformats.org/spreadsheetml/2006/main"} + + +def xlsx_rows(data): + z = zipfile.ZipFile(io.BytesIO(data)) + strings = ["".join(t.text or "" for t in si.iter("{%s}t" % XLSX_NS["m"])) for si in ET.fromstring(z.read("xl/sharedStrings.xml")).findall("m:si", XLSX_NS)] + sheet = ET.fromstring(z.read("xl/worksheets/sheet1.xml")) + for row in sheet.findall(".//m:row", XLSX_NS): + cells = [] + for c in row.findall("m:c", XLSX_NS): + v = c.find("m:v", XLSX_NS) + cells.append("" if v is None else strings[int(v.text)] if c.get("t") == "s" else v.text) + yield cells + + +def scb_codes(cache): + regions, municipalities = {}, {} + for cells in xlsx_rows(tsv.fetch(CODES, cache, "kommunlankod.xlsx", magic=b"PK")): + if len(cells) < 2 or not re.fullmatch(r"\d{2}|\d{4}", cells[0]): + continue + (regions if len(cells[0]) == 2 else municipalities)[cells[0]] = cells[1].strip() + return regions, municipalities + + +def scb_population(cache): + body = json.dumps(POPULATION_QUERY).encode() + data = tsv.fetch(POPULATION, cache, "befolkning.json", data=body, headers={"Content-Type": "application/json"}) + return {row["key"][0]: row["values"][0] for row in json.loads(data.decode("utf-8-sig"))["data"]} + + +def scb_tatorter(cache): + text = tsv.fetch(TATORTER, cache, "tatorter.csv").decode("utf-8") + by_name = collections.defaultdict(list) + for r in csv.DictReader(io.StringIO(text)): + by_name[r["tatort"]].append((r["kommun"], int(r["bef"]))) + return by_name + + +def geonames(cache): + z = zipfile.ZipFile(io.BytesIO(tsv.fetch(POSTAL_CODES, cache, "SE.zip", magic=b"PK"))) + rows = [] + for line in z.read("SE.txt").decode("utf-8").splitlines(): + f = line.split("\t") + lat, lon = (float(f[9]), float(f[10])) if f[9] and f[10] else (None, None) + rows.append({"code": f[1], "locality": f[2], "municipality": f[6], "lat": lat, "lon": lon}) + return rows + + +def nvdb_segments(cache, key): + path = cache / "nvdb-gatunamn.tsv" + if not path.exists(): + with open(path.with_suffix(".part"), "w", encoding="utf-8") as out: + change = "0" + while True: + query = ( + f'' + f'' + "" + "NamnGeometry.WKT-WGS84-3D" + ) + req = urllib.request.Request(NVDB, data=query.encode(), headers={"Content-Type": "text/xml"}) + with urllib.request.urlopen(req, timeout=600) as r: + result = json.load(r)["RESPONSE"]["RESULT"][0] + rows = result.get("Gatunamn", []) + for row in rows: + m = re.match(r"LINESTRING Z \(([-\d.]+) ([-\d.]+) ", row.get("Geometry", {}).get("WKT-WGS84-3D", "")) + name = " ".join(row.get("Namn", "").split()) + if m and name: + out.write(f"{name}\t{m.group(1)}\t{m.group(2)}\n") + change = result["INFO"]["LASTCHANGEID"] + if len(rows) < NVDB_PAGE: + break + path.with_suffix(".part").rename(path) + for line in path.read_text(encoding="utf-8").splitlines(): + name, lon, lat = line.split("\t") + yield name, float(lat), float(lon) + + +class Nearest: + """Nearest point by an equirectangular distance, over a degree grid.""" + + def __init__(self, points, cell=0.05): + self.cell = cell + self.grid = collections.defaultdict(list) + for lat, lon, value in points: + self.grid[(int(lat // cell), int(lon // cell))].append((lat, lon, value)) + + def find(self, lat, lon): + ci, cj = int(lat // self.cell), int(lon // self.cell) + best, best_d = None, math.inf + ring = 0 + while ring < 400: + for i in range(ci - ring, ci + ring + 1): + for j in range(cj - ring, cj + ring + 1): + if max(abs(i - ci), abs(j - cj)) != ring: + continue + for plat, plon, value in self.grid.get((i, j), ()): + d = (plat - lat) ** 2 + ((plon - lon) * math.cos(math.radians(lat))) ** 2 + if d < best_d: + best, best_d = value, d + if best is not None and math.sqrt(best_d) < ring * self.cell * math.cos(math.radians(lat)): + return best + ring += 1 + return best + + +def street_delivery(name, codes): + """The codes delivered to a street: the digit after the postort's own prefix says box, company or reply.""" + if name in ONE_POSITION: + return [c for c in codes if c[1] != "0"] + largest = collections.Counter(c[:3] for c in codes).most_common(1)[0][1] + length = 3 if largest * 2 >= len(codes) else 2 + return [c for c in codes if c[length] not in "018"] + + +def municipality_of(name, rows, tatorter, municipalities): + named = [code for code, n in municipalities.items() if n == name] + if named: + return named[0], "kommun" + voted = collections.Counter(r["municipality"] for r in rows if r["municipality"] in municipalities) + matches = tatorter.get(name, []) + if matches: + in_vote = [m for m in matches if voted and m[0] == voted.most_common(1)[0][0]] + return max(in_vote or matches, key=lambda m: m[1])[0], "tatort" + if voted: + return voted.most_common(1)[0][0], "codes" + return None, "unplaced" + + +def population_of(name, municipality, tatorter, municipalities, population): + matches = [m for m in tatorter.get(name, []) if m[0] == municipality] + if matches: + return str(max(m[1] for m in matches)) + return population[municipality] if municipalities[municipality] == name else str(UNMATCHED_POPULATION) + + +def well_cased(name): + return all(part[:1].isupper() and (len(part) == 1 or not part.isupper()) for part in re.split(r"[ -]", name)) + + +def localities(codes, tatorter, municipalities, population): + """Each postort with its municipality, weight, centroid and street-delivery codes.""" + by_locality = collections.defaultdict(list) + for r in codes: + by_locality[r["locality"]].append(r) + out, how = {}, collections.Counter() + for name, rows in by_locality.items(): + municipality, method = municipality_of(name, rows, tatorter, municipalities) + how[method] += 1 + kept = street_delivery(name, [r["code"].replace(" ", "") for r in rows]) + with_point = [r for r in rows if r["lat"] is not None] + if municipality is None or not kept or not with_point or not well_cased(name): + continue + lat = sum(r["lat"] for r in with_point) / len(with_point) + lon = sum(r["lon"] for r in with_point) / len(with_point) + out[name] = {"name": name, "municipality": municipality, "population": population_of(name, municipality, tatorter, municipalities, population), "lat": f"{lat:.4f}", "lon": f"{lon:.4f}", "codes": kept} + print(f"municipality by {dict(how)}; {len(by_locality) - len(out)} postorter dropped", file=sys.stderr) + return out + + +def streets(segments, codes, localities, per_locality): + """The names with most segments per locality, each segment at its nearest code centroid.""" + nearest = Nearest((r["lat"], r["lon"], r["locality"]) for r in codes if r["lat"] is not None and r["locality"] in localities) + count = collections.Counter() + for name, lat, lon in segments: + if name[0].isalpha(): + count[(nearest.find(lat, lon), name)] += 1 + of = collections.defaultdict(list) + for (locality, name), n in count.items(): + of[locality].append((n, name)) + return {locality: [{"name": name, "locality": locality, "segments": n} for n, name in sorted(named, key=lambda s: (-s[0], s[1]))[:per_locality]] for locality, named in of.items()} + + +def main(): + p = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + p.add_argument("--cache", default=str(CACHE)) + p.add_argument("--key-file", help="file holding the Trafikverket API key; TRAFIKVERKET_API_KEY otherwise") + p.add_argument("--out", default=str(OUT)) + p.add_argument("--streets-per-locality", type=int, default=10) + a = p.parse_args() + key = Path(a.key_file).read_text().strip() if a.key_file else os.environ.get("TRAFIKVERKET_API_KEY") + if not key: + sys.exit("set TRAFIKVERKET_API_KEY or pass --key-file") + cache, out = Path(a.cache), Path(a.out) + cache.mkdir(parents=True, exist_ok=True) + out.mkdir(parents=True, exist_ok=True) + + regions, municipalities = scb_codes(cache) + population = scb_population(cache) + codes = geonames(cache) + places = localities(codes, scb_tatorter(cache), municipalities, population) + named = streets(nvdb_segments(cache, key), codes, places, a.streets_per_locality) + places = {name: l for name, l in places.items() if name in named} + empty = sorted(m for m in municipalities if not any(l["municipality"] == m for l in places.values())) + if empty: + sys.exit(f"municipalities without a locality: {empty}") + + tsv.write(out / "region.tsv", ["code", "name", "population", "timezone"], [{"code": c, "name": n, "population": population[c], "timezone": TIMEZONE} for c, n in sorted(regions.items())]) + tsv.write(out / "municipality.tsv", ["code", "name", "region", "population"], [{"code": c, "name": n, "region": c[:2], "population": population[c]} for c, n in sorted(municipalities.items())]) + tsv.write(out / "locality.tsv", ["name", "municipality", "population", "lat", "lon"], [l for _, l in sorted(places.items())]) + tsv.write(out / "postal-code.tsv", ["code", "locality"], sorted(({"code": f"{c[:3]} {c[3:]}", "locality": l["name"]} for l in places.values() for c in l["codes"]), key=lambda r: r["code"])) + tsv.write(out / "street.tsv", ["name", "locality", "segments"], [s for locality in sorted(named) for s in named[locality]]) + + +if __name__ == "__main__": + main() diff --git a/data-import/geo-us.py b/data-import/geo-us.py new file mode 100644 index 0000000..9631d48 --- /dev/null +++ b/data-import/geo-us.py @@ -0,0 +1,201 @@ +#!/usr/bin/env python3 +"""Rebuild data/geo/US/*.tsv from the Census Bureau's Gazetteer, population estimates, ZCTA relationships and TIGER/Line files (public domain). + + data-import/geo-us.py [--cache DIR] [--min-population N] [--streets-per-locality N] [--out DIR] +""" +import argparse +import collections +import concurrent.futures +import csv +import io +import re +import struct +import sys +import zipfile +from pathlib import Path + +import tsv + +GAZETTEER = "https://www2.census.gov/geo/docs/maps-data/data/gazetteer/2026_Gazetteer/2026_Gaz_{}_national.zip" +POPULATION = "https://www2.census.gov/programs-surveys/popest/datasets/2020-2025/{}" +STATES = POPULATION.format("state/totals/NST-EST2025-ALLDATA.csv") +COUNTIES = POPULATION.format("counties/totals/co-est2025-alldata.csv") +PLACES = POPULATION.format("cities/totals/sub-est2025.csv") +ZCTA_PLACE = "https://www2.census.gov/geo/docs/maps-data/data/rel2020/zcta520/tab20_zcta520_place20_natl.txt" +TIGER = "https://www2.census.gov/geo/tiger/TIGER2025/{0}/tl_2025_{1}_{2}.zip" +OUT = Path(__file__).resolve().parent.parent / "data" / "geo" / "US" +CACHE = Path(__file__).resolve().parent / "cache" +ESTIMATE = "POPESTIMATE2025" +CDP = "57" +HIGHWAY = re.compile(r"\b(I- |Hwy |Highway |Loop |Rte |Route |Rd )\d") +SUFFIX = re.compile(r" (city and borough|city|town|village|borough|municipality|comunidad|zona urbana|metropolitan government|metro government|consolidated government|unified government|urban county|corporation|plantation)( \(balance\))?$") +# Places whose Census name is a merged government's; the postal city is what an address carries. +NAMES = {"1303440": "Athens", "1304204": "Augusta", "1349008": "Macon", "2148006": "Louisville", "3011397": "Butte", "4732742": "Hartsville", "4752006": "Nashville"} +TIMEZONES = { + "AK": "America/Anchorage", "AL": "America/Chicago", "AR": "America/Chicago", "AZ": "America/Phoenix", + "CA": "America/Los_Angeles", "CO": "America/Denver", "CT": "America/New_York", "DC": "America/New_York", + "DE": "America/New_York", "FL": "America/New_York", "GA": "America/New_York", "HI": "Pacific/Honolulu", + "IA": "America/Chicago", "ID": "America/Boise", "IL": "America/Chicago", "IN": "America/Indiana/Indianapolis", + "KS": "America/Chicago", "KY": "America/New_York", "LA": "America/Chicago", "MA": "America/New_York", + "MD": "America/New_York", "ME": "America/New_York", "MI": "America/Detroit", "MN": "America/Chicago", + "MO": "America/Chicago", "MS": "America/Chicago", "MT": "America/Denver", "NC": "America/New_York", + "ND": "America/Chicago", "NE": "America/Chicago", "NH": "America/New_York", "NJ": "America/New_York", + "NM": "America/Denver", "NV": "America/Los_Angeles", "NY": "America/New_York", "OH": "America/New_York", + "OK": "America/Chicago", "OR": "America/Los_Angeles", "PA": "America/New_York", "PR": "America/Puerto_Rico", + "RI": "America/New_York", "SC": "America/New_York", "SD": "America/Chicago", "TN": "America/Chicago", + "TX": "America/Chicago", "UT": "America/Denver", "VA": "America/New_York", "VT": "America/New_York", + "WA": "America/Los_Angeles", "WI": "America/Chicago", "WV": "America/New_York", "WY": "America/Denver", +} + + +def text(data): + try: + return data.decode("utf-8-sig") + except UnicodeDecodeError: + return data.decode("latin-1") + + +def gazetteer(cache, kind): + z = zipfile.ZipFile(io.BytesIO(tsv.fetch(GAZETTEER.format(kind), cache, f"gaz_{kind}.zip", magic=b"PK"))) + rows = text(z.read(z.namelist()[0])).splitlines() + header = [h.strip() for h in rows[0].split("|")] + return [dict(zip(header, (c.strip() for c in row.split("|")))) for row in rows[1:]] + + +def csv_rows(cache, url, name): + return list(csv.DictReader(io.StringIO(text(tsv.fetch(url, cache, name))))) + + +def dbf_rows(data, wanted): + """The records of a dBASE file, the wanted fields only.""" + count, header_len, record_len = struct.unpack("