Add the geo trees for SE and US as linked tables, and build each locale's address on them #20
@@ -1,3 +1,4 @@
|
||||
.claude
|
||||
*.out
|
||||
__pycache__/
|
||||
data-import/cache/
|
||||
|
||||
@@ -26,3 +26,9 @@ replacement, and each removed path, column or flag.
|
||||
`currency`, `flag`, `languages`, `numeric` and `tld` added; `misc.currency` the
|
||||
current ISO 4217 currencies with a minor unit, with `decimals` and `numeric`
|
||||
added, and its symbols from CLDR. `DATA-LICENSES.md` lists each source.
|
||||
- `geo.SE` and `geo.US`: five linked tables per country, `region`, `municipality`,
|
||||
`locality`, `postal-code` and `street`, weighted by population and address counts
|
||||
and built from SCB, GeoNames, Trafikverket NVDB and the US Census Bureau, and an
|
||||
`address` record over one consistent draw of them. `sv_SE.address` and
|
||||
`en_US.address` read those records, so `en_US.address.street` no longer carries
|
||||
`name` and `suffix`, and a locale folder loads only beside `geo`.
|
||||
|
||||
+7
-1
@@ -1,10 +1,16 @@
|
||||
# Data licenses
|
||||
|
||||
Every shipped dataset, its source, its licence and the attribution it asks for. A
|
||||
`data-import/` script rebuilds each sourced table; a curated one is hand-written.
|
||||
`data-import/` script rebuilds each sourced table, run as the README's
|
||||
[Development](README.md#development) section says; a curated one is hand-written.
|
||||
|
||||
| Table | Source | Licence | Attribution | Rebuild |
|
||||
|-------|--------|---------|-------------|---------|
|
||||
| `geo/SE/locality.tsv`, `postal-code.tsv` | [GeoNames](https://www.geonames.org/) postal codes for SE; populations from SCB tätorter 2023 | [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/); CC0 1.0 | "Postal codes from GeoNames, www.geonames.org" | `data-import/geo-se.py` |
|
||||
| `geo/SE/region.tsv`, `municipality.tsv` | [SCB](https://www.scb.se/) län and kommun codes 2026 and population 2024 | [CC0 1.0](https://creativecommons.org/publicdomain/zero/1.0/) | none required | `data-import/geo-se.py` |
|
||||
| `geo/SE/street.tsv` | [Trafikverket NVDB](https://www.trafikverket.se/) Gatunamn, through the open API | CC0 1.0 | none required | `data-import/geo-se.py` |
|
||||
| `geo/US/region.tsv`, `municipality.tsv`, `locality.tsv` | [Census Bureau](https://www.census.gov/) Gazetteer 2026 and population estimates 2025 | [public domain](https://www.usa.gov/government-works) | none required | `data-import/geo-us.py` |
|
||||
| `geo/US/postal-code.tsv`, `street.tsv` | Census Bureau ZCTA to place relationships 2020 and TIGER/Line 2025 address ranges and feature names | public domain | none required | `data-import/geo-us.py` |
|
||||
| `misc/country.tsv` | [datasets/country-codes](https://github.com/datasets/country-codes) | [PDDL 1.0](https://opendatacommons.org/licenses/pddl/1-0/) | none required | `data-import/country.py` |
|
||||
| `misc/currency.tsv` | [datasets/currency-codes](https://github.com/datasets/currency-codes); symbols from [Unicode CLDR](https://github.com/unicode-org/cldr) `en.xml` and `root.xml` | PDDL 1.0; [Unicode License v3](https://www.unicode.org/license.txt) | CLDR: "Copyright © 1991-2025 Unicode, Inc. Unicode and the Unicode Logo are registered trademarks of Unicode, Inc. in the United States and other countries." | `data-import/currency.py` |
|
||||
| `misc/httpstatus.tsv` | curated (IANA HTTP status codes are facts) | — | — | — |
|
||||
|
||||
@@ -18,6 +18,7 @@ fejkdata --seed 42 sv_SE.address # the same address every run
|
||||
fejkdata -n 3 --separator ', ' sv_SE.word # nät, barn, sol
|
||||
fejkdata --list # every path the data offers
|
||||
fejkdata 'misc.country[SE].capital' # Stockholm — a table's row, selected by key or name
|
||||
fejkdata 'geo.SE.locality[Lund].street' # Fjelievägen — a linked table, drawn inside the row
|
||||
fejkdata --data-path ./mydata sv_SE.word # layer a directory over the shipped data
|
||||
fejkdata --no-shipped-data -d ./mydata --list # only your data
|
||||
fejkdata 'name: {/sv_SE.person.last}' # name: <a surname> — an inline template
|
||||
@@ -227,6 +228,27 @@ Each locale carries `address`, `color`, `company`, `date`, `email`, `ip`,
|
||||
`misc.country[SE].capital` and `misc.currency[Euro].symbol` select a row;
|
||||
[`DATA-LICENSES.md`](DATA-LICENSES.md) names each table's source and licence.
|
||||
|
||||
A `geo` folder holds one tree per country under its alpha-2 code: five
|
||||
[linked tables](#linked-tables) named alike, and an `address` record over one
|
||||
consistent draw of them, which the locale's `address` reads.
|
||||
|
||||
| Table | `geo.SE` | `geo.US` | Weight |
|
||||
|-------|----------|----------|--------|
|
||||
| `region` | län, by code or name | state, by USPS abbreviation or name; `code` is the FIPS code | population |
|
||||
| `municipality` | kommun, by code or name | county, by FIPS code or name | population |
|
||||
| `locality` | postort, by name | incorporated place of 25,000 people or more with a postal code of its own, by GEOID or name; Hawaii has none | tätort population, the kommun's where the postort names it, else 200; place population |
|
||||
| `postal-code` | postnummer with street delivery, by code | ZCTA, by code | one; address ranges |
|
||||
| `street` | gatunamn, the ten with most road segments per postort | street name, the ten with most address ranges per place | segments; address ranges |
|
||||
|
||||
`geo.SE.region[Skåne län].municipality` draws a kommun in Skåne,
|
||||
`geo.SE.locality[Lund].street` a street in Lund, and
|
||||
`geo.US.region[IL].locality[Springfield]` settles which Springfield. A region row
|
||||
carries its `timezone`, the state's predominant zone, and a locality its `lat` and
|
||||
`lon`. What ports across countries is the five table names, the `name` column,
|
||||
selection by name, and the `address` record's columns `street`, `street-number`,
|
||||
`postal-code` and `locality`; every other column is the country's own, `code` on a
|
||||
Swedish region but `abbr` on a US one.
|
||||
|
||||
## Data format
|
||||
|
||||
Every value is a **node**, nestable without limit:
|
||||
@@ -983,6 +1005,41 @@ renamed or retyped line is a major.
|
||||
- **A path is walked once without drawing before it is walked for real.** A path
|
||||
that fails below its first level then moves no seeded stream, at the cost of one
|
||||
draw-free walk per call, which allocates nothing.
|
||||
- **A country's postal codes and streets are siblings under its locality.** No open
|
||||
source pairs a Swedish street with its postnummer, and pairing the US through its
|
||||
ZIPs would shape the two trees differently, so both draw inside the pinned
|
||||
locality and an address agrees at that level. A street's own code is the exact
|
||||
pairing to add when a source carries it.
|
||||
- **A locale's `address` reads its country's `geo` tree, so the shipped set loads
|
||||
whole.** `data/sv_SE` alone no longer loads: a test loads `data` and prefixes
|
||||
the locale, and `--no-shipped-data -d` takes the whole `data` folder or a set of
|
||||
one's own.
|
||||
- **The default embed holds every Swedish postort the import can place and give a
|
||||
street-delivery code and a street, and the US places of 25,000 or more.** Sweden
|
||||
fits whole in 700 KB; every US place of 10,000 would pass a
|
||||
megabyte and fetch 1,200 counties of TIGER files, so the threshold sits where the
|
||||
two countries match in size, and `--min-population` and
|
||||
`--streets-per-locality` on the import scripts build a fuller set. The two trees
|
||||
add about 20 ms to `New`, which loads the shipped set in about 45 ms.
|
||||
- **A locale's `address` restates its country record's format.** A record cannot
|
||||
read another whole and keep its columns, so `sv_SE.address` names the same four
|
||||
columns as `geo.SE.address`, each a reference into it, and the format appears
|
||||
twice; a column is spelled the same in both, `street-number`, so the two never
|
||||
disagree on a name.
|
||||
- **A postort's kommun comes from its name, its tätort or its codes, never from
|
||||
distance.** GeoNames leaves a fifth of Sweden's codes without a kommun and
|
||||
carries stale spellings; the nearest code across a border named the wrong kommun
|
||||
half the time it was tried, so a postort none of the three rules place is
|
||||
dropped, as is one not cased like a place name.
|
||||
- **A highway designation is not a street, and a US postal code belongs to the place
|
||||
holding most of its land inside places.** `I- 55 Bus` and `US Hwy 1` carry the
|
||||
most address ranges in many places and would head every address, so the import
|
||||
drops names spelled as a route. A ZCTA goes to the place its largest in-place part
|
||||
lies in, census-designated places left out since they never ship, and ships only
|
||||
when that place does, so a few dozen places whose every code lies mostly in a
|
||||
bigger neighbour ship no address; counting the land outside every place too would
|
||||
drop a quarter of the places, whose codes straddle unincorporated land, for a
|
||||
postal city the USPS mostly names the same way.
|
||||
- **`List` advertises direct descents only.** `region.municipality.locality` is
|
||||
listed, and `region.locality` resolves too but is not: the set of every descent
|
||||
through a chain of five tables is every subsequence of it, and the direct chain is
|
||||
@@ -1032,11 +1089,17 @@ REPIN=1 docker compose run --rm --user "$(id -u):$(id -g)" test
|
||||
|
||||
A shipped table built from a source is rebuilt by its script under
|
||||
[`data-import/`](data-import), one command per dataset, fetching the source named in
|
||||
[`DATA-LICENSES.md`](DATA-LICENSES.md):
|
||||
[`DATA-LICENSES.md`](DATA-LICENSES.md). Downloads are cached under
|
||||
`data-import/cache/`, so delete it to fetch afresh; `geo-us.py` fetches two
|
||||
TIGER/Line files per county it ships, a few hundred megabytes, and `geo-se.py` needs
|
||||
a Trafikverket API key, free at [data.trafikverket.se](https://data.trafikverket.se/),
|
||||
in `TRAFIKVERKET_API_KEY` or a `--key-file`:
|
||||
|
||||
```sh
|
||||
docker compose run --rm --user "$(id -u):$(id -g)" data-import data-import/country.py
|
||||
docker compose run --rm --user "$(id -u):$(id -g)" data-import data-import/currency.py
|
||||
docker compose run --rm --user "$(id -u):$(id -g)" data-import data-import/geo-us.py
|
||||
docker compose run --rm --user "$(id -u):$(id -g)" -e TRAFIKVERKET_API_KEY data-import data-import/geo-se.py
|
||||
```
|
||||
|
||||
To release, head `CHANGELOG.md` with the version's section in place of `Unreleased`
|
||||
@@ -1067,7 +1130,7 @@ datatype.go column datatypes: DataType, where datatype and null may sit, a c
|
||||
value.go the value proof: what a typed column or calc operand holds, checked at load
|
||||
data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge
|
||||
cmd/fejkdata/ the fejkdata CLI
|
||||
data/ shipped data (JSON, and a TSV per table), embedded at build: locale folders + a misc folder
|
||||
data/ shipped data (JSON, and a TSV per table), embedded at build: locale folders, geo, misc
|
||||
data-import/ the scripts that rebuild each sourced table (see DATA-LICENSES.md)
|
||||
release-tooling/ the release CI publishes from the changelog heading
|
||||
testdata/ the pinned shipped shape (see Versioning)
|
||||
|
||||
+6
-6
@@ -26,12 +26,12 @@ func benchPath(b *testing.B, dir, path string) {
|
||||
}
|
||||
}
|
||||
|
||||
func BenchmarkPerson(b *testing.B) { benchPath(b, "data/sv_SE", "person") }
|
||||
func BenchmarkAddress(b *testing.B) { benchPath(b, "data/sv_SE", "address") }
|
||||
func BenchmarkWord(b *testing.B) { benchPath(b, "data/sv_SE", "word") }
|
||||
func BenchmarkCreditcard(b *testing.B) { benchPath(b, "data/misc", "creditcard") }
|
||||
func BenchmarkSSN(b *testing.B) { benchPath(b, "data/sv_SE", "ssn") }
|
||||
func BenchmarkUUIDv7(b *testing.B) { benchPath(b, "data/misc", "uuid") }
|
||||
func BenchmarkPerson(b *testing.B) { benchPath(b, "data", "sv_SE.person") }
|
||||
func BenchmarkAddress(b *testing.B) { benchPath(b, "data", "sv_SE.address") }
|
||||
func BenchmarkWord(b *testing.B) { benchPath(b, "data", "sv_SE.word") }
|
||||
func BenchmarkCreditcard(b *testing.B) { benchPath(b, "data", "misc.creditcard") }
|
||||
func BenchmarkSSN(b *testing.B) { benchPath(b, "data", "sv_SE.ssn") }
|
||||
func BenchmarkUUIDv7(b *testing.B) { benchPath(b, "data", "misc.uuid") }
|
||||
|
||||
func tmpData(b *testing.B, name, body string) string {
|
||||
b.Helper()
|
||||
|
||||
@@ -14,6 +14,7 @@ import (
|
||||
const (
|
||||
svSE = "../../data/sv_SE"
|
||||
enUS = "../../data/en_US"
|
||||
misc = "../../data/misc"
|
||||
)
|
||||
|
||||
func runOut(args ...string) (int, string, string) {
|
||||
@@ -405,14 +406,14 @@ func TestRunTemplateMisuse(t *testing.T) {
|
||||
}
|
||||
|
||||
func TestRunNoShippedData(t *testing.T) {
|
||||
code, list, errb := runOut("--no-shipped-data", "-d", svSE, "--list")
|
||||
code, list, errb := runOut("--no-shipped-data", "-d", misc, "--list")
|
||||
if code != 0 {
|
||||
t.Fatalf("run = %d, stderr=%q", code, errb)
|
||||
}
|
||||
if strings.Contains(list, "en_US") || !strings.Contains(list, "person\n") {
|
||||
if strings.Contains(list, "sv_SE") || !strings.Contains(list, "uuid\n") {
|
||||
t.Errorf("--no-shipped-data --list = %q, want only the given dir", list)
|
||||
}
|
||||
code, out, _ := runOut("--no-shipped-data", "-d", svSE, "-s", "3", "person")
|
||||
code, out, _ := runOut("--no-shipped-data", "-d", misc, "-s", "3", "uuid")
|
||||
if code != 0 || strings.TrimSpace(out) == "" {
|
||||
t.Errorf("run = %d, out=%q", code, out)
|
||||
}
|
||||
|
||||
+7
-22
@@ -1,30 +1,24 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Rebuild data/misc/country.tsv from datasets/country-codes (PDDL).
|
||||
|
||||
data-import/country.py [--source URL_OR_FILE] [--out FILE]
|
||||
data-import/country.py [--source URL_OR_FILE] [--cache DIR] [--out FILE]
|
||||
"""
|
||||
import argparse
|
||||
import csv
|
||||
import io
|
||||
import re
|
||||
import sys
|
||||
import urllib.request
|
||||
from pathlib import Path
|
||||
|
||||
import tsv
|
||||
|
||||
SOURCE = "https://raw.githubusercontent.com/datasets/country-codes/main/data/country-codes.csv"
|
||||
OUT = Path(__file__).resolve().parent.parent / "data" / "misc" / "country.tsv"
|
||||
CACHE = Path(__file__).resolve().parent / "cache"
|
||||
COLUMNS = ["alpha2", "alpha3", "calling-code", "capital", "currency", "flag", "languages", "name", "numeric", "tld"]
|
||||
# Gaps in the source, keyed by alpha2.
|
||||
FIXUPS = {"TR": {"currency": "TRY"}}
|
||||
|
||||
|
||||
def read(source):
|
||||
if re.match(r"^https?://", source):
|
||||
with urllib.request.urlopen(source, timeout=60) as r:
|
||||
return r.read().decode("utf-8")
|
||||
return Path(source).read_text(encoding="utf-8")
|
||||
|
||||
|
||||
def flag(alpha2):
|
||||
return "".join(chr(0x1F1E6 + ord(c) - ord("A")) for c in alpha2)
|
||||
|
||||
@@ -58,23 +52,14 @@ def rows(text):
|
||||
yield row
|
||||
|
||||
|
||||
def write(out, table):
|
||||
lines = ["\t".join(COLUMNS)]
|
||||
for row in sorted(table, key=lambda r: r["alpha2"]):
|
||||
cells = [row[c] for c in COLUMNS]
|
||||
assert not any("\t" in c or "\n" in c for c in cells), row
|
||||
lines.append("\t".join(cells))
|
||||
Path(out).write_text("\n".join(lines) + "\n", encoding="utf-8")
|
||||
return len(lines) - 1
|
||||
|
||||
|
||||
def main():
|
||||
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
|
||||
p.add_argument("--cache", default=str(CACHE))
|
||||
p.add_argument("--source", default=SOURCE)
|
||||
p.add_argument("--out", default=str(OUT))
|
||||
a = p.parse_args()
|
||||
n = write(a.out, rows(read(a.source)))
|
||||
print(f"{a.out}: {n} rows", file=sys.stderr)
|
||||
table = rows(tsv.fetch(a.source, a.cache, "country-codes.csv").decode("utf-8"))
|
||||
tsv.write(a.out, COLUMNS, sorted(table, key=lambda r: r["alpha2"]))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
+8
-23
@@ -1,35 +1,28 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Rebuild data/misc/currency.tsv from datasets/currency-codes (PDDL) and CLDR's symbols (Unicode).
|
||||
|
||||
data-import/currency.py [--source URL_OR_FILE] [--symbols URL_OR_FILE ...] [--out FILE]
|
||||
data-import/currency.py [--source URL_OR_FILE] [--symbols URL_OR_FILE ...] [--cache DIR] [--out FILE]
|
||||
|
||||
The symbols come from the first locale file that has one, narrow symbols before wide.
|
||||
"""
|
||||
import argparse
|
||||
import csv
|
||||
import io
|
||||
import re
|
||||
import sys
|
||||
import urllib.request
|
||||
import xml.etree.ElementTree as ET
|
||||
from pathlib import Path
|
||||
|
||||
import tsv
|
||||
|
||||
SOURCE = "https://raw.githubusercontent.com/datasets/currency-codes/main/data/codes-all.csv"
|
||||
SYMBOLS = [
|
||||
"https://raw.githubusercontent.com/unicode-org/cldr/main/common/main/en.xml",
|
||||
"https://raw.githubusercontent.com/unicode-org/cldr/main/common/main/root.xml",
|
||||
]
|
||||
OUT = Path(__file__).resolve().parent.parent / "data" / "misc" / "currency.tsv"
|
||||
CACHE = Path(__file__).resolve().parent / "cache"
|
||||
COLUMNS = ["code", "decimals", "name", "numeric", "symbol"]
|
||||
|
||||
|
||||
def read(source):
|
||||
if re.match(r"^https?://", source):
|
||||
with urllib.request.urlopen(source, timeout=60) as r:
|
||||
return r.read().decode("utf-8")
|
||||
return Path(source).read_text(encoding="utf-8")
|
||||
|
||||
|
||||
def symbols(xml_texts):
|
||||
"""CLDR's symbol per code: the first locale's narrow symbol, else the first locale's wide one."""
|
||||
narrow, wide = {}, {}
|
||||
@@ -58,24 +51,16 @@ def rows(text, symbol):
|
||||
}
|
||||
|
||||
|
||||
def write(out, table):
|
||||
lines = ["\t".join(COLUMNS)]
|
||||
for row in sorted(table, key=lambda r: r["code"]):
|
||||
cells = [row[c] for c in COLUMNS]
|
||||
assert all(cells) and not any("\t" in c or "\n" in c for c in cells), row
|
||||
lines.append("\t".join(cells))
|
||||
Path(out).write_text("\n".join(lines) + "\n", encoding="utf-8")
|
||||
return len(lines) - 1
|
||||
|
||||
|
||||
def main():
|
||||
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
|
||||
p.add_argument("--cache", default=str(CACHE))
|
||||
p.add_argument("--source", default=SOURCE)
|
||||
p.add_argument("--symbols", nargs="+", default=SYMBOLS)
|
||||
p.add_argument("--out", default=str(OUT))
|
||||
a = p.parse_args()
|
||||
n = write(a.out, rows(read(a.source), symbols(read(s) for s in a.symbols)))
|
||||
print(f"{a.out}: {n} rows", file=sys.stderr)
|
||||
symbol = symbols(tsv.fetch(s, a.cache, Path(s).name).decode("utf-8") for s in a.symbols)
|
||||
table = rows(tsv.fetch(a.source, a.cache, "codes-all.csv").decode("utf-8"), symbol)
|
||||
tsv.write(a.out, COLUMNS, sorted(table, key=lambda r: r["code"]))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
@@ -0,0 +1,246 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Rebuild data/geo/SE/*.tsv from SCB (CC0), GeoNames (CC BY 4.0) and Trafikverket NVDB (CC0).
|
||||
|
||||
TRAFIKVERKET_API_KEY=… data-import/geo-se.py [--key-file FILE] [--cache DIR] [--streets-per-locality N] [--out DIR]
|
||||
"""
|
||||
import argparse
|
||||
import collections
|
||||
import csv
|
||||
import io
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
import urllib.request
|
||||
import xml.etree.ElementTree as ET
|
||||
import zipfile
|
||||
from pathlib import Path
|
||||
|
||||
import tsv
|
||||
|
||||
CODES = "https://www.scb.se/contentassets/7a89e48960f741e08918e489ea36354a/kommunlankod-2026.xlsx"
|
||||
POPULATION = "https://api.scb.se/OV0104/v1/doris/sv/ssd/START/BE/BE0101/BE0101A/BefolkningNy"
|
||||
POPULATION_QUERY = {
|
||||
"query": [
|
||||
{"code": "Region", "selection": {"filter": "all", "values": ["*"]}},
|
||||
{"code": "ContentsCode", "selection": {"filter": "item", "values": ["BE0101N1"]}},
|
||||
{"code": "Tid", "selection": {"filter": "top", "values": ["1"]}},
|
||||
],
|
||||
"response": {"format": "json"},
|
||||
}
|
||||
TATORTER = "https://geodata.scb.se/geoserver/stat/wfs?service=WFS&version=2.0.0&request=GetFeature&typeNames=stat:Tatorter_2023&outputFormat=csv&propertyName=tatort,kommun,bef"
|
||||
POSTAL_CODES = "https://download.geonames.org/export/zip/SE.zip"
|
||||
NVDB = "https://api.trafikinfo.trafikverket.se/v2/data.json"
|
||||
NVDB_PAGE = 50000
|
||||
OUT = Path(__file__).resolve().parent.parent / "data" / "geo" / "SE"
|
||||
CACHE = Path(__file__).resolve().parent / "cache"
|
||||
TIMEZONE = "Europe/Stockholm"
|
||||
ONE_POSITION = {"Stockholm", "Göteborg", "Malmö"}
|
||||
UNMATCHED_POPULATION = 200
|
||||
XLSX_NS = {"m": "http://schemas.openxmlformats.org/spreadsheetml/2006/main"}
|
||||
|
||||
|
||||
def xlsx_rows(data):
|
||||
z = zipfile.ZipFile(io.BytesIO(data))
|
||||
strings = ["".join(t.text or "" for t in si.iter("{%s}t" % XLSX_NS["m"])) for si in ET.fromstring(z.read("xl/sharedStrings.xml")).findall("m:si", XLSX_NS)]
|
||||
sheet = ET.fromstring(z.read("xl/worksheets/sheet1.xml"))
|
||||
for row in sheet.findall(".//m:row", XLSX_NS):
|
||||
cells = []
|
||||
for c in row.findall("m:c", XLSX_NS):
|
||||
v = c.find("m:v", XLSX_NS)
|
||||
cells.append("" if v is None else strings[int(v.text)] if c.get("t") == "s" else v.text)
|
||||
yield cells
|
||||
|
||||
|
||||
def scb_codes(cache):
|
||||
regions, municipalities = {}, {}
|
||||
for cells in xlsx_rows(tsv.fetch(CODES, cache, "kommunlankod.xlsx", magic=b"PK")):
|
||||
if len(cells) < 2 or not re.fullmatch(r"\d{2}|\d{4}", cells[0]):
|
||||
continue
|
||||
(regions if len(cells[0]) == 2 else municipalities)[cells[0]] = cells[1].strip()
|
||||
return regions, municipalities
|
||||
|
||||
|
||||
def scb_population(cache):
|
||||
body = json.dumps(POPULATION_QUERY).encode()
|
||||
data = tsv.fetch(POPULATION, cache, "befolkning.json", data=body, headers={"Content-Type": "application/json"})
|
||||
return {row["key"][0]: row["values"][0] for row in json.loads(data.decode("utf-8-sig"))["data"]}
|
||||
|
||||
|
||||
def scb_tatorter(cache):
|
||||
text = tsv.fetch(TATORTER, cache, "tatorter.csv").decode("utf-8")
|
||||
by_name = collections.defaultdict(list)
|
||||
for r in csv.DictReader(io.StringIO(text)):
|
||||
by_name[r["tatort"]].append((r["kommun"], int(r["bef"])))
|
||||
return by_name
|
||||
|
||||
|
||||
def geonames(cache):
|
||||
z = zipfile.ZipFile(io.BytesIO(tsv.fetch(POSTAL_CODES, cache, "SE.zip", magic=b"PK")))
|
||||
rows = []
|
||||
for line in z.read("SE.txt").decode("utf-8").splitlines():
|
||||
f = line.split("\t")
|
||||
lat, lon = (float(f[9]), float(f[10])) if f[9] and f[10] else (None, None)
|
||||
rows.append({"code": f[1], "locality": f[2], "municipality": f[6], "lat": lat, "lon": lon})
|
||||
return rows
|
||||
|
||||
|
||||
def nvdb_segments(cache, key):
|
||||
path = cache / "nvdb-gatunamn.tsv"
|
||||
if not path.exists():
|
||||
with open(path.with_suffix(".part"), "w", encoding="utf-8") as out:
|
||||
change = "0"
|
||||
while True:
|
||||
query = (
|
||||
f'<REQUEST><LOGIN authenticationkey="{key}"/>'
|
||||
f'<QUERY objecttype="Gatunamn" namespace="vägdata.nvdb_dk_o" schemaversion="1.0" limit="{NVDB_PAGE}" changeid="{change}">'
|
||||
"<FILTER><EQ name=\"Deleted\" value=\"false\"/></FILTER>"
|
||||
"<INCLUDE>Namn</INCLUDE><INCLUDE>Geometry.WKT-WGS84-3D</INCLUDE></QUERY></REQUEST>"
|
||||
)
|
||||
req = urllib.request.Request(NVDB, data=query.encode(), headers={"Content-Type": "text/xml"})
|
||||
with urllib.request.urlopen(req, timeout=600) as r:
|
||||
result = json.load(r)["RESPONSE"]["RESULT"][0]
|
||||
rows = result.get("Gatunamn", [])
|
||||
for row in rows:
|
||||
m = re.match(r"LINESTRING Z \(([-\d.]+) ([-\d.]+) ", row.get("Geometry", {}).get("WKT-WGS84-3D", ""))
|
||||
name = " ".join(row.get("Namn", "").split())
|
||||
if m and name:
|
||||
out.write(f"{name}\t{m.group(1)}\t{m.group(2)}\n")
|
||||
change = result["INFO"]["LASTCHANGEID"]
|
||||
if len(rows) < NVDB_PAGE:
|
||||
break
|
||||
path.with_suffix(".part").rename(path)
|
||||
for line in path.read_text(encoding="utf-8").splitlines():
|
||||
name, lon, lat = line.split("\t")
|
||||
yield name, float(lat), float(lon)
|
||||
|
||||
|
||||
class Nearest:
|
||||
"""Nearest point by an equirectangular distance, over a degree grid."""
|
||||
|
||||
def __init__(self, points, cell=0.05):
|
||||
self.cell = cell
|
||||
self.grid = collections.defaultdict(list)
|
||||
for lat, lon, value in points:
|
||||
self.grid[(int(lat // cell), int(lon // cell))].append((lat, lon, value))
|
||||
|
||||
def find(self, lat, lon):
|
||||
ci, cj = int(lat // self.cell), int(lon // self.cell)
|
||||
best, best_d = None, math.inf
|
||||
ring = 0
|
||||
while ring < 400:
|
||||
for i in range(ci - ring, ci + ring + 1):
|
||||
for j in range(cj - ring, cj + ring + 1):
|
||||
if max(abs(i - ci), abs(j - cj)) != ring:
|
||||
continue
|
||||
for plat, plon, value in self.grid.get((i, j), ()):
|
||||
d = (plat - lat) ** 2 + ((plon - lon) * math.cos(math.radians(lat))) ** 2
|
||||
if d < best_d:
|
||||
best, best_d = value, d
|
||||
if best is not None and math.sqrt(best_d) < ring * self.cell * math.cos(math.radians(lat)):
|
||||
return best
|
||||
ring += 1
|
||||
return best
|
||||
|
||||
|
||||
def street_delivery(name, codes):
|
||||
"""The codes delivered to a street: the digit after the postort's own prefix says box, company or reply."""
|
||||
if name in ONE_POSITION:
|
||||
return [c for c in codes if c[1] != "0"]
|
||||
largest = collections.Counter(c[:3] for c in codes).most_common(1)[0][1]
|
||||
length = 3 if largest * 2 >= len(codes) else 2
|
||||
return [c for c in codes if c[length] not in "018"]
|
||||
|
||||
|
||||
def municipality_of(name, rows, tatorter, municipalities):
|
||||
named = [code for code, n in municipalities.items() if n == name]
|
||||
if named:
|
||||
return named[0], "kommun"
|
||||
voted = collections.Counter(r["municipality"] for r in rows if r["municipality"] in municipalities)
|
||||
matches = tatorter.get(name, [])
|
||||
if matches:
|
||||
in_vote = [m for m in matches if voted and m[0] == voted.most_common(1)[0][0]]
|
||||
return max(in_vote or matches, key=lambda m: m[1])[0], "tatort"
|
||||
if voted:
|
||||
return voted.most_common(1)[0][0], "codes"
|
||||
return None, "unplaced"
|
||||
|
||||
|
||||
def population_of(name, municipality, tatorter, municipalities, population):
|
||||
matches = [m for m in tatorter.get(name, []) if m[0] == municipality]
|
||||
if matches:
|
||||
return str(max(m[1] for m in matches))
|
||||
return population[municipality] if municipalities[municipality] == name else str(UNMATCHED_POPULATION)
|
||||
|
||||
|
||||
def well_cased(name):
|
||||
return all(part[:1].isupper() and (len(part) == 1 or not part.isupper()) for part in re.split(r"[ -]", name))
|
||||
|
||||
|
||||
def localities(codes, tatorter, municipalities, population):
|
||||
"""Each postort with its municipality, weight, centroid and street-delivery codes."""
|
||||
by_locality = collections.defaultdict(list)
|
||||
for r in codes:
|
||||
by_locality[r["locality"]].append(r)
|
||||
out, how = {}, collections.Counter()
|
||||
for name, rows in by_locality.items():
|
||||
municipality, method = municipality_of(name, rows, tatorter, municipalities)
|
||||
how[method] += 1
|
||||
kept = street_delivery(name, [r["code"].replace(" ", "") for r in rows])
|
||||
with_point = [r for r in rows if r["lat"] is not None]
|
||||
if municipality is None or not kept or not with_point or not well_cased(name):
|
||||
continue
|
||||
lat = sum(r["lat"] for r in with_point) / len(with_point)
|
||||
lon = sum(r["lon"] for r in with_point) / len(with_point)
|
||||
out[name] = {"name": name, "municipality": municipality, "population": population_of(name, municipality, tatorter, municipalities, population), "lat": f"{lat:.4f}", "lon": f"{lon:.4f}", "codes": kept}
|
||||
print(f"municipality by {dict(how)}; {len(by_locality) - len(out)} postorter dropped", file=sys.stderr)
|
||||
return out
|
||||
|
||||
|
||||
def streets(segments, codes, localities, per_locality):
|
||||
"""The names with most segments per locality, each segment at its nearest code centroid."""
|
||||
nearest = Nearest((r["lat"], r["lon"], r["locality"]) for r in codes if r["lat"] is not None and r["locality"] in localities)
|
||||
count = collections.Counter()
|
||||
for name, lat, lon in segments:
|
||||
if name[0].isalpha():
|
||||
count[(nearest.find(lat, lon), name)] += 1
|
||||
of = collections.defaultdict(list)
|
||||
for (locality, name), n in count.items():
|
||||
of[locality].append((n, name))
|
||||
return {locality: [{"name": name, "locality": locality, "segments": n} for n, name in sorted(named, key=lambda s: (-s[0], s[1]))[:per_locality]] for locality, named in of.items()}
|
||||
|
||||
|
||||
def main():
|
||||
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
|
||||
p.add_argument("--cache", default=str(CACHE))
|
||||
p.add_argument("--key-file", help="file holding the Trafikverket API key; TRAFIKVERKET_API_KEY otherwise")
|
||||
p.add_argument("--out", default=str(OUT))
|
||||
p.add_argument("--streets-per-locality", type=int, default=10)
|
||||
a = p.parse_args()
|
||||
key = Path(a.key_file).read_text().strip() if a.key_file else os.environ.get("TRAFIKVERKET_API_KEY")
|
||||
if not key:
|
||||
sys.exit("set TRAFIKVERKET_API_KEY or pass --key-file")
|
||||
cache, out = Path(a.cache), Path(a.out)
|
||||
cache.mkdir(parents=True, exist_ok=True)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
regions, municipalities = scb_codes(cache)
|
||||
population = scb_population(cache)
|
||||
codes = geonames(cache)
|
||||
places = localities(codes, scb_tatorter(cache), municipalities, population)
|
||||
named = streets(nvdb_segments(cache, key), codes, places, a.streets_per_locality)
|
||||
places = {name: l for name, l in places.items() if name in named}
|
||||
empty = sorted(m for m in municipalities if not any(l["municipality"] == m for l in places.values()))
|
||||
if empty:
|
||||
sys.exit(f"municipalities without a locality: {empty}")
|
||||
|
||||
tsv.write(out / "region.tsv", ["code", "name", "population", "timezone"], [{"code": c, "name": n, "population": population[c], "timezone": TIMEZONE} for c, n in sorted(regions.items())])
|
||||
tsv.write(out / "municipality.tsv", ["code", "name", "region", "population"], [{"code": c, "name": n, "region": c[:2], "population": population[c]} for c, n in sorted(municipalities.items())])
|
||||
tsv.write(out / "locality.tsv", ["name", "municipality", "population", "lat", "lon"], [l for _, l in sorted(places.items())])
|
||||
tsv.write(out / "postal-code.tsv", ["code", "locality"], sorted(({"code": f"{c[:3]} {c[3:]}", "locality": l["name"]} for l in places.values() for c in l["codes"]), key=lambda r: r["code"]))
|
||||
tsv.write(out / "street.tsv", ["name", "locality", "segments"], [s for locality in sorted(named) for s in named[locality]])
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,201 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Rebuild data/geo/US/*.tsv from the Census Bureau's Gazetteer, population estimates, ZCTA relationships and TIGER/Line files (public domain).
|
||||
|
||||
data-import/geo-us.py [--cache DIR] [--min-population N] [--streets-per-locality N] [--out DIR]
|
||||
"""
|
||||
import argparse
|
||||
import collections
|
||||
import concurrent.futures
|
||||
import csv
|
||||
import io
|
||||
import re
|
||||
import struct
|
||||
import sys
|
||||
import zipfile
|
||||
from pathlib import Path
|
||||
|
||||
import tsv
|
||||
|
||||
GAZETTEER = "https://www2.census.gov/geo/docs/maps-data/data/gazetteer/2026_Gazetteer/2026_Gaz_{}_national.zip"
|
||||
POPULATION = "https://www2.census.gov/programs-surveys/popest/datasets/2020-2025/{}"
|
||||
STATES = POPULATION.format("state/totals/NST-EST2025-ALLDATA.csv")
|
||||
COUNTIES = POPULATION.format("counties/totals/co-est2025-alldata.csv")
|
||||
PLACES = POPULATION.format("cities/totals/sub-est2025.csv")
|
||||
ZCTA_PLACE = "https://www2.census.gov/geo/docs/maps-data/data/rel2020/zcta520/tab20_zcta520_place20_natl.txt"
|
||||
TIGER = "https://www2.census.gov/geo/tiger/TIGER2025/{0}/tl_2025_{1}_{2}.zip"
|
||||
OUT = Path(__file__).resolve().parent.parent / "data" / "geo" / "US"
|
||||
CACHE = Path(__file__).resolve().parent / "cache"
|
||||
ESTIMATE = "POPESTIMATE2025"
|
||||
CDP = "57"
|
||||
HIGHWAY = re.compile(r"\b(I- |Hwy |Highway |Loop |Rte |Route |Rd )\d")
|
||||
SUFFIX = re.compile(r" (city and borough|city|town|village|borough|municipality|comunidad|zona urbana|metropolitan government|metro government|consolidated government|unified government|urban county|corporation|plantation)( \(balance\))?$")
|
||||
# Places whose Census name is a merged government's; the postal city is what an address carries.
|
||||
NAMES = {"1303440": "Athens", "1304204": "Augusta", "1349008": "Macon", "2148006": "Louisville", "3011397": "Butte", "4732742": "Hartsville", "4752006": "Nashville"}
|
||||
TIMEZONES = {
|
||||
"AK": "America/Anchorage", "AL": "America/Chicago", "AR": "America/Chicago", "AZ": "America/Phoenix",
|
||||
"CA": "America/Los_Angeles", "CO": "America/Denver", "CT": "America/New_York", "DC": "America/New_York",
|
||||
"DE": "America/New_York", "FL": "America/New_York", "GA": "America/New_York", "HI": "Pacific/Honolulu",
|
||||
"IA": "America/Chicago", "ID": "America/Boise", "IL": "America/Chicago", "IN": "America/Indiana/Indianapolis",
|
||||
"KS": "America/Chicago", "KY": "America/New_York", "LA": "America/Chicago", "MA": "America/New_York",
|
||||
"MD": "America/New_York", "ME": "America/New_York", "MI": "America/Detroit", "MN": "America/Chicago",
|
||||
"MO": "America/Chicago", "MS": "America/Chicago", "MT": "America/Denver", "NC": "America/New_York",
|
||||
"ND": "America/Chicago", "NE": "America/Chicago", "NH": "America/New_York", "NJ": "America/New_York",
|
||||
"NM": "America/Denver", "NV": "America/Los_Angeles", "NY": "America/New_York", "OH": "America/New_York",
|
||||
"OK": "America/Chicago", "OR": "America/Los_Angeles", "PA": "America/New_York", "PR": "America/Puerto_Rico",
|
||||
"RI": "America/New_York", "SC": "America/New_York", "SD": "America/Chicago", "TN": "America/Chicago",
|
||||
"TX": "America/Chicago", "UT": "America/Denver", "VA": "America/New_York", "VT": "America/New_York",
|
||||
"WA": "America/Los_Angeles", "WI": "America/Chicago", "WV": "America/New_York", "WY": "America/Denver",
|
||||
}
|
||||
|
||||
|
||||
def text(data):
|
||||
try:
|
||||
return data.decode("utf-8-sig")
|
||||
except UnicodeDecodeError:
|
||||
return data.decode("latin-1")
|
||||
|
||||
|
||||
def gazetteer(cache, kind):
|
||||
z = zipfile.ZipFile(io.BytesIO(tsv.fetch(GAZETTEER.format(kind), cache, f"gaz_{kind}.zip", magic=b"PK")))
|
||||
rows = text(z.read(z.namelist()[0])).splitlines()
|
||||
header = [h.strip() for h in rows[0].split("|")]
|
||||
return [dict(zip(header, (c.strip() for c in row.split("|")))) for row in rows[1:]]
|
||||
|
||||
|
||||
def csv_rows(cache, url, name):
|
||||
return list(csv.DictReader(io.StringIO(text(tsv.fetch(url, cache, name)))))
|
||||
|
||||
|
||||
def dbf_rows(data, wanted):
|
||||
"""The records of a dBASE file, the wanted fields only."""
|
||||
count, header_len, record_len = struct.unpack("<xxxxIHH", data[:12])
|
||||
fields, pos = [], 32
|
||||
while data[pos] != 0x0D:
|
||||
name = data[pos:pos + 11].split(b"\0")[0].decode()
|
||||
fields.append((name, data[pos + 16]))
|
||||
pos += 32
|
||||
pos = header_len
|
||||
for _ in range(count):
|
||||
record = data[pos:pos + record_len]
|
||||
pos += record_len
|
||||
if record[:1] == b"*":
|
||||
continue
|
||||
row, at = {}, 1
|
||||
for name, length in fields:
|
||||
if name in wanted:
|
||||
row[name] = record[at:at + length].decode("utf-8", "replace").strip()
|
||||
at += length
|
||||
yield row
|
||||
|
||||
|
||||
def tiger_zip(cache, kind, county):
|
||||
return tsv.fetch(TIGER.format(kind.upper(), county, kind), cache, f"tl_{county}_{kind}.zip", magic=b"PK")
|
||||
|
||||
|
||||
def tiger(cache, kind, county, wanted):
|
||||
z = zipfile.ZipFile(io.BytesIO(tiger_zip(cache, kind, county)))
|
||||
return dbf_rows(z.read(f"tl_2025_{county}_{kind}.dbf"), wanted)
|
||||
|
||||
|
||||
def place_name(geoid, name):
|
||||
if geoid in NAMES:
|
||||
return NAMES[geoid]
|
||||
stripped = SUFFIX.sub("", name)
|
||||
if stripped == name:
|
||||
print(f"{geoid}: no suffix stripped from {name!r}", file=sys.stderr)
|
||||
return stripped
|
||||
|
||||
|
||||
def localities(cache, min_population, counties):
|
||||
"""Each shipped place with its county, population and centroid."""
|
||||
place_population, county_part = {}, collections.defaultdict(list)
|
||||
for r in csv_rows(cache, PLACES, "sub-est2025.csv"):
|
||||
if r["SUMLEV"] == "162":
|
||||
place_population[r["STATE"] + r["PLACE"]] = int(r[ESTIMATE])
|
||||
elif r["SUMLEV"] == "157":
|
||||
county_part[r["STATE"] + r["PLACE"]].append((int(r[ESTIMATE]), r["STATE"] + r["COUNTY"]))
|
||||
out = {}
|
||||
for r in gazetteer(cache, "place"):
|
||||
geoid, population = r["GEOID"], place_population.get(r["GEOID"], 0)
|
||||
if r["FUNCSTAT"] not in ("A", "F", "N") or r["LSAD"] == CDP or population < min_population or not county_part.get(geoid):
|
||||
continue
|
||||
county = max(county_part[geoid])[1]
|
||||
if county not in counties:
|
||||
print(f"{geoid} {r['NAME']}: county {county} unknown, dropped", file=sys.stderr)
|
||||
continue
|
||||
out[geoid] = {"code": geoid, "name": place_name(geoid, r["NAME"]), "municipality": county, "population": population, "lat": r["INTPTLAT"], "lon": r["INTPTLONG"]}
|
||||
return out
|
||||
|
||||
|
||||
def postal_codes(cache, localities):
|
||||
"""Each ZCTA whose largest part inside an incorporated place lies in a shipped place."""
|
||||
parts = {}
|
||||
for r in csv.DictReader(io.StringIO(text(tsv.fetch(ZCTA_PLACE, cache, "zcta-place.txt"))), delimiter="|"):
|
||||
if r["GEOID_ZCTA5_20"] and r["GEOID_PLACE_20"] and not r["NAMELSAD_PLACE_20"].endswith(" CDP"):
|
||||
parts.setdefault(r["GEOID_ZCTA5_20"], []).append((int(r["AREALAND_PART"]), r["GEOID_PLACE_20"]))
|
||||
largest = {zcta: max(p)[1] for zcta, p in parts.items()}
|
||||
return {zcta: place for zcta, place in largest.items() if place in localities}
|
||||
|
||||
|
||||
def streets(cache, counties, locality_of_zcta, per_locality):
|
||||
"""The names with most TIGER address ranges per place, and the address ranges per ZCTA."""
|
||||
with concurrent.futures.ThreadPoolExecutor(3) as pool:
|
||||
list(pool.map(lambda c: (tiger_zip(cache, "addr", c), tiger_zip(cache, "featnames", c)), counties))
|
||||
addresses, count = collections.Counter(), collections.Counter()
|
||||
for county in counties:
|
||||
zips = collections.defaultdict(set)
|
||||
for r in tiger(cache, "addr", county, {"TLID", "ZIP"}):
|
||||
if r["ZIP"] in locality_of_zcta:
|
||||
zips[r["TLID"]].add(r["ZIP"])
|
||||
addresses[r["ZIP"]] += 1
|
||||
for r in tiger(cache, "featnames", county, {"TLID", "FULLNAME", "PAFLAG"}):
|
||||
if r["PAFLAG"] == "P" and r["FULLNAME"] and not HIGHWAY.search(r["FULLNAME"]):
|
||||
for z in zips.get(r["TLID"], ()):
|
||||
count[(locality_of_zcta[z], r["FULLNAME"])] += 1
|
||||
of = collections.defaultdict(list)
|
||||
for (locality, name), n in count.items():
|
||||
of[locality].append((n, name))
|
||||
named = {locality: [{"name": name, "locality": locality, "addresses": n} for n, name in sorted(ranked, key=lambda s: (-s[0], s[1]))[:per_locality]] for locality, ranked in of.items()}
|
||||
return named, addresses
|
||||
|
||||
|
||||
def main():
|
||||
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
|
||||
p.add_argument("--cache", default=str(CACHE))
|
||||
p.add_argument("--min-population", type=int, default=25000)
|
||||
p.add_argument("--out", default=str(OUT))
|
||||
p.add_argument("--streets-per-locality", type=int, default=10)
|
||||
a = p.parse_args()
|
||||
cache, out = Path(a.cache), Path(a.out)
|
||||
cache.mkdir(parents=True, exist_ok=True)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
state_population = {r["STATE"]: r[ESTIMATE] for r in csv_rows(cache, STATES, "nst-est2025.csv") if r["SUMLEV"] == "040"}
|
||||
regions = {r["USPS"]: {"abbr": r["USPS"], "code": r["GEOID"], "name": r["NAME"], "population": state_population[r["GEOID"]], "timezone": TIMEZONES[r["USPS"]]} for r in gazetteer(cache, "state")}
|
||||
county_population = {r["STATE"] + r["COUNTY"]: r[ESTIMATE] for r in csv_rows(cache, COUNTIES, "co-est2025.csv") if r["SUMLEV"] == "050"}
|
||||
counties, unestimated = {}, collections.Counter()
|
||||
for r in gazetteer(cache, "counties"):
|
||||
if r["GEOID"] not in county_population:
|
||||
unestimated[r["USPS"]] += 1
|
||||
continue
|
||||
counties[r["GEOID"]] = {"code": r["GEOID"], "name": r["NAME"], "region": r["USPS"], "population": county_population[r["GEOID"]]}
|
||||
print(f"counties without a population estimate, dropped: {dict(unestimated)}", file=sys.stderr)
|
||||
|
||||
places = localities(cache, a.min_population, counties)
|
||||
locality_of_zcta = postal_codes(cache, places)
|
||||
named, addresses = streets(cache, sorted({l["municipality"] for l in places.values()}), locality_of_zcta, a.streets_per_locality)
|
||||
for geoid in [l for l in places if l not in named]:
|
||||
print(f"{geoid} {places[geoid]['name']}: no streets, dropped", file=sys.stderr)
|
||||
del places[geoid]
|
||||
kept_counties = {l["municipality"] for l in places.values()}
|
||||
kept_regions = {counties[c]["region"] for c in kept_counties}
|
||||
|
||||
tsv.write(out / "region.tsv", ["abbr", "code", "name", "population", "timezone"], [r for _, r in sorted(regions.items()) if r["abbr"] in kept_regions])
|
||||
tsv.write(out / "municipality.tsv", ["code", "name", "region", "population"], [c for _, c in sorted(counties.items()) if c["code"] in kept_counties])
|
||||
tsv.write(out / "locality.tsv", ["code", "name", "municipality", "population", "lat", "lon"], [l for _, l in sorted(places.items())])
|
||||
tsv.write(out / "postal-code.tsv", ["code", "locality", "addresses"], [{"code": z, "locality": l, "addresses": addresses[z]} for z, l in sorted(locality_of_zcta.items()) if addresses[z] and l in places])
|
||||
tsv.write(out / "street.tsv", ["name", "locality", "addresses"], [s for locality in sorted(named) for s in named[locality]])
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,40 @@
|
||||
"""A source fetched once into the cache, and a table written as the loader admits it."""
|
||||
import re
|
||||
import sys
|
||||
import time
|
||||
import urllib.request
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def fetch(source, cache, name, magic=b"", data=None, headers=None):
|
||||
"""The bytes of a URL, downloaded into cache/name once, or of a local file."""
|
||||
if not re.match(r"^https?://", source):
|
||||
return Path(source).read_bytes()
|
||||
path = Path(cache) / name
|
||||
for attempt in range(1, 6):
|
||||
if path.exists():
|
||||
return path.read_bytes()
|
||||
req = urllib.request.Request(source, data=data, headers={"User-Agent": "fejkdata data-import", **(headers or {})})
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=600) as r:
|
||||
body = r.read()
|
||||
except OSError:
|
||||
body = b""
|
||||
if body and body.startswith(magic) and b"Request Rejected" not in body[:512]:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_bytes(body)
|
||||
elif attempt < 5:
|
||||
time.sleep(10 * attempt)
|
||||
sys.exit(f"{source}: no valid download in 5 attempts")
|
||||
|
||||
|
||||
def write(path, columns, rows):
|
||||
"""Write the rows as a TSV; every cell must be non-empty and free of tabs, newlines and braces."""
|
||||
lines = ["\t".join(columns)]
|
||||
for row in rows:
|
||||
cells = [str(row[c]) for c in columns]
|
||||
if not all(cells) or any(re.search(r"[\t\n{}]", c) for c in cells):
|
||||
raise ValueError(f"{path}: a cell is empty or holds a tab, newline or brace: {row}")
|
||||
lines.append("\t".join(cells))
|
||||
Path(path).write_text("\n".join(lines) + "\n", encoding="utf-8")
|
||||
print(f"{path}: {len(lines) - 1} rows", file=sys.stderr)
|
||||
+5
-14
@@ -1,17 +1,8 @@
|
||||
{
|
||||
"format": "{street-number} {street}\n{locality}, {region} {postal-code}",
|
||||
"street": {
|
||||
"format": "{name} {suffix}",
|
||||
"name": ["Adams", "Ashby", "Aspen", "Bay", "Birch", "Bridge", "Cedar", "Chestnut", "Church", "Clark", "Cypress", "Dogwood", "Elm", "Forest", "Franklin", "Garden", "Grove", "Hawthorn", "Hickory", "Highland", "Jackson", "Jefferson", "Juniper", "Lake", "Laurel", "Liberty", "Lincoln", "Madison", "Magnolia", "Maple", "Market", "Meadow", "Mill", "Oak", "Park", "Pine", "Poplar", "Prospect", "Ridge", "River", "Spruce", "Sunset", "Sycamore", "Union", "Walnut", "Washington", "Willow", "Wilson"],
|
||||
"suffix": ["Avenue", "Boulevard", "Circle", "Court", "Drive", "Lane", "Place", "Road", "Street", "Terrace", "Trail", "Way"]
|
||||
},
|
||||
"street-number": [
|
||||
"{int(10,99)}",
|
||||
{ "format": "{int(100,999)}", "weight": 2 },
|
||||
{ "format": "{int(1000,9999)}", "weight": 0.5 },
|
||||
{ "format": "{int(10,99)}{int(100,999)}", "weight": 0.3 }
|
||||
],
|
||||
"locality": ["Albany", "Atlanta", "Austin", "Baltimore", "Boston", "Charlotte", "Chicago", "Cincinnati", "Cleveland", "Columbus", "Dallas", "Denver", "Detroit", "El Paso", "Fort Worth", "Fresno", "Houston", "Indianapolis", "Jacksonville", "Kansas City", "Las Vegas", "Long Beach", "Los Angeles", "Memphis", "Mesa", "Miami", "Milwaukee", "Minneapolis", "Nashville", "New Orleans", "Oakland", "Oklahoma City", "Omaha", "Orlando", "Philadelphia", "Phoenix", "Pittsburgh", "Portland", "Raleigh", "Sacramento", "San Antonio", "San Diego", "San Jose", "Seattle", "St. Louis", "Tampa", "Tucson", "Tulsa"],
|
||||
"region": ["AL", "AZ", "CA", "CO", "CT", "FL", "GA", "IL", "IN", "KY", "LA", "MA", "MD", "MI", "MN", "MO", "NC", "NJ", "NV", "NY", "OH", "OK", "OR", "PA", "TN", "TX", "VA", "WA", "WI"],
|
||||
"postal-code": ["{int(10000,99999)}", { "format": "{int(10000,99999)}-{digits(4)}", "weight": 0.3 }]
|
||||
"street": "{/geo.US.address.street}",
|
||||
"street-number": "{/geo.US.address.street-number}",
|
||||
"locality": "{/geo.US.address.locality}",
|
||||
"region": "{/geo.US.address.region}",
|
||||
"postal-code": "{/geo.US.address.postal-code}"
|
||||
}
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
{
|
||||
"format": "{street} {street-number}\n{postal-code} {locality}",
|
||||
"street": "{.street.name}",
|
||||
"street-number": [
|
||||
"{int(1,9)}",
|
||||
"{int(10,99)}",
|
||||
{ "format": "{int(100,999)}", "weight": 0.2 },
|
||||
{ "format": "{int(1,9)}{upper(1)}", "weight": 0.2 },
|
||||
{ "format": "{int(10,99)}{upper(1)}", "weight": 0.1 },
|
||||
{ "format": "{int(100,999)}{upper(1)}", "weight": 0.05 }
|
||||
],
|
||||
"postal-code": "{.postal-code.code}",
|
||||
"locality": "{.locality.name}"
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
{ "format": "{name}", "rows": "locality.tsv", "key": "name", "parent": "municipality", "weight": "population" }
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1 @@
|
||||
{ "format": "{name}", "rows": "municipality.tsv", "key": "code", "name": "name", "parent": "region", "weight": "population" }
|
||||
@@ -0,0 +1,291 @@
|
||||
code name region population
|
||||
0114 Upplands Väsby 01 50323
|
||||
0115 Vallentuna 01 35119
|
||||
0117 Österåker 01 49787
|
||||
0120 Värmdö 01 46635
|
||||
0123 Järfälla 01 88950
|
||||
0125 Ekerö 01 28910
|
||||
0126 Huddinge 01 114304
|
||||
0127 Botkyrka 01 95905
|
||||
0128 Salem 01 17507
|
||||
0136 Haninge 01 100895
|
||||
0138 Tyresö 01 49179
|
||||
0139 Upplands-Bro 01 32868
|
||||
0140 Nykvarn 01 12342
|
||||
0160 Täby 01 77744
|
||||
0162 Danderyd 01 32425
|
||||
0163 Sollentuna 01 77624
|
||||
0180 Stockholm 01 995574
|
||||
0181 Södertälje 01 102911
|
||||
0182 Nacka 01 112112
|
||||
0183 Sundbyberg 01 56274
|
||||
0184 Solna 01 85789
|
||||
0186 Lidingö 01 48377
|
||||
0187 Vaxholm 01 11822
|
||||
0188 Norrtälje 01 66585
|
||||
0191 Sigtuna 01 52767
|
||||
0192 Nynäshamn 01 30579
|
||||
0305 Håbo 03 22973
|
||||
0319 Älvkarleby 03 9552
|
||||
0330 Knivsta 03 21193
|
||||
0331 Heby 03 14345
|
||||
0360 Tierp 03 21104
|
||||
0380 Uppsala 03 248016
|
||||
0381 Enköping 03 48591
|
||||
0382 Östhammar 03 22138
|
||||
0428 Vingåker 04 8750
|
||||
0461 Gnesta 04 11458
|
||||
0480 Nyköping 04 58344
|
||||
0481 Oxelösund 04 12031
|
||||
0482 Flen 04 15362
|
||||
0483 Katrineholm 04 34154
|
||||
0484 Eskilstuna 04 107203
|
||||
0486 Strängnäs 04 39313
|
||||
0488 Trosa 04 14927
|
||||
0509 Ödeshög 05 5237
|
||||
0512 Ydre 05 3626
|
||||
0513 Kinda 05 9957
|
||||
0560 Boxholm 05 5517
|
||||
0561 Åtvidaberg 05 11467
|
||||
0562 Finspång 05 21623
|
||||
0563 Valdemarsvik 05 7525
|
||||
0580 Linköping 05 168035
|
||||
0581 Norrköping 05 144980
|
||||
0582 Söderköping 05 14789
|
||||
0583 Motala 05 43505
|
||||
0584 Vadstena 05 7490
|
||||
0586 Mjölby 05 28695
|
||||
0604 Aneby 06 6797
|
||||
0617 Gnosjö 06 9131
|
||||
0642 Mullsjö 06 7594
|
||||
0643 Habo 06 13456
|
||||
0662 Gislaved 06 28936
|
||||
0665 Vaggeryd 06 14825
|
||||
0680 Jönköping 06 147654
|
||||
0682 Nässjö 06 31587
|
||||
0683 Värnamo 06 34542
|
||||
0684 Sävsjö 06 11563
|
||||
0685 Vetlanda 06 27528
|
||||
0686 Eksjö 06 17792
|
||||
0687 Tranås 06 18604
|
||||
0760 Uppvidinge 07 9061
|
||||
0761 Lessebo 07 8289
|
||||
0763 Tingsryd 07 11966
|
||||
0764 Alvesta 07 19830
|
||||
0765 Älmhult 07 17653
|
||||
0767 Markaryd 07 9938
|
||||
0780 Växjö 07 98334
|
||||
0781 Ljungby 07 28280
|
||||
0821 Högsby 08 5321
|
||||
0834 Torsås 08 6984
|
||||
0840 Mörbylånga 08 16224
|
||||
0860 Hultsfred 08 13673
|
||||
0861 Mönsterås 08 13069
|
||||
0862 Emmaboda 08 9006
|
||||
0880 Kalmar 08 72704
|
||||
0881 Nybro 08 19951
|
||||
0882 Oskarshamn 08 26923
|
||||
0883 Västervik 08 36447
|
||||
0884 Vimmerby 08 15384
|
||||
0885 Borgholm 08 10666
|
||||
0980 Gotland 09 60971
|
||||
1060 Olofström 10 13000
|
||||
1080 Karlskrona 10 66301
|
||||
1081 Ronneby 10 28741
|
||||
1082 Karlshamn 10 31751
|
||||
1083 Sölvesborg 10 17430
|
||||
1214 Svalöv 12 14543
|
||||
1230 Staffanstorp 12 27303
|
||||
1231 Burlöv 12 20101
|
||||
1233 Vellinge 12 37816
|
||||
1256 Östra Göinge 12 13978
|
||||
1257 Örkelljunga 12 10277
|
||||
1260 Bjuv 12 15985
|
||||
1261 Kävlinge 12 32477
|
||||
1262 Lomma 12 24715
|
||||
1263 Svedala 12 23581
|
||||
1264 Skurup 12 17099
|
||||
1265 Sjöbo 12 19337
|
||||
1266 Hörby 12 15562
|
||||
1267 Höör 12 17518
|
||||
1270 Tomelilla 12 13639
|
||||
1272 Bromölla 12 12470
|
||||
1273 Osby 12 12947
|
||||
1275 Perstorp 12 7235
|
||||
1276 Klippan 12 17714
|
||||
1277 Åstorp 12 16449
|
||||
1278 Båstad 12 16026
|
||||
1280 Malmö 12 365644
|
||||
1281 Lund 12 131590
|
||||
1282 Landskrona 12 47309
|
||||
1283 Helsingborg 12 152091
|
||||
1284 Höganäs 12 28430
|
||||
1285 Eslöv 12 34922
|
||||
1286 Ystad 12 32106
|
||||
1287 Trelleborg 12 47269
|
||||
1290 Kristianstad 12 86379
|
||||
1291 Simrishamn 12 18890
|
||||
1292 Ängelholm 12 45110
|
||||
1293 Hässleholm 12 52114
|
||||
1315 Hylte 13 10196
|
||||
1380 Halmstad 13 106084
|
||||
1381 Laholm 13 26595
|
||||
1382 Falkenberg 13 47337
|
||||
1383 Varberg 13 69070
|
||||
1384 Kungsbacka 13 85792
|
||||
1401 Härryda 14 40003
|
||||
1402 Partille 14 41060
|
||||
1407 Öckerö 14 12771
|
||||
1415 Stenungsund 14 27851
|
||||
1419 Tjörn 14 16092
|
||||
1421 Orust 14 15352
|
||||
1427 Sotenäs 14 9104
|
||||
1430 Munkedal 14 10354
|
||||
1435 Tanum 14 12773
|
||||
1438 Dals-Ed 14 4606
|
||||
1439 Färgelanda 14 6376
|
||||
1440 Ale 14 32576
|
||||
1441 Lerum 14 43570
|
||||
1442 Vårgårda 14 12474
|
||||
1443 Bollebygd 14 9802
|
||||
1444 Grästorp 14 5555
|
||||
1445 Essunga 14 5560
|
||||
1446 Karlsborg 14 7023
|
||||
1447 Gullspång 14 5031
|
||||
1452 Tranemo 14 11839
|
||||
1460 Bengtsfors 14 9076
|
||||
1461 Mellerud 14 9052
|
||||
1462 Lilla Edet 14 14442
|
||||
1463 Mark 14 35155
|
||||
1465 Svenljunga 14 10747
|
||||
1466 Herrljunga 14 9497
|
||||
1470 Vara 14 16088
|
||||
1471 Götene 14 13286
|
||||
1472 Tibro 14 11338
|
||||
1473 Töreboda 14 9043
|
||||
1480 Göteborg 14 608993
|
||||
1481 Mölndal 14 71420
|
||||
1482 Kungälv 14 50313
|
||||
1484 Lysekil 14 13907
|
||||
1485 Uddevalla 14 57010
|
||||
1486 Strömstad 14 13482
|
||||
1487 Vänersborg 14 40041
|
||||
1488 Trollhättan 14 59003
|
||||
1489 Alingsås 14 42722
|
||||
1490 Borås 14 114872
|
||||
1491 Ulricehamn 14 24985
|
||||
1492 Åmål 14 11906
|
||||
1493 Mariestad 14 24583
|
||||
1494 Lidköping 14 40425
|
||||
1495 Skara 14 18707
|
||||
1496 Skövde 14 57995
|
||||
1497 Hjo 14 9350
|
||||
1498 Tidaholm 14 12805
|
||||
1499 Falköping 14 32806
|
||||
1715 Kil 17 12061
|
||||
1730 Eda 17 8412
|
||||
1737 Torsby 17 11321
|
||||
1760 Storfors 17 3788
|
||||
1761 Hammarö 17 16992
|
||||
1762 Munkfors 17 3625
|
||||
1763 Forshaga 17 11520
|
||||
1764 Grums 17 9004
|
||||
1765 Årjäng 17 9825
|
||||
1766 Sunne 17 13356
|
||||
1780 Karlstad 17 98084
|
||||
1781 Kristinehamn 17 23756
|
||||
1782 Filipstad 17 9776
|
||||
1783 Hagfors 17 11418
|
||||
1784 Arvika 17 25547
|
||||
1785 Säffle 17 14899
|
||||
1814 Lekeberg 18 8606
|
||||
1860 Laxå 18 5423
|
||||
1861 Hallsberg 18 16120
|
||||
1862 Degerfors 18 9278
|
||||
1863 Hällefors 18 6321
|
||||
1864 Ljusnarsberg 18 4369
|
||||
1880 Örebro 18 160140
|
||||
1881 Kumla 18 22681
|
||||
1882 Askersund 18 11477
|
||||
1883 Karlskoga 18 30180
|
||||
1884 Nora 18 10639
|
||||
1885 Lindesberg 18 23141
|
||||
1904 Skinnskatteberg 19 4256
|
||||
1907 Surahammar 19 9845
|
||||
1960 Kungsör 19 8694
|
||||
1961 Hallstahammar 19 16653
|
||||
1962 Norberg 19 5452
|
||||
1980 Västerås 19 160634
|
||||
1981 Sala 19 22843
|
||||
1982 Fagersta 19 13072
|
||||
1983 Köping 19 25729
|
||||
1984 Arboga 19 13980
|
||||
2021 Vansbro 20 6752
|
||||
2023 Malung-Sälen 20 10254
|
||||
2026 Gagnef 20 10384
|
||||
2029 Leksand 20 16137
|
||||
2031 Rättvik 20 10998
|
||||
2034 Orsa 20 6851
|
||||
2039 Älvdalen 20 6882
|
||||
2061 Smedjebacken 20 10823
|
||||
2062 Mora 20 20540
|
||||
2080 Falun 20 59945
|
||||
2081 Borlänge 20 51425
|
||||
2082 Säter 20 11223
|
||||
2083 Hedemora 20 15281
|
||||
2084 Avesta 20 22417
|
||||
2085 Ludvika 20 26634
|
||||
2101 Ockelbo 21 5715
|
||||
2104 Hofors 21 9281
|
||||
2121 Ovanåker 21 11341
|
||||
2132 Nordanstig 21 9262
|
||||
2161 Ljusdal 21 18445
|
||||
2180 Gävle 21 103838
|
||||
2181 Sandviken 21 38360
|
||||
2182 Söderhamn 21 24545
|
||||
2183 Bollnäs 21 26243
|
||||
2184 Hudiksvall 21 37528
|
||||
2260 Ånge 22 9044
|
||||
2262 Timrå 22 17521
|
||||
2280 Härnösand 22 24515
|
||||
2281 Sundsvall 22 99048
|
||||
2282 Kramfors 22 17491
|
||||
2283 Sollefteå 22 18396
|
||||
2284 Örnsköldsvik 22 55443
|
||||
2303 Ragunda 23 5146
|
||||
2305 Bräcke 23 6035
|
||||
2309 Krokom 23 15680
|
||||
2313 Strömsund 23 11023
|
||||
2321 Åre 23 12693
|
||||
2326 Berg 23 7108
|
||||
2361 Härjedalen 23 10175
|
||||
2380 Östersund 23 64979
|
||||
2401 Nordmaling 24 6942
|
||||
2403 Bjurholm 24 2359
|
||||
2404 Vindeln 24 5417
|
||||
2409 Robertsfors 24 6690
|
||||
2417 Norsjö 24 3968
|
||||
2418 Malå 24 2962
|
||||
2421 Storuman 24 5577
|
||||
2422 Sorsele 24 2357
|
||||
2425 Dorotea 24 2294
|
||||
2460 Vännäs 24 9132
|
||||
2462 Vilhelmina 24 6229
|
||||
2463 Åsele 24 2694
|
||||
2480 Umeå 24 134249
|
||||
2481 Lycksele 24 12118
|
||||
2482 Skellefteå 24 78150
|
||||
2505 Arvidsjaur 25 6089
|
||||
2506 Arjeplog 25 2599
|
||||
2510 Jokkmokk 25 4701
|
||||
2513 Överkalix 25 3201
|
||||
2514 Kalix 25 15391
|
||||
2518 Övertorneå 25 4057
|
||||
2521 Pajala 25 5857
|
||||
2523 Gällivare 25 17233
|
||||
2560 Älvsbyn 25 7774
|
||||
2580 Luleå 25 79645
|
||||
2581 Piteå 25 42447
|
||||
2582 Boden 25 28049
|
||||
2583 Haparanda 25 9151
|
||||
2584 Kiruna 25 22426
|
||||
|
@@ -0,0 +1 @@
|
||||
{ "format": "{code}", "rows": "postal-code.tsv", "key": "code", "parent": "locality" }
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1 @@
|
||||
{ "format": "{name}", "rows": "region.tsv", "key": "code", "name": "name", "weight": "population" }
|
||||
@@ -0,0 +1,22 @@
|
||||
code name population timezone
|
||||
01 Stockholms län 2473307 Europe/Stockholm
|
||||
03 Uppsala län 407912 Europe/Stockholm
|
||||
04 Södermanlands län 301542 Europe/Stockholm
|
||||
05 Östergötlands län 472446 Europe/Stockholm
|
||||
06 Jönköpings län 370009 Europe/Stockholm
|
||||
07 Kronobergs län 203351 Europe/Stockholm
|
||||
08 Kalmar län 246352 Europe/Stockholm
|
||||
09 Gotlands län 60971 Europe/Stockholm
|
||||
10 Blekinge län 157223 Europe/Stockholm
|
||||
12 Skåne län 1428626 Europe/Stockholm
|
||||
13 Hallands län 345074 Europe/Stockholm
|
||||
14 Västra Götalands län 1772821 Europe/Stockholm
|
||||
17 Värmlands län 283384 Europe/Stockholm
|
||||
18 Örebro län 308375 Europe/Stockholm
|
||||
19 Västmanlands län 281158 Europe/Stockholm
|
||||
20 Dalarnas län 286546 Europe/Stockholm
|
||||
21 Gävleborgs län 284558 Europe/Stockholm
|
||||
22 Västernorrlands län 241458 Europe/Stockholm
|
||||
23 Jämtlands län 132839 Europe/Stockholm
|
||||
24 Västerbottens län 281138 Europe/Stockholm
|
||||
25 Norrbottens län 248620 Europe/Stockholm
|
||||
|
@@ -0,0 +1 @@
|
||||
{ "format": "{name}", "rows": "street.tsv", "parent": "locality", "weight": "segments" }
|
||||
+14765
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,13 @@
|
||||
{
|
||||
"format": "{street-number} {street}\n{locality}, {region} {postal-code}",
|
||||
"street": "{.street.name}",
|
||||
"street-number": [
|
||||
"{int(10,99)}",
|
||||
{ "format": "{int(100,999)}", "weight": 2 },
|
||||
{ "format": "{int(1000,9999)}", "weight": 0.5 },
|
||||
{ "format": "{int(10,99)}{int(100,999)}", "weight": 0.3 }
|
||||
],
|
||||
"locality": "{.locality.name}",
|
||||
"region": "{.region.abbr}",
|
||||
"postal-code": "{.postal-code.code}"
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
{ "format": "{name}", "rows": "locality.tsv", "key": "code", "name": "name", "parent": "municipality", "weight": "population" }
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1 @@
|
||||
{ "format": "{name}", "rows": "municipality.tsv", "key": "code", "name": "name", "parent": "region", "weight": "population" }
|
||||
@@ -0,0 +1,666 @@
|
||||
code name region population
|
||||
01001 Autauga County AL 61920
|
||||
01003 Baldwin County AL 267761
|
||||
01031 Coffee County AL 56953
|
||||
01055 Etowah County AL 103886
|
||||
01069 Houston County AL 110318
|
||||
01073 Jefferson County AL 665742
|
||||
01077 Lauderdale County AL 97135
|
||||
01081 Lee County AL 189881
|
||||
01083 Limestone County AL 122928
|
||||
01089 Madison County AL 433516
|
||||
01097 Mobile County AL 411658
|
||||
01101 Montgomery County AL 225891
|
||||
01103 Morgan County AL 126483
|
||||
01113 Russell County AL 58898
|
||||
01117 Shelby County AL 238552
|
||||
01125 Tuscaloosa County AL 241368
|
||||
02020 Anchorage Municipality AK 287155
|
||||
02090 Fairbanks North Star Borough AK 93972
|
||||
02110 Juneau City and Borough AK 31609
|
||||
04003 Cochise County AZ 126332
|
||||
04005 Coconino County AZ 144368
|
||||
04013 Maricopa County AZ 4689558
|
||||
04015 Mohave County AZ 228102
|
||||
04019 Pima County AZ 1074685
|
||||
04021 Pinal County AZ 539380
|
||||
04025 Yavapai County AZ 252552
|
||||
04027 Yuma County AZ 224449
|
||||
05007 Benton County AR 332554
|
||||
05031 Craighead County AR 116957
|
||||
05045 Faulkner County AR 133979
|
||||
05051 Garland County AR 99695
|
||||
05055 Greene County AR 47411
|
||||
05069 Jefferson County AR 62987
|
||||
05085 Lonoke County AR 76664
|
||||
05091 Miller County AR 42357
|
||||
05115 Pope County AR 64976
|
||||
05119 Pulaski County AR 404611
|
||||
05125 Saline County AR 133288
|
||||
05131 Sebastian County AR 130641
|
||||
05143 Washington County AR 271213
|
||||
06001 Alameda County CA 1636630
|
||||
06007 Butte County CA 209211
|
||||
06013 Contra Costa County CA 1170070
|
||||
06019 Fresno County CA 1035456
|
||||
06023 Humboldt County CA 131647
|
||||
06025 Imperial County CA 181411
|
||||
06029 Kern County CA 927068
|
||||
06031 Kings County CA 154327
|
||||
06037 Los Angeles County CA 9694934
|
||||
06039 Madera County CA 167927
|
||||
06041 Marin County CA 253694
|
||||
06047 Merced County CA 297260
|
||||
06053 Monterey County CA 433729
|
||||
06055 Napa County CA 132949
|
||||
06059 Orange County CA 3149507
|
||||
06061 Placer County CA 442081
|
||||
06065 Riverside County CA 2544916
|
||||
06067 Sacramento County CA 1618460
|
||||
06069 San Benito County CA 70082
|
||||
06071 San Bernardino County CA 2224091
|
||||
06073 San Diego County CA 3282248
|
||||
06075 San Francisco County CA 826079
|
||||
06077 San Joaquin County CA 823815
|
||||
06079 San Luis Obispo County CA 282367
|
||||
06081 San Mateo County CA 743568
|
||||
06083 Santa Barbara County CA 442065
|
||||
06085 Santa Clara County CA 1914391
|
||||
06087 Santa Cruz County CA 258852
|
||||
06089 Shasta County CA 181648
|
||||
06095 Solano County CA 455376
|
||||
06097 Sonoma County CA 486444
|
||||
06099 Stanislaus County CA 557719
|
||||
06101 Sutter County CA 98787
|
||||
06107 Tulare County CA 485146
|
||||
06111 Ventura County CA 830851
|
||||
06113 Yolo County CA 224410
|
||||
08001 Adams County CO 554668
|
||||
08005 Arapahoe County CO 673820
|
||||
08013 Boulder County CO 328560
|
||||
08014 Broomfield County CO 79174
|
||||
08031 Denver County CO 740613
|
||||
08035 Douglas County CO 399396
|
||||
08041 El Paso County CO 757040
|
||||
08059 Jefferson County CO 580451
|
||||
08069 Larimer County CO 377292
|
||||
08077 Mesa County CO 162845
|
||||
08101 Pueblo County CO 169277
|
||||
08123 Weld County CO 378426
|
||||
09110 Capitol Planning Region CT 994115
|
||||
09120 Greater Bridgeport Planning Region CT 337697
|
||||
09130 Lower Connecticut River Valley Planning Region CT 177311
|
||||
09140 Naugatuck Valley Planning Region CT 463349
|
||||
09160 Northwest Hills Planning Region CT 114690
|
||||
09170 South Central Connecticut Planning Region CT 578741
|
||||
09180 Southeastern Connecticut Planning Region CT 284015
|
||||
09190 Western Connecticut Planning Region CT 640482
|
||||
10001 Kent County DE 194786
|
||||
10003 New Castle County DE 588026
|
||||
11001 District of Columbia DC 693645
|
||||
12001 Alachua County FL 290028
|
||||
12005 Bay County FL 204479
|
||||
12009 Brevard County FL 663982
|
||||
12011 Broward County FL 2013317
|
||||
12031 Duval County FL 1062963
|
||||
12033 Escambia County FL 333834
|
||||
12035 Flagler County FL 140360
|
||||
12057 Hillsborough County FL 1574115
|
||||
12061 Indian River County FL 172799
|
||||
12069 Lake County FL 456068
|
||||
12071 Lee County FL 875607
|
||||
12073 Leon County FL 299048
|
||||
12081 Manatee County FL 468200
|
||||
12083 Marion County FL 442660
|
||||
12086 Miami-Dade County FL 2802029
|
||||
12091 Okaloosa County FL 221810
|
||||
12095 Orange County FL 1528002
|
||||
12097 Osceola County FL 481718
|
||||
12099 Palm Beach County FL 1575726
|
||||
12103 Pinellas County FL 948563
|
||||
12105 Polk County FL 874790
|
||||
12111 St. Lucie County FL 402449
|
||||
12115 Sarasota County FL 479958
|
||||
12117 Seminole County FL 491884
|
||||
12127 Volusia County FL 606573
|
||||
13015 Bartow County GA 120800
|
||||
13021 Bibb County GA 157556
|
||||
13031 Bulloch County GA 86949
|
||||
13045 Carroll County GA 131036
|
||||
13051 Chatham County GA 311855
|
||||
13057 Cherokee County GA 299273
|
||||
13059 Clarke County GA 129921
|
||||
13067 Cobb County GA 793345
|
||||
13077 Coweta County GA 160240
|
||||
13089 DeKalb County GA 774394
|
||||
13095 Dougherty County GA 82616
|
||||
13097 Douglas County GA 154293
|
||||
13113 Fayette County GA 125156
|
||||
13115 Floyd County GA 101378
|
||||
13121 Fulton County GA 1098791
|
||||
13135 Gwinnett County GA 1018099
|
||||
13139 Hall County GA 226568
|
||||
13151 Henry County GA 264922
|
||||
13153 Houston County GA 178214
|
||||
13179 Liberty County GA 70313
|
||||
13185 Lowndes County GA 122867
|
||||
13215 Muscogee County GA 202171
|
||||
13245 Richmond County GA 206559
|
||||
13285 Troup County GA 72844
|
||||
13313 Whitfield County GA 106212
|
||||
16001 Ada County ID 546141
|
||||
16005 Bannock County ID 91591
|
||||
16019 Bonneville County ID 135771
|
||||
16027 Canyon County ID 275123
|
||||
16055 Kootenai County ID 191864
|
||||
16057 Latah County ID 41842
|
||||
16065 Madison County ID 55172
|
||||
16069 Nez Perce County ID 42905
|
||||
16083 Twin Falls County ID 97539
|
||||
17001 Adams County IL 64267
|
||||
17007 Boone County IL 53568
|
||||
17019 Champaign County IL 209972
|
||||
17031 Cook County IL 5194625
|
||||
17037 DeKalb County IL 101835
|
||||
17043 DuPage County IL 934298
|
||||
17089 Kane County IL 525757
|
||||
17093 Kendall County IL 145470
|
||||
17095 Knox County IL 47767
|
||||
17097 Lake County IL 719339
|
||||
17111 McHenry County IL 317751
|
||||
17113 McLean County IL 171419
|
||||
17115 Macon County IL 99300
|
||||
17119 Madison County IL 263110
|
||||
17143 Peoria County IL 178553
|
||||
17161 Rock Island County IL 141869
|
||||
17163 St. Clair County IL 250708
|
||||
17167 Sangamon County IL 194170
|
||||
17179 Tazewell County IL 130049
|
||||
17183 Vermilion County IL 71259
|
||||
17197 Will County IL 712253
|
||||
17201 Winnebago County IL 283674
|
||||
18003 Allen County IN 402329
|
||||
18005 Bartholomew County IN 85729
|
||||
18011 Boone County IN 80689
|
||||
18019 Clark County IN 130451
|
||||
18035 Delaware County IN 113106
|
||||
18039 Elkhart County IN 208774
|
||||
18043 Floyd County IN 82153
|
||||
18053 Grant County IN 66524
|
||||
18057 Hamilton County IN 387036
|
||||
18059 Hancock County IN 90969
|
||||
18063 Hendricks County IN 193510
|
||||
18067 Howard County IN 83904
|
||||
18081 Johnson County IN 174262
|
||||
18089 Lake County IN 504612
|
||||
18091 LaPorte County IN 111294
|
||||
18095 Madison County IN 135088
|
||||
18097 Marion County IN 992196
|
||||
18105 Monroe County IN 143345
|
||||
18127 Porter County IN 176049
|
||||
18141 St. Joseph County IN 272861
|
||||
18157 Tippecanoe County IN 190456
|
||||
18163 Vanderburgh County IN 181995
|
||||
18167 Vigo County IN 106512
|
||||
18177 Wayne County IN 66169
|
||||
19013 Black Hawk County IA 131532
|
||||
19033 Cerro Gordo County IA 42372
|
||||
19049 Dallas County IA 118457
|
||||
19061 Dubuque County IA 99381
|
||||
19103 Johnson County IA 160044
|
||||
19113 Linn County IA 232028
|
||||
19127 Marshall County IA 39890
|
||||
19153 Polk County IA 516546
|
||||
19155 Pottawattamie County IA 92996
|
||||
19163 Scott County IA 175259
|
||||
19169 Story County IA 101291
|
||||
19179 Wapello County IA 35210
|
||||
19193 Woodbury County IA 106649
|
||||
20045 Douglas County KS 120920
|
||||
20055 Finney County KS 37505
|
||||
20057 Ford County KS 33993
|
||||
20091 Johnson County KS 636906
|
||||
20103 Leavenworth County KS 84590
|
||||
20155 Reno County KS 61539
|
||||
20161 Riley County KS 72598
|
||||
20169 Saline County KS 53377
|
||||
20173 Sedgwick County KS 538433
|
||||
20177 Shawnee County KS 178607
|
||||
20209 Wyandotte County KS 170597
|
||||
21015 Boone County KY 145316
|
||||
21047 Christian County KY 70115
|
||||
21059 Daviess County KY 104898
|
||||
21067 Fayette County KY 329751
|
||||
21073 Franklin County KY 52649
|
||||
21093 Hardin County KY 113482
|
||||
21101 Henderson County KY 44255
|
||||
21111 Jefferson County KY 795222
|
||||
21113 Jessamine County KY 57147
|
||||
21117 Kenton County KY 175779
|
||||
21145 McCracken County KY 67553
|
||||
21151 Madison County KY 101696
|
||||
21209 Scott County KY 62262
|
||||
21227 Warren County KY 149375
|
||||
22015 Bossier Parish LA 131867
|
||||
22017 Caddo Parish LA 224226
|
||||
22019 Calcasieu Parish LA 208466
|
||||
22033 East Baton Rouge Parish LA 456180
|
||||
22045 Iberia Parish LA 66846
|
||||
22051 Jefferson Parish LA 431398
|
||||
22071 Orleans Parish LA 362154
|
||||
22073 Ouachita Parish LA 158542
|
||||
22079 Rapides Parish LA 125877
|
||||
22103 St. Tammany Parish LA 279108
|
||||
22109 Terrebonne Parish LA 104163
|
||||
23001 Androscoggin County ME 116487
|
||||
23005 Cumberland County ME 317222
|
||||
23019 Penobscot County ME 157967
|
||||
24003 Anne Arundel County MD 603380
|
||||
24021 Frederick County MD 302883
|
||||
24031 Montgomery County MD 1074582
|
||||
24033 Prince George's County MD 970374
|
||||
24043 Washington County MD 157731
|
||||
24045 Wicomico County MD 106899
|
||||
24510 Baltimore city MD 569997
|
||||
25001 Barnstable County MA 233539
|
||||
25003 Berkshire County MA 128224
|
||||
25005 Bristol County MA 593640
|
||||
25009 Essex County MA 826653
|
||||
25013 Hampden County MA 464338
|
||||
25015 Hampshire County MA 164065
|
||||
25017 Middlesex County MA 1669979
|
||||
25021 Norfolk County MA 739749
|
||||
25023 Plymouth County MA 546829
|
||||
25025 Suffolk County MA 791891
|
||||
25027 Worcester County MA 888502
|
||||
26017 Bay County MI 102123
|
||||
26025 Calhoun County MI 133408
|
||||
26049 Genesee County MI 401093
|
||||
26065 Ingham County MI 289709
|
||||
26075 Jackson County MI 159552
|
||||
26077 Kalamazoo County MI 263795
|
||||
26081 Kent County MI 675232
|
||||
26099 Macomb County MI 886221
|
||||
26111 Midland County MI 83754
|
||||
26121 Muskegon County MI 177901
|
||||
26125 Oakland County MI 1288337
|
||||
26139 Ottawa County MI 308459
|
||||
26145 Saginaw County MI 187688
|
||||
26147 St. Clair County MI 160486
|
||||
26161 Washtenaw County MI 370214
|
||||
26163 Wayne County MI 1769038
|
||||
27003 Anoka County MN 381605
|
||||
27013 Blue Earth County MN 70634
|
||||
27019 Carver County MN 114379
|
||||
27027 Clay County MN 67734
|
||||
27037 Dakota County MN 457710
|
||||
27053 Hennepin County MN 1284784
|
||||
27099 Mower County MN 40971
|
||||
27109 Olmsted County MN 166731
|
||||
27123 Ramsey County MN 541623
|
||||
27131 Rice County MN 69939
|
||||
27137 St. Louis County MN 200518
|
||||
27139 Scott County MN 159017
|
||||
27141 Sherburne County MN 104194
|
||||
27145 Stearns County MN 164110
|
||||
27147 Steele County MN 37464
|
||||
27163 Washington County MN 286895
|
||||
27169 Winona County MN 50523
|
||||
28033 DeSoto County MS 197918
|
||||
28035 Forrest County MS 79034
|
||||
28047 Harrison County MS 217136
|
||||
28049 Hinds County MS 211888
|
||||
28071 Lafayette County MS 59597
|
||||
28075 Lauderdale County MS 70317
|
||||
28081 Lee County MS 83731
|
||||
28089 Madison County MS 116298
|
||||
28105 Oktibbeha County MS 51896
|
||||
28121 Rankin County MS 162181
|
||||
28151 Washington County MS 40446
|
||||
29019 Boone County MO 191746
|
||||
29021 Buchanan County MO 83540
|
||||
29031 Cape Girardeau County MO 83999
|
||||
29037 Cass County MO 115859
|
||||
29043 Christian County MO 96725
|
||||
29047 Clay County MO 265032
|
||||
29051 Cole County MO 77908
|
||||
29077 Greene County MO 309286
|
||||
29095 Jackson County MO 732994
|
||||
29097 Jasper County MO 127428
|
||||
29183 St. Charles County MO 426499
|
||||
29189 St. Louis County MO 990911
|
||||
29510 St. Louis city MO 278144
|
||||
30013 Cascade County MT 85029
|
||||
30029 Flathead County MT 115429
|
||||
30031 Gallatin County MT 128740
|
||||
30049 Lewis and Clark County MT 75331
|
||||
30063 Missoula County MT 123513
|
||||
30093 Silver Bow County MT 36118
|
||||
30111 Yellowstone County MT 172692
|
||||
31001 Adams County NE 31071
|
||||
31019 Buffalo County NE 51172
|
||||
31053 Dodge County NE 38057
|
||||
31055 Douglas County NE 606460
|
||||
31079 Hall County NE 63633
|
||||
31109 Lancaster County NE 334049
|
||||
31119 Madison County NE 36106
|
||||
31141 Platte County NE 35649
|
||||
31153 Sarpy County NE 208303
|
||||
32003 Clark County NV 2407226
|
||||
32019 Lyon County NV 65088
|
||||
32031 Washoe County NV 509386
|
||||
32510 Carson City NV 58571
|
||||
33011 Hillsborough County NH 433415
|
||||
33013 Merrimack County NH 158078
|
||||
33017 Strafford County NH 135043
|
||||
34001 Atlantic County NJ 278657
|
||||
34003 Bergen County NJ 977026
|
||||
34007 Camden County NJ 535799
|
||||
34011 Cumberland County NJ 157148
|
||||
34013 Essex County NJ 896379
|
||||
34017 Hudson County NJ 735033
|
||||
34021 Mercer County NJ 399289
|
||||
34023 Middlesex County NJ 883335
|
||||
34025 Monmouth County NJ 651035
|
||||
34031 Passaic County NJ 531624
|
||||
34039 Union County NJ 601863
|
||||
35001 Bernalillo County NM 667601
|
||||
35005 Chaves County NM 63364
|
||||
35009 Curry County NM 46655
|
||||
35013 Doña Ana County NM 229091
|
||||
35015 Eddy County NM 62509
|
||||
35025 Lea County NM 74749
|
||||
35035 Otero County NM 70368
|
||||
35043 Sandoval County NM 159565
|
||||
35045 San Juan County NM 120340
|
||||
35049 Santa Fe County NM 156907
|
||||
36001 Albany County NY 321225
|
||||
36007 Broome County NY 195736
|
||||
36011 Cayuga County NY 74365
|
||||
36013 Chautauqua County NY 124126
|
||||
36015 Chemung County NY 80415
|
||||
36027 Dutchess County NY 300708
|
||||
36029 Erie County NY 946741
|
||||
36047 Kings County NY 2653963
|
||||
36055 Monroe County NY 750506
|
||||
36059 Nassau County NY 1398939
|
||||
36063 Niagara County NY 208912
|
||||
36065 Oneida County NY 226392
|
||||
36067 Onondaga County NY 466584
|
||||
36071 Orange County NY 417669
|
||||
36083 Rensselaer County NY 160510
|
||||
36091 Saratoga County NY 241343
|
||||
36093 Schenectady County NY 162581
|
||||
36103 Suffolk County NY 1546090
|
||||
36109 Tompkins County NY 104047
|
||||
36119 Westchester County NY 1015743
|
||||
37001 Alamance County NC 186177
|
||||
37019 Brunswick County NC 174702
|
||||
37021 Buncombe County NC 277417
|
||||
37025 Cabarrus County NC 249725
|
||||
37035 Catawba County NC 170172
|
||||
37049 Craven County NC 105025
|
||||
37051 Cumberland County NC 338473
|
||||
37057 Davidson County NC 180182
|
||||
37063 Durham County NC 347240
|
||||
37067 Forsyth County NC 401718
|
||||
37071 Gaston County NC 246558
|
||||
37081 Guilford County NC 562234
|
||||
37097 Iredell County NC 211798
|
||||
37101 Johnston County NC 256448
|
||||
37105 Lee County NC 70258
|
||||
37119 Mecklenburg County NC 1233383
|
||||
37127 Nash County NC 99365
|
||||
37129 New Hanover County NC 245959
|
||||
37133 Onslow County NC 217175
|
||||
37135 Orange County NC 152498
|
||||
37147 Pitt County NC 182936
|
||||
37151 Randolph County NC 149516
|
||||
37159 Rowan County NC 155096
|
||||
37179 Union County NC 267674
|
||||
37183 Wake County NC 1257235
|
||||
37191 Wayne County NC 122278
|
||||
37195 Wilson County NC 81150
|
||||
38015 Burleigh County ND 103251
|
||||
38017 Cass County ND 201794
|
||||
38035 Grand Forks County ND 74501
|
||||
38059 Morton County ND 34601
|
||||
38089 Stark County ND 34013
|
||||
38101 Ward County ND 68233
|
||||
38105 Williams County ND 41767
|
||||
39003 Allen County OH 100881
|
||||
39009 Athens County OH 63197
|
||||
39017 Butler County OH 400128
|
||||
39023 Clark County OH 135340
|
||||
39035 Cuyahoga County OH 1232925
|
||||
39041 Delaware County OH 242032
|
||||
39045 Fairfield County OH 169752
|
||||
39049 Franklin County OH 1361536
|
||||
39057 Greene County OH 174322
|
||||
39061 Hamilton County OH 838418
|
||||
39063 Hancock County OH 75034
|
||||
39085 Lake County OH 232217
|
||||
39089 Licking County OH 185564
|
||||
39093 Lorain County OH 323219
|
||||
39095 Lucas County OH 423347
|
||||
39099 Mahoning County OH 224706
|
||||
39101 Marion County OH 65115
|
||||
39103 Medina County OH 185025
|
||||
39109 Miami County OH 112634
|
||||
39113 Montgomery County OH 539598
|
||||
39119 Muskingum County OH 87014
|
||||
39133 Portage County OH 163404
|
||||
39139 Richland County OH 124893
|
||||
39151 Stark County OH 373771
|
||||
39153 Summit County OH 538376
|
||||
39155 Trumbull County OH 198972
|
||||
39159 Union County OH 73446
|
||||
39165 Warren County OH 257181
|
||||
39169 Wayne County OH 116758
|
||||
39173 Wood County OH 134176
|
||||
40027 Cleveland County OK 303973
|
||||
40031 Comanche County OK 122158
|
||||
40047 Garfield County OK 61779
|
||||
40101 Muskogee County OK 66708
|
||||
40109 Oklahoma County OK 822125
|
||||
40119 Payne County OK 83889
|
||||
40125 Pottawatomie County OK 75102
|
||||
40143 Tulsa County OK 698782
|
||||
40147 Washington County OK 54037
|
||||
41003 Benton County OR 97728
|
||||
41005 Clackamas County OR 426280
|
||||
41017 Deschutes County OR 213072
|
||||
41029 Jackson County OR 221795
|
||||
41033 Josephine County OR 87867
|
||||
41039 Lane County OR 381584
|
||||
41043 Linn County OR 132843
|
||||
41047 Marion County OR 355777
|
||||
41051 Multnomah County OR 795391
|
||||
41067 Washington County OR 611708
|
||||
41071 Yamhill County OR 110024
|
||||
42003 Allegheny County PA 1225035
|
||||
42011 Berks County PA 440072
|
||||
42013 Blair County PA 119541
|
||||
42027 Centre County PA 157393
|
||||
42043 Dauphin County PA 293351
|
||||
42045 Delaware County PA 580937
|
||||
42049 Erie County PA 265832
|
||||
42069 Lackawanna County PA 216502
|
||||
42071 Lancaster County PA 563159
|
||||
42075 Lebanon County PA 146380
|
||||
42077 Lehigh County PA 384383
|
||||
42079 Luzerne County PA 332126
|
||||
42081 Lycoming County PA 112587
|
||||
42091 Montgomery County PA 877643
|
||||
42095 Northampton County PA 324411
|
||||
42101 Philadelphia County PA 1574281
|
||||
42133 York County PA 473197
|
||||
44003 Kent County RI 173495
|
||||
44007 Providence County RI 678179
|
||||
45003 Aiken County SC 181515
|
||||
45007 Anderson County SC 219930
|
||||
45013 Beaufort County SC 204433
|
||||
45015 Berkeley County SC 274666
|
||||
45019 Charleston County SC 436200
|
||||
45035 Dorchester County SC 178397
|
||||
45041 Florence County SC 138504
|
||||
45045 Greenville County SC 583125
|
||||
45051 Horry County SC 427551
|
||||
45063 Lexington County SC 317588
|
||||
45077 Pickens County SC 139198
|
||||
45079 Richland County SC 434956
|
||||
45083 Spartanburg County SC 380857
|
||||
45085 Sumter County SC 105067
|
||||
45091 York County SC 306887
|
||||
46011 Brookings County SD 37635
|
||||
46013 Brown County SD 37561
|
||||
46099 Minnehaha County SD 212691
|
||||
46103 Pennington County SD 116792
|
||||
47001 Anderson County TN 82066
|
||||
47003 Bedford County TN 55273
|
||||
47009 Blount County TN 143820
|
||||
47011 Bradley County TN 115465
|
||||
47037 Davidson County TN 745904
|
||||
47063 Hamblen County TN 68843
|
||||
47065 Hamilton County TN 390833
|
||||
47093 Knox County TN 511453
|
||||
47113 Madison County TN 100790
|
||||
47119 Maury County TN 118131
|
||||
47125 Montgomery County TN 249935
|
||||
47141 Putnam County TN 86612
|
||||
47149 Rutherford County TN 386352
|
||||
47157 Shelby County TN 910226
|
||||
47163 Sullivan County TN 163759
|
||||
47165 Sumner County TN 215538
|
||||
47179 Washington County TN 141199
|
||||
47187 Williamson County TN 272061
|
||||
47189 Wilson County TN 175033
|
||||
48005 Angelina County TX 88154
|
||||
48027 Bell County TX 402248
|
||||
48029 Bexar County TX 2160088
|
||||
48037 Bowie County TX 92696
|
||||
48039 Brazoria County TX 419080
|
||||
48041 Brazos County TX 249088
|
||||
48061 Cameron County TX 433946
|
||||
48085 Collin County TX 1297179
|
||||
48091 Comal County TX 209166
|
||||
48099 Coryell County TX 85592
|
||||
48113 Dallas County TX 2661397
|
||||
48121 Denton County TX 1069346
|
||||
48135 Ector County TX 173801
|
||||
48139 Ellis County TX 240867
|
||||
48141 El Paso County TX 877858
|
||||
48157 Fort Bend County TX 975191
|
||||
48167 Galveston County TX 372207
|
||||
48181 Grayson County TX 153613
|
||||
48183 Gregg County TX 126095
|
||||
48187 Guadalupe County TX 201111
|
||||
48201 Harris County TX 5045026
|
||||
48209 Hays County TX 304390
|
||||
48215 Hidalgo County TX 921549
|
||||
48231 Hunt County TX 123336
|
||||
48245 Jefferson County TX 254321
|
||||
48251 Johnson County TX 218048
|
||||
48257 Kaufman County TX 209235
|
||||
48265 Kerr County TX 54037
|
||||
48277 Lamar County TX 51503
|
||||
48303 Lubbock County TX 328906
|
||||
48309 McLennan County TX 272020
|
||||
48323 Maverick County TX 58823
|
||||
48329 Midland County TX 187855
|
||||
48339 Montgomery County TX 781194
|
||||
48347 Nacogdoches County TX 66035
|
||||
48349 Navarro County TX 57181
|
||||
48355 Nueces County TX 352992
|
||||
48367 Parker County TX 184767
|
||||
48381 Randall County TX 152351
|
||||
48397 Rockwall County TX 140738
|
||||
48423 Smith County TX 252549
|
||||
48439 Tarrant County TX 2248466
|
||||
48441 Taylor County TX 150077
|
||||
48451 Tom Green County TX 120602
|
||||
48453 Travis County TX 1389670
|
||||
48465 Val Verde County TX 47835
|
||||
48469 Victoria County TX 92656
|
||||
48471 Walker County TX 83842
|
||||
48479 Webb County TX 281224
|
||||
48485 Wichita County TX 129555
|
||||
48491 Williamson County TX 752827
|
||||
49005 Cache County UT 145000
|
||||
49011 Davis County UT 381227
|
||||
49021 Iron County UT 67141
|
||||
49035 Salt Lake County UT 1220916
|
||||
49045 Tooele County UT 87461
|
||||
49049 Utah County UT 759859
|
||||
49053 Washington County UT 213670
|
||||
49057 Weber County UT 278174
|
||||
50007 Chittenden County VT 169115
|
||||
51059 Fairfax County VA 1167873
|
||||
51107 Loudoun County VA 449749
|
||||
51121 Montgomery County VA 98434
|
||||
51510 Alexandria city VA 160662
|
||||
51540 Charlottesville city VA 44388
|
||||
51550 Chesapeake city VA 255332
|
||||
51590 Danville city VA 41647
|
||||
51600 Fairfax city VA 26772
|
||||
51630 Fredericksburg city VA 30393
|
||||
51650 Hampton city VA 137315
|
||||
51660 Harrisonburg city VA 50839
|
||||
51680 Lynchburg city VA 81347
|
||||
51683 Manassas city VA 44332
|
||||
51700 Newport News city VA 183230
|
||||
51710 Norfolk city VA 231013
|
||||
51730 Petersburg city VA 33734
|
||||
51740 Portsmouth city VA 96777
|
||||
51760 Richmond city VA 237257
|
||||
51770 Roanoke city VA 99111
|
||||
51775 Salem city VA 25816
|
||||
51790 Staunton city VA 26801
|
||||
51800 Suffolk city VA 104699
|
||||
51810 Virginia Beach city VA 453737
|
||||
51840 Winchester city VA 28272
|
||||
53005 Benton County WA 221722
|
||||
53007 Chelan County WA 81941
|
||||
53011 Clark County WA 532119
|
||||
53015 Cowlitz County WA 114885
|
||||
53021 Franklin County WA 102612
|
||||
53025 Grant County WA 105727
|
||||
53033 King County WA 2344939
|
||||
53035 Kitsap County WA 283374
|
||||
53053 Pierce County WA 946288
|
||||
53057 Skagit County WA 132975
|
||||
53061 Snohomish County WA 870656
|
||||
53063 Spokane County WA 558344
|
||||
53067 Thurston County WA 304261
|
||||
53071 Walla Walla County WA 62361
|
||||
53073 Whatcom County WA 236392
|
||||
53075 Whitman County WA 48512
|
||||
53077 Yakima County WA 259185
|
||||
54011 Cabell County WV 91183
|
||||
54039 Kanawha County WV 172381
|
||||
54061 Monongalia County WV 107991
|
||||
54069 Ohio County WV 40496
|
||||
54107 Wood County WV 82385
|
||||
55009 Brown County WI 275803
|
||||
55025 Dane County WI 590375
|
||||
55031 Douglas County WI 43990
|
||||
55035 Eau Claire County WI 109033
|
||||
55039 Fond du Lac County WI 104669
|
||||
55059 Kenosha County WI 168448
|
||||
55063 La Crosse County WI 121339
|
||||
55071 Manitowoc County WI 81710
|
||||
55073 Marathon County WI 139432
|
||||
55079 Milwaukee County WI 924216
|
||||
55087 Outagamie County WI 195894
|
||||
55089 Ozaukee County WI 94346
|
||||
55097 Portage County WI 71943
|
||||
55101 Racine County WI 198919
|
||||
55105 Rock County WI 166472
|
||||
55117 Sheboygan County WI 118047
|
||||
55131 Washington County WI 139238
|
||||
55133 Waukesha County WI 417210
|
||||
55139 Winnebago County WI 174218
|
||||
56001 Albany County WY 38558
|
||||
56005 Campbell County WY 48145
|
||||
56021 Laramie County WY 102938
|
||||
56025 Natrona County WY 80526
|
||||
|
@@ -0,0 +1 @@
|
||||
{ "format": "{code}", "rows": "postal-code.tsv", "key": "code", "parent": "locality", "weight": "addresses" }
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1 @@
|
||||
{ "format": "{name}", "rows": "region.tsv", "key": "abbr", "name": "name", "weight": "population" }
|
||||
@@ -0,0 +1,51 @@
|
||||
abbr code name population timezone
|
||||
AK 02 Alaska 737270 America/Anchorage
|
||||
AL 01 Alabama 5193088 America/Chicago
|
||||
AR 05 Arkansas 3114791 America/Chicago
|
||||
AZ 04 Arizona 7623818 America/Phoenix
|
||||
CA 06 California 39355309 America/Los_Angeles
|
||||
CO 08 Colorado 6012561 America/Denver
|
||||
CT 09 Connecticut 3688496 America/New_York
|
||||
DC 11 District of Columbia 693645 America/New_York
|
||||
DE 10 Delaware 1059952 America/New_York
|
||||
FL 12 Florida 23462518 America/New_York
|
||||
GA 13 Georgia 11302748 America/New_York
|
||||
IA 19 Iowa 3238387 America/Chicago
|
||||
ID 16 Idaho 2029733 America/Boise
|
||||
IL 17 Illinois 12719141 America/Chicago
|
||||
IN 18 Indiana 6973333 America/Indiana/Indianapolis
|
||||
KS 20 Kansas 2977220 America/Chicago
|
||||
KY 21 Kentucky 4606864 America/New_York
|
||||
LA 22 Louisiana 4618189 America/Chicago
|
||||
MA 25 Massachusetts 7154084 America/New_York
|
||||
MD 24 Maryland 6265347 America/New_York
|
||||
ME 23 Maine 1414874 America/New_York
|
||||
MI 26 Michigan 10127884 America/Detroit
|
||||
MN 27 Minnesota 5830405 America/Chicago
|
||||
MO 29 Missouri 6270541 America/Chicago
|
||||
MS 28 Mississippi 2954160 America/Chicago
|
||||
MT 30 Montana 1144694 America/Denver
|
||||
NC 37 North Carolina 11197968 America/New_York
|
||||
ND 38 North Dakota 799358 America/Chicago
|
||||
NE 31 Nebraska 2018006 America/Chicago
|
||||
NH 33 New Hampshire 1415342 America/New_York
|
||||
NJ 34 New Jersey 9548215 America/New_York
|
||||
NM 35 New Mexico 2125498 America/Denver
|
||||
NV 32 Nevada 3282188 America/Los_Angeles
|
||||
NY 36 New York 20002427 America/New_York
|
||||
OH 39 Ohio 11900510 America/New_York
|
||||
OK 40 Oklahoma 4123288 America/Chicago
|
||||
OR 41 Oregon 4273586 America/Los_Angeles
|
||||
PA 42 Pennsylvania 13059432 America/New_York
|
||||
RI 44 Rhode Island 1114521 America/New_York
|
||||
SC 45 South Carolina 5570274 America/New_York
|
||||
SD 46 South Dakota 935094 America/Chicago
|
||||
TN 47 Tennessee 7315076 America/Chicago
|
||||
TX 48 Texas 31709821 America/Chicago
|
||||
UT 49 Utah 3538904 America/Denver
|
||||
VA 51 Virginia 8880107 America/New_York
|
||||
VT 50 Vermont 644663 America/New_York
|
||||
WA 53 Washington 8001020 America/Los_Angeles
|
||||
WI 55 Wisconsin 5972787 America/Chicago
|
||||
WV 54 West Virginia 1766147 America/New_York
|
||||
WY 56 Wyoming 588753 America/Denver
|
||||
|
@@ -0,0 +1 @@
|
||||
{ "format": "{name}", "rows": "street.tsv", "parent": "locality", "weight": "addresses" }
|
||||
+15841
File diff suppressed because it is too large
Load Diff
+4
-18
@@ -1,21 +1,7 @@
|
||||
{
|
||||
"format": "{street} {street-number}\n{postal-code} {locality}",
|
||||
"street": [
|
||||
{
|
||||
"format": "{first}{last}",
|
||||
"first": ["Ängs", "Bergs", "Björk", "Drottning", "Eke", "Furu", "Hamn", "Köpmans", "Kungs", "Linné", "Norra", "Nybro", "Oden", "Öster", "Park", "Skogs", "Skol", "Slotts", "Söder", "Stations", "Stor", "Strand", "Trädgårds", "Vasa", "Väster"],
|
||||
"last": ["gatan", "vägen", "stigen", "gränd", "backen", "torget", "allén"]
|
||||
},
|
||||
["Avenyn", "Birger Jarlsgatan", "Promenaden", "Staby", "Sveavägen", "Vintjärn"]
|
||||
],
|
||||
"street-number": [
|
||||
"{int(1,9)}",
|
||||
"{int(10,99)}",
|
||||
{ "format": "{int(100,999)}", "weight": 0.2 },
|
||||
{ "format": "{int(1,9)}{upper(1)}", "weight": 0.2 },
|
||||
{ "format": "{int(10,99)}{upper(1)}", "weight": 0.1 },
|
||||
{ "format": "{int(100,999)}{upper(1)}", "weight": 0.05 }
|
||||
],
|
||||
"postal-code": "{int(100,999)} {int(10,99)}",
|
||||
"locality": ["Alingsås", "Alvesta", "Ängelholm", "Arboga", "Arvika", "Avesta", "Boden", "Bollnäs", "Borås", "Borlänge", "Enköping", "Eskilstuna", "Eslöv", "Fagersta", "Falkenberg", "Falköping", "Falun", "Finspång", "Gällivare", "Gävle", "Göteborg", "Halmstad", "Haparanda", "Härnösand", "Hässleholm", "Helsingborg", "Huddinge", "Hudiksvall", "Jönköping", "Kalmar", "Karlshamn", "Karlskoga", "Karlskrona", "Karlstad", "Katrineholm", "Kiruna", "Köping", "Kramfors", "Kristianstad", "Kristinehamn", "Landskrona", "Lidingö", "Lidköping", "Lindesberg", "Linköping", "Ljungby", "Ludvika", "Luleå", "Lund", "Lycksele", "Malmö", "Mariestad", "Mjölby", "Mölndal", "Mora", "Motala", "Nacka", "Nässjö", "Norrköping", "Norrtälje", "Nyköping", "Nynäshamn", "Örebro", "Örnsköldsvik", "Oskarshamn", "Östersund", "Piteå", "Rabbalshede", "Ronneby", "Säffle", "Sandviken", "Sävsjö", "Sigtuna", "Skara", "Skellefteå", "Skövde", "Söderhamn", "Södertälje", "Sollentuna", "Solna", "Sölvesborg", "Stockholm", "Strängnäs", "Sundbyberg", "Sundsvall", "Täby", "Tierp", "Tranås", "Trelleborg", "Trollhättan", "Uddevalla", "Ulricehamn", "Umeå", "Upplands Väsby", "Uppsala", "Vänersborg", "Varberg", "Värnamo", "Västerås", "Västervik", "Växjö", "Vetlanda", "Vimmerby", "Visby", "Ystad"]
|
||||
"street": "{/geo.SE.address.street}",
|
||||
"street-number": "{/geo.SE.address.street-number}",
|
||||
"postal-code": "{/geo.SE.address.postal-code}",
|
||||
"locality": "{/geo.SE.address.locality}"
|
||||
}
|
||||
|
||||
+27
-28
@@ -51,14 +51,13 @@ func TestShippedDataCategories(t *testing.T) {
|
||||
regexp.MustCompile(`^\d{1,3}( \d{3})?(,\d{2})? kr$`)},
|
||||
}
|
||||
|
||||
en := newGenerator(t, "data/en_US", WithSeed(1))
|
||||
sv := newGenerator(t, "data/sv_SE", WithSeed(1))
|
||||
f := newGenerator(t, "data", WithSeed(1))
|
||||
for _, c := range cases {
|
||||
for i := 0; i < 200; i++ {
|
||||
if v := fake(t, en, c.path); !c.en.MatchString(v) {
|
||||
if v := fake(t, f, "en_US."+c.path); !c.en.MatchString(v) {
|
||||
t.Fatalf("en_US %s = %q, want %s", c.path, v, c.en)
|
||||
}
|
||||
if v := fake(t, sv, c.path); !c.sv.MatchString(v) {
|
||||
if v := fake(t, f, "sv_SE."+c.path); !c.sv.MatchString(v) {
|
||||
t.Fatalf("sv_SE %s = %q, want %s", c.path, v, c.sv)
|
||||
}
|
||||
}
|
||||
@@ -134,8 +133,8 @@ func TestShippedMiscReferenceData(t *testing.T) {
|
||||
// TestSwedishPersonNamesHaveNoTripleLetter pins an orthographic rule the shape
|
||||
// regexes miss: no generated name repeats a character three times over.
|
||||
func TestSwedishPersonNamesHaveNoTripleLetter(t *testing.T) {
|
||||
f := newGenerator(t, "data/sv_SE", WithSeed(11))
|
||||
for _, path := range []string{"person", "person.last"} {
|
||||
f := newGenerator(t, "data", WithSeed(11))
|
||||
for _, path := range []string{"sv_SE.person", "sv_SE.person.last"} {
|
||||
for i := 0; i < 20000; i++ {
|
||||
name := fake(t, f, path)
|
||||
r := []rune(name)
|
||||
@@ -152,10 +151,10 @@ func TestSwedishPersonNamesHaveNoTripleLetter(t *testing.T) {
|
||||
// is a real calendar date (so month-length variants never emit e.g. Apr 31 or
|
||||
// Feb 30) and the trailing digit is a valid Luhn checksum over the other nine.
|
||||
func TestSwedishPersonnummer(t *testing.T) {
|
||||
sv := newGenerator(t, "data/sv_SE", WithSeed(1))
|
||||
sv := newGenerator(t, "data", WithSeed(1))
|
||||
sawLongMonthEnd := false
|
||||
for i := 0; i < 2000; i++ {
|
||||
v := fake(t, sv, "ssn")
|
||||
v := fake(t, sv, "sv_SE.ssn")
|
||||
d := digitsOnly(v)
|
||||
if len(d) != 10 {
|
||||
t.Fatalf("ssn %q has %d digits, want 10", v, len(d))
|
||||
@@ -176,49 +175,49 @@ func TestSwedishPersonnummer(t *testing.T) {
|
||||
}
|
||||
|
||||
func TestShippedSwedishPhone(t *testing.T) {
|
||||
f := newGenerator(t, "data/sv_SE", WithSeed(11))
|
||||
f := newGenerator(t, "data", WithSeed(11))
|
||||
re := regexp.MustCompile(`^0\d{1,2}-\d{3} \d{2} \d{2}$`)
|
||||
for i := 0; i < 50; i++ {
|
||||
if n := fake(t, f, "phone"); !re.MatchString(n) {
|
||||
if n := fake(t, f, "sv_SE.phone"); !re.MatchString(n) {
|
||||
t.Fatalf("phone %q does not match %s", n, re)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestShippedSwedishAddress(t *testing.T) {
|
||||
f := newGenerator(t, "data/sv_SE", WithSeed(3))
|
||||
f := newGenerator(t, "data", WithSeed(3))
|
||||
digit := regexp.MustCompile(`\d`)
|
||||
for i := 0; i < 30; i++ {
|
||||
a := fake(t, f, "address")
|
||||
a := fake(t, f, "sv_SE.address")
|
||||
if !regexp.MustCompile(`\n`).MatchString(a) || !digit.MatchString(a) {
|
||||
t.Fatalf("address %q is not a multi-line address with a number", a)
|
||||
}
|
||||
}
|
||||
|
||||
locality := regexp.MustCompile(`^\p{L}+( \p{L}+)*$`)
|
||||
locality := regexp.MustCompile(`^\p{L}+([ -]\p{L}+)*$`)
|
||||
for i := 0; i < 30; i++ {
|
||||
if c := fake(t, f, "address.locality"); !locality.MatchString(c) {
|
||||
if c := fake(t, f, "sv_SE.address.locality"); !locality.MatchString(c) {
|
||||
t.Fatalf("locality %q is not a Swedish place name", c)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestShippedPersonHasParts(t *testing.T) {
|
||||
for _, dir := range []string{"data/sv_SE", "data/en_US"} {
|
||||
f := newGenerator(t, dir, WithSeed(7))
|
||||
f := newGenerator(t, "data", WithSeed(7))
|
||||
for _, path := range []string{"sv_SE.person", "en_US.person"} {
|
||||
for i := 0; i < 30; i++ {
|
||||
if name := fake(t, f, "person"); len(name) < 3 || !regexp.MustCompile(`\S \S`).MatchString(name) {
|
||||
t.Fatalf("%s person %q lacks first and last name", dir, name)
|
||||
if name := fake(t, f, path); len(name) < 3 || !regexp.MustCompile(`\S \S`).MatchString(name) {
|
||||
t.Fatalf("%s %q lacks first and last name", path, name)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestShippedUSPhone(t *testing.T) {
|
||||
f := newGenerator(t, "data/en_US", WithSeed(11))
|
||||
f := newGenerator(t, "data", WithSeed(11))
|
||||
re := regexp.MustCompile(`^(\(\d{3}\) \d{3}-\d{4}|\d{3}-\d{3}-\d{4})$`)
|
||||
for i := 0; i < 50; i++ {
|
||||
if n := fake(t, f, "phone"); !re.MatchString(n) {
|
||||
if n := fake(t, f, "en_US.phone"); !re.MatchString(n) {
|
||||
t.Fatalf("phone %q does not match %s", n, re)
|
||||
}
|
||||
}
|
||||
@@ -270,11 +269,11 @@ func luhnValid(s string) bool {
|
||||
// tests so shipped name lists can grow without re-enumerating them here.
|
||||
var swedishName = regexp.MustCompile(`^\p{L}+([ -]\p{L}+)*$`)
|
||||
|
||||
func TestShippedStreetComposition(t *testing.T) {
|
||||
// street is a choice of composed {first}{last} templates and literal names.
|
||||
f := newGenerator(t, "data/sv_SE", WithSeed(5))
|
||||
func TestShippedStreetIsARegisteredName(t *testing.T) {
|
||||
f := newGenerator(t, "data", WithSeed(5))
|
||||
street := regexp.MustCompile(`^\p{L}[\p{L}\d:.-]*([ -][\p{L}\d:.-]+)*$`)
|
||||
for i := 0; i < 300; i++ {
|
||||
if s := fake(t, f, "address.street"); !swedishName.MatchString(s) {
|
||||
if s := fake(t, f, "sv_SE.address.street"); !street.MatchString(s) {
|
||||
t.Fatalf("street %q is not a Swedish street name", s)
|
||||
}
|
||||
}
|
||||
@@ -283,9 +282,9 @@ func TestShippedStreetComposition(t *testing.T) {
|
||||
func TestShippedLastNameComposition(t *testing.T) {
|
||||
// last is a choice of patronymic {first}sson templates, compound
|
||||
// {first}{last} templates and literal surnames.
|
||||
f := newGenerator(t, "data/sv_SE", WithSeed(6))
|
||||
f := newGenerator(t, "data", WithSeed(6))
|
||||
for i := 0; i < 300; i++ {
|
||||
if s := fake(t, f, "person.last"); !swedishName.MatchString(s) {
|
||||
if s := fake(t, f, "sv_SE.person.last"); !swedishName.MatchString(s) {
|
||||
t.Fatalf("last name %q is not a Swedish surname", s)
|
||||
}
|
||||
}
|
||||
@@ -293,10 +292,10 @@ func TestShippedLastNameComposition(t *testing.T) {
|
||||
|
||||
func TestShippedStreetNumberFormats(t *testing.T) {
|
||||
// Reachable via a hyphenated path; covers all five weighted number variants.
|
||||
f := newGenerator(t, "data/sv_SE", WithSeed(8))
|
||||
f := newGenerator(t, "data", WithSeed(8))
|
||||
re := regexp.MustCompile(`^[1-9]\d{0,2}[A-Z]?$`)
|
||||
for i := 0; i < 300; i++ {
|
||||
if n := fake(t, f, "address.street-number"); !re.MatchString(n) {
|
||||
if n := fake(t, f, "sv_SE.address.street-number"); !re.MatchString(n) {
|
||||
t.Fatalf("street-number %q does not match %s", n, re)
|
||||
}
|
||||
}
|
||||
|
||||
+4
-4
@@ -64,18 +64,18 @@ func TestNewReportsAnEntropyFailure(t *testing.T) {
|
||||
}
|
||||
|
||||
func TestWithSeedIsDeterministic(t *testing.T) {
|
||||
a, b := newGenerator(t, "data/sv_SE", WithSeed(42)), newGenerator(t, "data/sv_SE", WithSeed(42))
|
||||
a, b := newGenerator(t, "data", WithSeed(42)), newGenerator(t, "data", WithSeed(42))
|
||||
for i := 0; i < 50; i++ {
|
||||
if x, y := fake(t, a, "person"), fake(t, b, "person"); x != y {
|
||||
if x, y := fake(t, a, "sv_SE.person"), fake(t, b, "sv_SE.person"); x != y {
|
||||
t.Fatalf("same seed diverged at %d: %q != %q", i, x, y)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestDifferentSeedsDiffer(t *testing.T) {
|
||||
a, b := newGenerator(t, "data/en_US", WithSeed(1)), newGenerator(t, "data/en_US", WithSeed(2))
|
||||
a, b := newGenerator(t, "data", WithSeed(1)), newGenerator(t, "data", WithSeed(2))
|
||||
for i := 0; i < 50; i++ {
|
||||
if fake(t, a, "person") != fake(t, b, "person") {
|
||||
if fake(t, a, "en_US.person") != fake(t, b, "en_US.person") {
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
@@ -63,6 +63,9 @@ func contained(n node) []namedNode {
|
||||
case *table:
|
||||
return append([]namedNode{{node: n.format}}, named(n.fields)...)
|
||||
case *column:
|
||||
if len(n.t.tokens) == 0 {
|
||||
return nil
|
||||
}
|
||||
var out []namedNode
|
||||
for r := 0; r < n.t.rows(); r++ {
|
||||
if cell := n.t.cellNode(r, n.i); cell != nil {
|
||||
|
||||
+2
-2
@@ -130,8 +130,8 @@ func TestPathKeyIsUnambiguous(t *testing.T) {
|
||||
// single-variant choice always picks the same item, so it needs no every-variant
|
||||
// guard and the error can name the field that is missing.
|
||||
func TestMissingFieldNamesItself(t *testing.T) {
|
||||
f := newGenerator(t, "data/sv_SE", WithSeed(1))
|
||||
_, err := f.Fake("person.typo")
|
||||
f := newGenerator(t, "data", WithSeed(1))
|
||||
_, err := f.Fake("sv_SE.person.typo")
|
||||
if err == nil || !strings.Contains(err.Error(), `no field "typo"`) {
|
||||
t.Errorf("Fake(person.typo) = %v, want it to name the missing field", err)
|
||||
}
|
||||
|
||||
@@ -285,7 +285,7 @@ func (t *table) checkCells() error {
|
||||
if (col == t.key || col == t.name) && strings.ContainsAny(cell, inSelector) {
|
||||
return fmt.Errorf("line %d: %s %q contains %q, which a selector cannot spell", row+2, t.columns[col], cell, cell[strings.IndexAny(cell, inSelector):][:1])
|
||||
}
|
||||
if !strings.ContainsAny(cell, "{}") {
|
||||
if strings.IndexByte(cell, '{') < 0 && strings.IndexByte(cell, '}') < 0 {
|
||||
continue
|
||||
}
|
||||
n, err := compileString(cell)
|
||||
|
||||
@@ -44,6 +44,20 @@ var (
|
||||
regionOf = map[string]string{"0180": "01", "0184": "01", "1280": "12", "1281": "12", "1480": "14"}
|
||||
)
|
||||
|
||||
func siblings() map[string]string {
|
||||
return with(geo(), map[string]string{
|
||||
"postal-code.json": `{"format":"{code}","rows":"postal-code.tsv","key":"code","parent":"locality"}`,
|
||||
"postal-code.tsv": "code\tlocality\n111 20\tL1\n111 21\tL1\n171 41\tL2\n211 20\tL3\n221 00\tL4\n223 50\tL4\n411 01\tL5\n417 05\tL6\n247 45\tL7\n170 71\tL8\n",
|
||||
"street.json": `{"format":"{name}","rows":"street.tsv","parent":"locality","weight":"segments"}`,
|
||||
"street.tsv": "name\tlocality\tsegments\nDrottninggatan\tL1\t30\nSveavägen\tL1\t12\nRåsundavägen\tL2\t8\nStorgatan\tL3\t20\nStora Södergatan\tL4\t9\nKlostergatan\tL4\t4\nAvenyn\tL5\t15\nHisingsgatan\tL6\t3\nSandbyvägen\tL7\t2\nSandbyvägen\tL8\t2\n",
|
||||
})
|
||||
}
|
||||
|
||||
var (
|
||||
localityOfCode = map[string]string{"111 20": "L1", "111 21": "L1", "171 41": "L2", "211 20": "L3", "221 00": "L4", "223 50": "L4", "411 01": "L5", "417 05": "L6", "247 45": "L7", "170 71": "L8"}
|
||||
localityOfStreet = map[string][]string{"Drottninggatan": {"L1"}, "Sveavägen": {"L1"}, "Råsundavägen": {"L2"}, "Storgatan": {"L3"}, "Stora Södergatan": {"L4"}, "Klostergatan": {"L4"}, "Avenyn": {"L5"}, "Hisingsgatan": {"L6"}, "Sandbyvägen": {"L7", "L8"}}
|
||||
)
|
||||
|
||||
func with(files map[string]string, more map[string]string) map[string]string {
|
||||
out := map[string]string{}
|
||||
for k, v := range files {
|
||||
@@ -271,6 +285,41 @@ func TestLinkedTablesDrawApartAcrossGroupsAndRepeats(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestSiblingTablesDrawInsideOneAncestor(t *testing.T) {
|
||||
files := with(siblings(), map[string]string{
|
||||
"addr.json": `{"format":"{s}|{p}|{l}","s":"{/street.name}","p":"{/postal-code.code}","l":"{/locality.code}"}`,
|
||||
})
|
||||
f := newGenerator(t, writeFiles(t, files), WithSeed(3))
|
||||
seen := map[string]bool{}
|
||||
for i := 0; i < 300; i++ {
|
||||
parts := strings.Split(fake(t, f, "addr"), "|")
|
||||
if s, p, l := parts[0], parts[1], parts[2]; !slices.Contains(localityOfStreet[s], l) || localityOfCode[p] != l {
|
||||
t.Fatalf("addr = %v, want the street and the postal code inside the locality", parts)
|
||||
}
|
||||
seen[parts[2]] = true
|
||||
}
|
||||
if len(seen) < 4 {
|
||||
t.Fatalf("only %v drawn in 300 renders", seen)
|
||||
}
|
||||
for i := 0; i < 100; i++ {
|
||||
if s := fake(t, f, "locality[L4].street.name"); s != "Stora Södergatan" && s != "Klostergatan" {
|
||||
t.Fatalf("locality[L4].street.name = %q, outside L4", s)
|
||||
}
|
||||
if p := fake(t, f, "region[01].postal-code"); localityOfCode[p] != "L1" && localityOfCode[p] != "L2" && localityOfCode[p] != "L8" {
|
||||
t.Fatalf("region[01].postal-code = %q, outside region 01", p)
|
||||
}
|
||||
}
|
||||
if _, err := f.Fake("street[Avenyn]"); err == nil || !strings.Contains(err.Error(), "no key") {
|
||||
t.Fatalf("Fake(street[Avenyn]) = %v, want no column to select by", err)
|
||||
}
|
||||
paths := f.List()
|
||||
for _, p := range []string{"locality.postal-code", "locality.street", "region.municipality.locality.street.name"} {
|
||||
if !slices.Contains(paths, p) {
|
||||
t.Fatalf("List() lacks %s: %v", p, paths)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestTableSelectorInAReference(t *testing.T) {
|
||||
files := with(geo(), map[string]string{
|
||||
"x.json": `{"format":"{a} {b} {c}","a":"{/region[12].name}","b":"{/region[12].municipality.code}","c":"{/region[12].locality.code}"}`,
|
||||
|
||||
Vendored
+143
-4
@@ -1,11 +1,9 @@
|
||||
en_US.address format "{street-number} {street}\n{locality}, {region} {postal-code}"
|
||||
en_US.address format "{street-number} {street}\n{locality}, {region} {postal-code}" reads geo.US.address
|
||||
en_US.address.locality string
|
||||
en_US.address.postal-code string
|
||||
en_US.address.region string
|
||||
en_US.address.street string
|
||||
en_US.address.street-number string
|
||||
en_US.address.street.name
|
||||
en_US.address.street.suffix
|
||||
en_US.color
|
||||
en_US.company format "{base} {suffix}"
|
||||
en_US.company.base string
|
||||
@@ -52,6 +50,147 @@ en_US.version.n string
|
||||
en_US.version.pre string
|
||||
en_US.version.suffix string
|
||||
en_US.word
|
||||
geo.SE.address format "{street} {street-number}\n{postal-code} {locality}" reads geo.SE.locality geo.SE.postal-code geo.SE.street
|
||||
geo.SE.address.locality string
|
||||
geo.SE.address.postal-code string
|
||||
geo.SE.address.street string
|
||||
geo.SE.address.street-number string
|
||||
geo.SE.locality format "{name}" key name weight population parent municipality
|
||||
geo.SE.locality.lat string
|
||||
geo.SE.locality.lon string
|
||||
geo.SE.locality.municipality string
|
||||
geo.SE.locality.name string
|
||||
geo.SE.locality.population string
|
||||
geo.SE.locality.postal-code
|
||||
geo.SE.locality.postal-code.code
|
||||
geo.SE.locality.postal-code.locality
|
||||
geo.SE.locality.street
|
||||
geo.SE.locality.street.locality
|
||||
geo.SE.locality.street.name
|
||||
geo.SE.locality.street.segments
|
||||
geo.SE.municipality format "{name}" key code name name weight population parent region
|
||||
geo.SE.municipality.code string
|
||||
geo.SE.municipality.locality
|
||||
geo.SE.municipality.locality.lat
|
||||
geo.SE.municipality.locality.lon
|
||||
geo.SE.municipality.locality.municipality
|
||||
geo.SE.municipality.locality.name
|
||||
geo.SE.municipality.locality.population
|
||||
geo.SE.municipality.locality.postal-code
|
||||
geo.SE.municipality.locality.postal-code.code
|
||||
geo.SE.municipality.locality.postal-code.locality
|
||||
geo.SE.municipality.locality.street
|
||||
geo.SE.municipality.locality.street.locality
|
||||
geo.SE.municipality.locality.street.name
|
||||
geo.SE.municipality.locality.street.segments
|
||||
geo.SE.municipality.name string
|
||||
geo.SE.municipality.population string
|
||||
geo.SE.municipality.region string
|
||||
geo.SE.postal-code format "{code}" key code parent locality
|
||||
geo.SE.postal-code.code string
|
||||
geo.SE.postal-code.locality string
|
||||
geo.SE.region format "{name}" key code name name weight population
|
||||
geo.SE.region.code string
|
||||
geo.SE.region.municipality
|
||||
geo.SE.region.municipality.code
|
||||
geo.SE.region.municipality.locality
|
||||
geo.SE.region.municipality.locality.lat
|
||||
geo.SE.region.municipality.locality.lon
|
||||
geo.SE.region.municipality.locality.municipality
|
||||
geo.SE.region.municipality.locality.name
|
||||
geo.SE.region.municipality.locality.population
|
||||
geo.SE.region.municipality.locality.postal-code
|
||||
geo.SE.region.municipality.locality.postal-code.code
|
||||
geo.SE.region.municipality.locality.postal-code.locality
|
||||
geo.SE.region.municipality.locality.street
|
||||
geo.SE.region.municipality.locality.street.locality
|
||||
geo.SE.region.municipality.locality.street.name
|
||||
geo.SE.region.municipality.locality.street.segments
|
||||
geo.SE.region.municipality.name
|
||||
geo.SE.region.municipality.population
|
||||
geo.SE.region.municipality.region
|
||||
geo.SE.region.name string
|
||||
geo.SE.region.population string
|
||||
geo.SE.region.timezone string
|
||||
geo.SE.street format "{name}" weight segments parent locality
|
||||
geo.SE.street.locality string
|
||||
geo.SE.street.name string
|
||||
geo.SE.street.segments string
|
||||
geo.US.address format "{street-number} {street}\n{locality}, {region} {postal-code}" reads geo.US.locality geo.US.postal-code geo.US.region geo.US.street
|
||||
geo.US.address.locality string
|
||||
geo.US.address.postal-code string
|
||||
geo.US.address.region string
|
||||
geo.US.address.street string
|
||||
geo.US.address.street-number string
|
||||
geo.US.locality format "{name}" key code name name weight population parent municipality
|
||||
geo.US.locality.code string
|
||||
geo.US.locality.lat string
|
||||
geo.US.locality.lon string
|
||||
geo.US.locality.municipality string
|
||||
geo.US.locality.name string
|
||||
geo.US.locality.population string
|
||||
geo.US.locality.postal-code
|
||||
geo.US.locality.postal-code.addresses
|
||||
geo.US.locality.postal-code.code
|
||||
geo.US.locality.postal-code.locality
|
||||
geo.US.locality.street
|
||||
geo.US.locality.street.addresses
|
||||
geo.US.locality.street.locality
|
||||
geo.US.locality.street.name
|
||||
geo.US.municipality format "{name}" key code name name weight population parent region
|
||||
geo.US.municipality.code string
|
||||
geo.US.municipality.locality
|
||||
geo.US.municipality.locality.code
|
||||
geo.US.municipality.locality.lat
|
||||
geo.US.municipality.locality.lon
|
||||
geo.US.municipality.locality.municipality
|
||||
geo.US.municipality.locality.name
|
||||
geo.US.municipality.locality.population
|
||||
geo.US.municipality.locality.postal-code
|
||||
geo.US.municipality.locality.postal-code.addresses
|
||||
geo.US.municipality.locality.postal-code.code
|
||||
geo.US.municipality.locality.postal-code.locality
|
||||
geo.US.municipality.locality.street
|
||||
geo.US.municipality.locality.street.addresses
|
||||
geo.US.municipality.locality.street.locality
|
||||
geo.US.municipality.locality.street.name
|
||||
geo.US.municipality.name string
|
||||
geo.US.municipality.population string
|
||||
geo.US.municipality.region string
|
||||
geo.US.postal-code format "{code}" key code weight addresses parent locality
|
||||
geo.US.postal-code.addresses string
|
||||
geo.US.postal-code.code string
|
||||
geo.US.postal-code.locality string
|
||||
geo.US.region format "{name}" key abbr name name weight population
|
||||
geo.US.region.abbr string
|
||||
geo.US.region.code string
|
||||
geo.US.region.municipality
|
||||
geo.US.region.municipality.code
|
||||
geo.US.region.municipality.locality
|
||||
geo.US.region.municipality.locality.code
|
||||
geo.US.region.municipality.locality.lat
|
||||
geo.US.region.municipality.locality.lon
|
||||
geo.US.region.municipality.locality.municipality
|
||||
geo.US.region.municipality.locality.name
|
||||
geo.US.region.municipality.locality.population
|
||||
geo.US.region.municipality.locality.postal-code
|
||||
geo.US.region.municipality.locality.postal-code.addresses
|
||||
geo.US.region.municipality.locality.postal-code.code
|
||||
geo.US.region.municipality.locality.postal-code.locality
|
||||
geo.US.region.municipality.locality.street
|
||||
geo.US.region.municipality.locality.street.addresses
|
||||
geo.US.region.municipality.locality.street.locality
|
||||
geo.US.region.municipality.locality.street.name
|
||||
geo.US.region.municipality.name
|
||||
geo.US.region.municipality.population
|
||||
geo.US.region.municipality.region
|
||||
geo.US.region.name string
|
||||
geo.US.region.population string
|
||||
geo.US.region.timezone string
|
||||
geo.US.street format "{name}" weight addresses parent locality
|
||||
geo.US.street.addresses string
|
||||
geo.US.street.locality string
|
||||
geo.US.street.name string
|
||||
misc.car
|
||||
misc.car.maker
|
||||
misc.car.model
|
||||
@@ -93,7 +232,7 @@ misc.timezone
|
||||
misc.useragent
|
||||
misc.uuid format "{hex(8)}-{hex(4)}-4{hex(3)}-{variant}{hex(3)}-{hex(12)}"
|
||||
misc.uuid.variant string
|
||||
sv_SE.address format "{street} {street-number}\n{postal-code} {locality}"
|
||||
sv_SE.address format "{street} {street-number}\n{postal-code} {locality}" reads geo.SE.address
|
||||
sv_SE.address.locality string
|
||||
sv_SE.address.postal-code string
|
||||
sv_SE.address.street string
|
||||
|
||||
@@ -53,19 +53,25 @@ countries; the README maps each to the native term.
|
||||
|
||||
| Table | SE | US | Weight |
|
||||
|---|---|---|---|
|
||||
| `region` | län (21) | state (56) | population |
|
||||
| `municipality` | kommun (290) | county (3,234) | population |
|
||||
| `locality` | postort (~1,780), tätort population | place (~19,500) | population |
|
||||
| `postal-code` | postnummer (~10,500 deliverable) | ZCTA (33,791) | address count or 1 |
|
||||
| `street` | gatunamn, top N per locality | street name, top N per place | address count |
|
||||
| `region` | län (21) | state and DC (50; Hawaii has no incorporated place) | population |
|
||||
| `municipality` | kommun (290) | county with a shipped place | population |
|
||||
| `locality` | postort, tätort population | place of 25,000+ | population |
|
||||
| `postal-code` | postnummer with street delivery | ZCTA of a shipped place | one; address ranges |
|
||||
| `street` | gatunamn, top 10 per postort | street name, top 10 per place | segments; address ranges |
|
||||
|
||||
- `geo.SE.address` is a record over one consistent draw: street, number, postal
|
||||
code, locality; `geo.SE.locality[Lund].address` stays inside Lund. Each region
|
||||
row carries its timezone, each locality its centroid.
|
||||
- Shipped in step 2, README Data. `geo.SE.address` is a record over one consistent
|
||||
draw. Each region row carries its timezone, each locality its centroid.
|
||||
- Let `geo.SE.locality[Lund].address` descend from a selected row into the
|
||||
template beside the family: the path step must reach a sibling category and the
|
||||
outer selector's pins seed every draw group of the render.
|
||||
- Ship the fuller sets, every US place of 10,000 and more streets per locality, as
|
||||
packs; `--min-population` and `--streets-per-locality` on the scripts build them.
|
||||
- Fill the 398 Swedish localities weighted 200 from SCB småorter before v0.1.0.
|
||||
- Give the address records one column set across countries: `region` and
|
||||
`municipality` as building-block columns on `geo.SE.address` too, in step 6.
|
||||
- v0.1.0 ships SE and US; then NO, DK, FI, NL, FR, AU, CA, ES, GB, DE.
|
||||
- SE streets come from NVDB per kommun and postnummer from GeoNames per postort,
|
||||
box codes dropped by the third-digit rule; the pairing is approximate. Revisit
|
||||
an application to Lantmäteriet for the exact pairing after v0.1.0.
|
||||
- Revisit an application to Lantmäteriet for the exact street to postnummer
|
||||
pairing after v0.1.0; today a street goes to the nearest postal code centroid.
|
||||
|
||||
#### Builtins the data cannot express
|
||||
|
||||
@@ -157,8 +163,8 @@ address, phone, national id, company and date names each.
|
||||
#### Order of work
|
||||
|
||||
1. Table node, key and name selection, parent links, consistent draws, the
|
||||
choice-of-rows fence, `DATA-LICENSES.md`, `data-import/`.
|
||||
2. `geo/SE` and `geo/US`, and `address` in both locales on top of them.
|
||||
choice-of-rows fence, `DATA-LICENSES.md`, `data-import/` — done.
|
||||
2. `geo/SE` and `geo/US`, and `address` in both locales on top of them — done.
|
||||
3. Weighted person names and valid ids in both locales; `date()`.
|
||||
4. `misc` conversions and the new `misc` tables.
|
||||
5. The remaining locale categories: company, phone, finance, vehicle, words.
|
||||
|
||||
Reference in New Issue
Block a user