Weighted person names, valid personal ids and date() in both locales #21

Merged
lilleman merged 38 commits from person-ids-date into main 2026-09-18 20:45:51 +02:00
12 changed files with 40 additions and 36 deletions
Showing only changes of commit 5fb5e5f4c8 - Show all commits
+1 -1
View File
@@ -1,6 +1,6 @@
# Rules
- Go runs only through `docker compose run --rm <test|vet|fmt|build>`; the merge gate is `docker build .` plus CI's changelog check.
- Go runs only through `docker compose run --rm <test|vet|fmt|build|cyclo>`; the merge gate is `docker build .` plus CI's changelog check.
- Tests first, in their own commit; the implementation follows in the next. A re-pin of seeded output or of `testdata/shipped_shape.txt` is its own commit.
- A change under `data/` or to the shape pin adds its `CHANGELOG.md` entry under `Unreleased` in the same PR, and so does a change to a flag, an exit code, an exported name, a fence, a builtin or the lowest Go; what is major is the README's Versioning table.
- One-line commit messages: no ticket prefix, no repo name, no authorship trailers.
+1 -1
View File
@@ -14,7 +14,7 @@ Every shipped dataset, its source, its licence and the attribution it asks for.
| `sv_SE/first-name.tsv`, `last-name.tsv` | [SCB](https://www.scb.se/) names with at least two bearers, 31 December 2022 | CC0 1.0 | "Källa: SCB" | `data-import/names-se.py` |
| `en_US/first-name.tsv` | [SSA](https://www.ssa.gov/oact/babynames/) baby names, births 1930 to 2020, through [hackerb9/ssa-baby-names](https://github.com/hackerb9/ssa-baby-names) | public domain | none required | `data-import/names-us.py` |
| `en_US/last-name.tsv` | Census Bureau surnames occurring 100 or more times, 2010 | public domain | none required | `data-import/names-us.py` |
| `sv_SE/sex.tsv`, `sv_SE/birth-number.tsv`, `en_US/sex.tsv` | curated (Skatteverket's test birth numbers are facts) | — | — | — |
| `sv_SE/sex.tsv`, `sv_SE/birth-number.tsv`, `sv_SE/title.tsv`, `en_US/sex.tsv`, `en_US/title.tsv` | curated (Skatteverket's test birth numbers are facts) | — | — | — |
| `misc/country.tsv` | [datasets/country-codes](https://github.com/datasets/country-codes) | [PDDL 1.0](https://opendatacommons.org/licenses/pddl/1-0/) | none required | `data-import/country.py` |
| `misc/currency.tsv` | [datasets/currency-codes](https://github.com/datasets/currency-codes); symbols from [Unicode CLDR](https://github.com/unicode-org/cldr) `en.xml` and `root.xml` | PDDL 1.0; [Unicode License v3](https://www.unicode.org/license.txt) | CLDR: "Copyright © 1991-2025 Unicode, Inc. Unicode and the Unicode Logo are registered trademarks of Unicode, Inc. in the United States and other countries." | `data-import/currency.py` |
| `misc/httpstatus.tsv` | curated (IANA HTTP status codes are facts) | — | — | — |
+5 -4
View File
@@ -1098,7 +1098,9 @@ App developers writing tests and fixtures, in Go and at a shell:
whose `sex` column says `female`, which is the disagreement the record exists to
prevent; the tables this set already has are what a title needs, so `en_US.title`
links to `sex` as `first-name` does. Swedish has no everyday sexed honorific, so
`sv_SE.title` is a table as well but carries no `parent`.
`sv_SE.title` is a table as well but carries no `parent`. Its weight column is
`share`, not the `count` a name table carries, because the values are a curated
proportion rather than bearers anyone counted.
- **A table owns the spelling of a selector on it.** A reference reaches a table by a
path that carries no selector — `sv_SE.person.first` reads `first-name` through
`sex` — so the walk that resolved a name cannot say where a reader would type one.
@@ -1125,9 +1127,8 @@ App developers writing tests and fixtures, in Go and at a shell:
- **The US given names come from a mirror of the SSA file.** ssa.gov refuses a
client outside the US, so `names-us.py` reads a GitHub copy that ends at 2020,
which a count over the births since 1930 barely feels; `--names` takes the
official zip. The SSA's placeholder rows — `Unknown`, `Baby`, `Infant` — are
top-1000 entries that name nobody, so the import drops them by name rather than by
a rank a regeneration would move.
official zip. The SSA's placeholder rows are top-1000 entries that name nobody, so
the import drops them by name rather than by a rank a regeneration would move.
- **`List` advertises direct descents only.** `region.municipality.locality` is
listed, and `region.locality` resolves too but is not: the set of every descent
through a chain of five tables is every subsequence of it, and the direct chain is
+3 -1
View File
@@ -17,8 +17,8 @@ import zipfile
from pathlib import Path
import source
import xlsx
import tsv
import xlsx
CODES = "https://www.scb.se/contentassets/7a89e48960f741e08918e489ea36354a/kommunlankod-2026.xlsx"
POPULATION = "https://api.scb.se/OV0104/v1/doris/sv/ssd/START/BE/BE0101/BE0101A/BefolkningNy"
@@ -39,6 +39,8 @@ CACHE = Path(__file__).resolve().parent / "cache"
TIMEZONE = "Europe/Stockholm"
ONE_POSITION = {"Stockholm", "Göteborg", "Malmö"}
UNMATCHED_POPULATION = 200
def scb_codes(cache):
regions, municipalities = {}, {}
for cells in xlsx.rows(source.fetch(CODES, cache, "kommunlankod.xlsx", magic=b"PK")):
+1 -1
View File
@@ -10,8 +10,8 @@ import re
from pathlib import Path
import source
import xlsx
import tsv
import xlsx
SOURCE = "https://www.scb.se/contentassets/9fe7dbb460994c72b835163dbc491ef9/namn-med-minst-tva-barare-31-december-2022.xlsx"
OUT = Path(__file__).resolve().parent.parent / "data" / "sv_SE"
+3 -2
View File
@@ -11,9 +11,9 @@ def fetch(source, cache, name, magic=b"", data=None, headers=None):
if not re.match(r"^https?://", source):
return Path(source).read_bytes()
path = Path(cache) / name
for attempt in range(1, 6):
if path.exists():
return path.read_bytes()
for attempt in range(1, 6):
req = urllib.request.Request(source, data=data, headers={"User-Agent": "fejkdata data-import", **(headers or {})})
try:
with urllib.request.urlopen(req, timeout=600) as r:
@@ -23,6 +23,7 @@ def fetch(source, cache, name, magic=b"", data=None, headers=None):
if body and body.startswith(magic) and b"Request Rejected" not in body[:512]:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_bytes(body)
elif attempt < 5:
return body
if attempt < 5:
time.sleep(10 * attempt)
sys.exit(f"{source}: no valid download in 5 attempts")
+1 -1
View File
@@ -837,7 +837,6 @@ Cronholm 393
Cronqvist 388
Cronvall 199
Cruz 207
Da Silva 287
Daabas 247
Dabrowski 222
Dag 327
@@ -4875,6 +4874,7 @@ Zimmermann 230
Zingmark 337
Zivkovic 455
af Ekenstam 204
da Silva 287
de Geer 200
du Rietz 282
von Bahr 188
1 name count
837 Cronqvist 388
838 Cronvall 199
839 Cruz 207
Da Silva 287
840 Daabas 247
841 Dabrowski 222
842 Dag 327
4874 Zingmark 337
4875 Zivkovic 455
4876 af Ekenstam 204
4877 da Silva 287
4878 de Geer 200
4879 du Rietz 282
4880 von Bahr 188
+1 -1
View File
@@ -66,7 +66,7 @@ func layoutArity(name string, n int, a []string) error {
}
hint := ""
if len(a) > n && !holdsQuotedLayout(a[n-1:]) {
hint = fmt.Sprintf("; a layout holding a comma is quoted: '%s'", strings.Join(a[n-1:], ", "))
hint = fmt.Sprintf("; a layout holding a comma is quoted: '%s'", strings.Trim(strings.Join(a[n-1:], ", "), `'"`))
}
return fmt.Errorf("%s takes %d argument%s, got %d%s", name, n, plural(n), len(a), hint)
}
+3 -3
View File
@@ -337,9 +337,6 @@ func (t *table) compileFormat(format string) error {
return t.format.compileFormat()
}
// linkTables binds every table's parent to the table beside it, and proves the
// links: a parent has a key, every link cell is one, every parent row is linked
// to, no chain of parents closes, and no child is named like a parent's column.
// setTablePaths gives every table the path a selector on it is written at, before a
// link or a draw can name one.
func setTablePaths(root map[string]node) {
@@ -357,6 +354,9 @@ func setTablePaths(root map[string]node) {
walk("", root)
}
// linkTables binds every table's parent to the table beside it, and proves the
// links: a parent has a key, every link cell is one, every parent row is linked
// to, no chain of parents closes, and no child is named like a parent's column.
func linkTables(root map[string]node) error {
var walk func(dir string, children map[string]node) error
walk = func(dir string, children map[string]node) error {
+3 -3
View File
@@ -166,9 +166,9 @@ address, phone, national id, company and date names each.
2. `geo/SE` and `geo/US`, and `address` in both locales on top of them — done.
3. Weighted person names and valid ids in both locales; `date()` — done. Left for
later: middle names, and drawing a shipped `personnummer` inside a *selected*
sex, both of which want a draw group that shares its family's pins — until then
the conflict error can name a rewrite that returns a different value when the
read it conflicts with sits inside another category. Give the national ids one
sex, both of which want a draw group that shares its family's pins; stop the
conflict error naming a rewrite that returns a different value where the read it
conflicts with sits inside another category. Give the national ids one
record shape in step 6, and report a struct column's draw conflict with the path
spelling a tag takes rather than the reference spelling.
4. `misc` conversions and the new `misc` tables.