diff --git a/AGENTS.md b/AGENTS.md index 5826bcb..abdbe46 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,6 +1,6 @@ # Rules -- Go runs only through `docker compose run --rm `; the merge gate is `docker build .` plus CI's changelog check. +- Go runs only through `docker compose run --rm `; the merge gate is `docker build .` plus CI's changelog check. - Tests first, in their own commit; the implementation follows in the next. A re-pin of seeded output or of `testdata/shipped_shape.txt` is its own commit. - A change under `data/` or to the shape pin adds its `CHANGELOG.md` entry under `Unreleased` in the same PR, and so does a change to a flag, an exit code, an exported name, a fence, a builtin or the lowest Go; what is major is the README's Versioning table. - One-line commit messages: no ticket prefix, no repo name, no authorship trailers. diff --git a/CHANGELOG.md b/CHANGELOG.md index 919441d..048d111 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -44,7 +44,7 @@ replacement, and each removed path, column or flag. name both sexes carry is a row under each. `person` reads them, so its columns are `first`, `last`, `prefix` and `sex`; `person.femalefirst` and `person.malefirst` are no longer paths — draw `sex[f].first-name` and `sex[m].first-name` instead. - `en_US.title` links to `sex` as well, so `en_US.person.prefix` draws `Mr` or `Ms` + `en_US.title` links to `sex` as well, so `en_US.person.prefix` draws `Mr` or `Ms` without contradicting the record's `sex`; `sv_SE.title` is an unsexed table of the same shape. - `sv_SE.personnummer` and `sv_SE.samordningsnummer`, Skatteverket's test series diff --git a/DATA-LICENSES.md b/DATA-LICENSES.md index ea9fd1f..1436f95 100644 --- a/DATA-LICENSES.md +++ b/DATA-LICENSES.md @@ -14,7 +14,7 @@ Every shipped dataset, its source, its licence and the attribution it asks for. | `sv_SE/first-name.tsv`, `last-name.tsv` | [SCB](https://www.scb.se/) names with at least two bearers, 31 December 2022 | CC0 1.0 | "Källa: SCB" | `data-import/names-se.py` | | `en_US/first-name.tsv` | [SSA](https://www.ssa.gov/oact/babynames/) baby names, births 1930 to 2020, through [hackerb9/ssa-baby-names](https://github.com/hackerb9/ssa-baby-names) | public domain | none required | `data-import/names-us.py` | | `en_US/last-name.tsv` | Census Bureau surnames occurring 100 or more times, 2010 | public domain | none required | `data-import/names-us.py` | -| `sv_SE/sex.tsv`, `sv_SE/birth-number.tsv`, `en_US/sex.tsv` | curated (Skatteverket's test birth numbers are facts) | — | — | — | +| `sv_SE/sex.tsv`, `sv_SE/birth-number.tsv`, `sv_SE/title.tsv`, `en_US/sex.tsv`, `en_US/title.tsv` | curated (Skatteverket's test birth numbers are facts) | — | — | — | | `misc/country.tsv` | [datasets/country-codes](https://github.com/datasets/country-codes) | [PDDL 1.0](https://opendatacommons.org/licenses/pddl/1-0/) | none required | `data-import/country.py` | | `misc/currency.tsv` | [datasets/currency-codes](https://github.com/datasets/currency-codes); symbols from [Unicode CLDR](https://github.com/unicode-org/cldr) `en.xml` and `root.xml` | PDDL 1.0; [Unicode License v3](https://www.unicode.org/license.txt) | CLDR: "Copyright © 1991-2025 Unicode, Inc. Unicode and the Unicode Logo are registered trademarks of Unicode, Inc. in the United States and other countries." | `data-import/currency.py` | | `misc/httpstatus.tsv` | curated (IANA HTTP status codes are facts) | — | — | — | diff --git a/README.md b/README.md index 6799a1e..4837c45 100644 --- a/README.md +++ b/README.md @@ -1097,8 +1097,10 @@ App developers writing tests and fixtures, in Go and at a shell: - **A title is a table under `sex`.** A prefix drawn apart would put `Mr` on a record whose `sex` column says `female`, which is the disagreement the record exists to prevent; the tables this set already has are what a title needs, so `en_US.title` - links to `sex` as `first-name` does. Swedish has no everyday sexed honorific, so - `sv_SE.title` is a table as well but carries no `parent`. + links to `sex` as `first-name` does. Swedish has no everyday sexed honorific, so + `sv_SE.title` is a table as well but carries no `parent`. Its weight column is + `share`, not the `count` a name table carries, because the values are a curated + proportion rather than bearers anyone counted. - **A table owns the spelling of a selector on it.** A reference reaches a table by a path that carries no selector — `sv_SE.person.first` reads `first-name` through `sex` — so the walk that resolved a name cannot say where a reader would type one. @@ -1124,10 +1126,9 @@ App developers writing tests and fixtures, in Go and at a shell: than computed from the date drawn. - **The US given names come from a mirror of the SSA file.** ssa.gov refuses a client outside the US, so `names-us.py` reads a GitHub copy that ends at 2020, - which a count over the births since 1930 barely feels; `--names` takes the - official zip. The SSA's placeholder rows — `Unknown`, `Baby`, `Infant` — are - top-1000 entries that name nobody, so the import drops them by name rather than by - a rank a regeneration would move. + which a count over the births since 1930 barely feels; `--names` takes the + official zip. The SSA's placeholder rows are top-1000 entries that name nobody, so + the import drops them by name rather than by a rank a regeneration would move. - **`List` advertises direct descents only.** `region.municipality.locality` is listed, and `region.locality` resolves too but is not: the set of every descent through a chain of five tables is every subsequence of it, and the direct chain is diff --git a/builtins_test.go b/builtins_test.go index 92c585f..adcf039 100644 --- a/builtins_test.go +++ b/builtins_test.go @@ -297,21 +297,21 @@ func TestBuiltinDateAndTime(t *testing.T) { // the layout is quoted, names a field, and for time names no date field. func TestBuiltinDateArgs(t *testing.T) { for tmpl, want := range map[string]string{ - `"{date(1990-13-01,1990-12-31,'2006-01-02')}"`: "1990-13-01", - `"{date(1990-12-31,1990-01-01,'2006-01-02')}"`: "is after", - `"{date(1990-01-01,1990-01-01,'2006-01-02')}"`: "write it as text", - `"{date(1990-01-01,1990-12-31,2006-01-02)}"`: "'2006-01-02'", - `"{date(1990-01-01,1990-12-31,'January 2, 2006)}"`: "'", + `"{date(1990-13-01,1990-12-31,'2006-01-02')}"`: "1990-13-01", + `"{date(1990-12-31,1990-01-01,'2006-01-02')}"`: "is after", + `"{date(1990-01-01,1990-01-01,'2006-01-02')}"`: "write it as text", + `"{date(1990-01-01,1990-12-31,2006-01-02)}"`: "'2006-01-02'", + `"{date(1990-01-01,1990-12-31,'January 2, 2006)}"`: "'", `"{date(1990-01-01,1990-12-31,\"January 2, 2006\")}"`: "quoted: 'January 2, 2006'", - `"{date(1990-01-01,1990-12-31,'x')}"`: "text", - `"{date(1990-01-01,1990-12-31,'')}"`: "text", - `"{date(1990-01-01,1990-12-31)}"`: "3 arguments", - `"{date(1990-01-01,1990-12-31,January 2, 2006)}"`: "'January 2, 2006'", - `"{time(3:04 PM, Mon)}"`: "'3:04 PM, Mon'", - `"{time(15:04)}"`: "'15:04'", - `"{time('2006-01-02 15:04')}"`: "date(", - `"{time('x')}"`: "text", - `"{date(1990-01-01,1990-12-31,'15:04')}"`: "time('15:04')", + `"{date(1990-01-01,1990-12-31,'x')}"`: "text", + `"{date(1990-01-01,1990-12-31,'')}"`: "text", + `"{date(1990-01-01,1990-12-31)}"`: "3 arguments", + `"{date(1990-01-01,1990-12-31,January 2, 2006)}"`: "'January 2, 2006'", + `"{time(3:04 PM, Mon)}"`: "'3:04 PM, Mon'", + `"{time(15:04)}"`: "'15:04'", + `"{time('2006-01-02 15:04')}"`: "date(", + `"{time('x')}"`: "text", + `"{date(1990-01-01,1990-12-31,'15:04')}"`: "time('15:04')", } { _, err := compile(parse(t, tmpl)) if err == nil || !strings.Contains(err.Error(), want) { diff --git a/data-import/geo-se.py b/data-import/geo-se.py index c0d627f..d4fa0e0 100644 --- a/data-import/geo-se.py +++ b/data-import/geo-se.py @@ -17,8 +17,8 @@ import zipfile from pathlib import Path import source -import xlsx import tsv +import xlsx CODES = "https://www.scb.se/contentassets/7a89e48960f741e08918e489ea36354a/kommunlankod-2026.xlsx" POPULATION = "https://api.scb.se/OV0104/v1/doris/sv/ssd/START/BE/BE0101/BE0101A/BefolkningNy" @@ -39,6 +39,8 @@ CACHE = Path(__file__).resolve().parent / "cache" TIMEZONE = "Europe/Stockholm" ONE_POSITION = {"Stockholm", "Göteborg", "Malmö"} UNMATCHED_POPULATION = 200 + + def scb_codes(cache): regions, municipalities = {}, {} for cells in xlsx.rows(source.fetch(CODES, cache, "kommunlankod.xlsx", magic=b"PK")): diff --git a/data-import/names-se.py b/data-import/names-se.py index 6ad6428..eef8b6e 100644 --- a/data-import/names-se.py +++ b/data-import/names-se.py @@ -10,8 +10,8 @@ import re from pathlib import Path import source -import xlsx import tsv +import xlsx SOURCE = "https://www.scb.se/contentassets/9fe7dbb460994c72b835163dbc491ef9/namn-med-minst-tva-barare-31-december-2022.xlsx" OUT = Path(__file__).resolve().parent.parent / "data" / "sv_SE" diff --git a/data-import/source.py b/data-import/source.py index dbcda52..121e9d0 100644 --- a/data-import/source.py +++ b/data-import/source.py @@ -11,9 +11,9 @@ def fetch(source, cache, name, magic=b"", data=None, headers=None): if not re.match(r"^https?://", source): return Path(source).read_bytes() path = Path(cache) / name + if path.exists(): + return path.read_bytes() for attempt in range(1, 6): - if path.exists(): - return path.read_bytes() req = urllib.request.Request(source, data=data, headers={"User-Agent": "fejkdata data-import", **(headers or {})}) try: with urllib.request.urlopen(req, timeout=600) as r: @@ -23,6 +23,7 @@ def fetch(source, cache, name, magic=b"", data=None, headers=None): if body and body.startswith(magic) and b"Request Rejected" not in body[:512]: path.parent.mkdir(parents=True, exist_ok=True) path.write_bytes(body) - elif attempt < 5: + return body + if attempt < 5: time.sleep(10 * attempt) sys.exit(f"{source}: no valid download in 5 attempts") diff --git a/data/sv_SE/last-name.tsv b/data/sv_SE/last-name.tsv index 294345f..669c520 100644 --- a/data/sv_SE/last-name.tsv +++ b/data/sv_SE/last-name.tsv @@ -837,7 +837,6 @@ Cronholm 393 Cronqvist 388 Cronvall 199 Cruz 207 -Da Silva 287 Daabas 247 Dabrowski 222 Dag 327 @@ -4875,6 +4874,7 @@ Zimmermann 230 Zingmark 337 Zivkovic 455 af Ekenstam 204 +da Silva 287 de Geer 200 du Rietz 282 von Bahr 188 diff --git a/layout.go b/layout.go index a8a03c2..a389bc8 100644 --- a/layout.go +++ b/layout.go @@ -66,7 +66,7 @@ func layoutArity(name string, n int, a []string) error { } hint := "" if len(a) > n && !holdsQuotedLayout(a[n-1:]) { - hint = fmt.Sprintf("; a layout holding a comma is quoted: '%s'", strings.Join(a[n-1:], ", ")) + hint = fmt.Sprintf("; a layout holding a comma is quoted: '%s'", strings.Trim(strings.Join(a[n-1:], ", "), `'"`)) } return fmt.Errorf("%s takes %d argument%s, got %d%s", name, n, plural(n), len(a), hint) } diff --git a/table.go b/table.go index 526e0c0..1d3dc50 100644 --- a/table.go +++ b/table.go @@ -337,9 +337,6 @@ func (t *table) compileFormat(format string) error { return t.format.compileFormat() } -// linkTables binds every table's parent to the table beside it, and proves the -// links: a parent has a key, every link cell is one, every parent row is linked -// to, no chain of parents closes, and no child is named like a parent's column. // setTablePaths gives every table the path a selector on it is written at, before a // link or a draw can name one. func setTablePaths(root map[string]node) { @@ -357,6 +354,9 @@ func setTablePaths(root map[string]node) { walk("", root) } +// linkTables binds every table's parent to the table beside it, and proves the +// links: a parent has a key, every link cell is one, every parent row is linked +// to, no chain of parents closes, and no child is named like a parent's column. func linkTables(root map[string]node) error { var walk func(dir string, children map[string]node) error walk = func(dir string, children map[string]node) error { diff --git a/todo.md b/todo.md index 33cc624..4e4b815 100644 --- a/todo.md +++ b/todo.md @@ -166,9 +166,9 @@ address, phone, national id, company and date names each. 2. `geo/SE` and `geo/US`, and `address` in both locales on top of them — done. 3. Weighted person names and valid ids in both locales; `date()` — done. Left for later: middle names, and drawing a shipped `personnummer` inside a *selected* - sex, both of which want a draw group that shares its family's pins — until then - the conflict error can name a rewrite that returns a different value when the - read it conflicts with sits inside another category. Give the national ids one + sex, both of which want a draw group that shares its family's pins; stop the + conflict error naming a rewrite that returns a different value where the read it + conflicts with sits inside another category. Give the national ids one record shape in step 6, and report a struct column's draw conflict with the path spelling a tag takes rather than the reference spelling. 4. `misc` conversions and the new `misc` tables.