Weighted person names, valid personal ids and date() in both locales #21

Merged
lilleman merged 38 commits from person-ids-date into main 2026-09-18 20:45:51 +02:00
65 changed files with 19454 additions and 507 deletions
+2 -2
View File
@@ -1,6 +1,6 @@
# Rules
- Go runs only through `docker compose run --rm <test|vet|fmt|build>`; the merge gate is `docker build .` plus CI's changelog check.
- Go runs only through `docker compose run --rm <test|vet|fmt|build|cyclo>`; the merge gate is `docker build .` plus CI's changelog check.
- Tests first, in their own commit; the implementation follows in the next. A re-pin of seeded output or of `testdata/shipped_shape.txt` is its own commit.
- A change under `data/` or to the shape pin adds its `CHANGELOG.md` entry under `Unreleased` in the same PR, and so does a change to a flag, an exit code, an exported name, a fence, a builtin or the lowest Go; what is major is the README's Versioning table.
- One-line commit messages: no ticket prefix, no repo name, no authorship trailers.
@@ -8,4 +8,4 @@
- One spelling per result: reject the other at `New`, and let the error name the spelling to use.
- A standing choice a reader would relitigate goes under Decisions in the README, not in a comment.
- A README example is a `json` block that loads and renders as a category; `readme_test.go` runs every one.
- Cyclomatic complexity is gated at 14: the table-shaped dispatches (`eachToken`, `calc.factor`, `walkPath`) sit at 13–14 and stay whole; anything else that reaches 14 is decomposed.
- Cyclomatic complexity is gated at 14: the table-shaped dispatches (`linkParent`, `compileTemplate`, `renderEdges`) sit at it and stay whole; a function that would pass it is decomposed.
+31 -2
View File
@@ -13,8 +13,9 @@ replacement, and each removed path, column or flag.
`key`, `name`, `weight` and `parent`; a path selects a row by key or name,
`misc.country[SE]`, and descends to a linked table by name; linked tables draw
consistently within one render and draw group. `rows` is an option, so no
template may carry a field of that name. Refused at `New`: a `name` without a
`key`, a name spelling another row's key, a table named like a column of any
template may carry a field of that name. A `name` without a `key` resolves
inside the table's `parent`. Refused at `New`: a `name` without a `key` or a
`parent`, a name repeating inside one parent row, a name spelling another row's key, a table named like a column of any
table above it, a table whose format or cell references its own family, and,
within one render and draw group, a path drawing a table another path selects a
row of, or two paths pinning different rows of one table.
@@ -32,3 +33,31 @@ replacement, and each removed path, column or flag.
`address` record over one consistent draw of them. `sv_SE.address` and
`en_US.address` read those records, so `en_US.address.street` no longer carries
`name` and `suffix`, and a locale folder loads only beside `geo`.
- `{date(from,to,'layout')}` and `{time('layout')}`: a second between two days, or
within one, in a single-quoted Go layout, drawn in UTC; `from` may equal `to`.
`sv_SE.date`, `en_US.date`, `sv_SE.time` and `en_US.time` render through them, so
`date.year`, `date.month`, `date.day`, `time.hour`, `time.minute`, `time.minute.t`,
`time.sec` and `time.ampm` are no longer paths — a part of a date is now its own
`{date(…,'2006')}`. `misc.datetime` is an RFC 3339 instant.
- `sex`, `first-name` and `last-name` tables in `sv_SE` and `en_US`, weighted by
bearers from SCB, the SSA and the Census Bureau; `first-name` links to `sex`, and a
name both sexes carry is a row under each. `person` reads them, so its columns are
`first`, `last`, `prefix` and `sex`; `person.femalefirst` and `person.malefirst` are
no longer paths — draw `sex[f].first-name` and `sex[m].first-name` instead.
`en_US.title` links to `sex` as well, so `en_US.person.prefix` draws `Mr` or `Ms`
without contradicting the record's `sex`; `sv_SE.title` is an unsexed table of the
same shape.
- `sv_SE.personnummer` and `sv_SE.samordningsnummer`, Skatteverket's test series
from a `sv_SE.birth-number` table under `sex`, in place of `sv_SE.ssn`, whose
`ssn.mmdd`, `ssn.mmdd.m` and `ssn.mmdd.d` go with it. `en_US.ssn` now draws the
ranges the SSA assigns and carries the columns `area`, `group` and `serial`, so
`--format csv en_US.ssn` writes a header where it used to fail; `en_US.itin` is new
and carries the same three. `person.prefix` is null where a person has no title,
where it used to be an empty string, so `--format sql` writes `NULL` and a `string`
struct field reading it becomes `*string`.
- `ErrNoColumns` is exported, so a caller can tell the one record fence a path can
answer from the rest.
- An error names a spelling that runs: a layout is named single-quoted and free of
its own quotes, a row of a table with no key is named as the path that selects it,
`sv_SE.sex[f].first-name[Kim]`, and a category with no columns names the record
that gives it one.
+4
View File
@@ -11,6 +11,10 @@ Every shipped dataset, its source, its licence and the attribution it asks for.
| `geo/SE/street.tsv` | [Trafikverket NVDB](https://www.trafikverket.se/) Gatunamn, through the open API | CC0 1.0 | none required | `data-import/geo-se.py` |
| `geo/US/region.tsv`, `municipality.tsv`, `locality.tsv` | [Census Bureau](https://www.census.gov/) Gazetteer 2026 and population estimates 2025 | [public domain](https://www.usa.gov/government-works) | none required | `data-import/geo-us.py` |
| `geo/US/postal-code.tsv`, `street.tsv` | Census Bureau ZCTA to place relationships 2020 and TIGER/Line 2025 address ranges and feature names | public domain | none required | `data-import/geo-us.py` |
| `sv_SE/first-name.tsv`, `last-name.tsv` | [SCB](https://www.scb.se/) names with at least two bearers, 31 December 2022 | CC0 1.0 | "Källa: SCB" | `data-import/names-se.py` |
| `en_US/first-name.tsv` | [SSA](https://www.ssa.gov/oact/babynames/) baby names, births 1930 to 2020, through [hackerb9/ssa-baby-names](https://github.com/hackerb9/ssa-baby-names) | public domain | none required | `data-import/names-us.py` |
| `en_US/last-name.tsv` | Census Bureau surnames occurring 100 or more times, 2010 | public domain | none required | `data-import/names-us.py` |
| `sv_SE/sex.tsv`, `sv_SE/birth-number.tsv`, `sv_SE/title.tsv`, `en_US/sex.tsv`, `en_US/title.tsv` | curated (Skatteverket's test birth numbers are facts) | — | — | — |
| `misc/country.tsv` | [datasets/country-codes](https://github.com/datasets/country-codes) | [PDDL 1.0](https://opendatacommons.org/licenses/pddl/1-0/) | none required | `data-import/country.py` |
| `misc/currency.tsv` | [datasets/currency-codes](https://github.com/datasets/currency-codes); symbols from [Unicode CLDR](https://github.com/unicode-org/cldr) `en.xml` and `root.xml` | PDDL 1.0; [Unicode License v3](https://www.unicode.org/license.txt) | CLDR: "Copyright © 1991-2025 Unicode, Inc. Unicode and the Unicode Logo are registered trademarks of Unicode, Inc. in the United States and other countries." | `data-import/currency.py` |
| `misc/httpstatus.tsv` | curated (IANA HTTP status codes are facts) | — | — | — |
+168 -63
View File
@@ -22,6 +22,7 @@ fejkdata 'geo.SE.locality[Lund].street' # Fjelievägen — a linked tabl
fejkdata --data-path ./mydata sv_SE.word # layer a directory over the shipped data
fejkdata --no-shipped-data -d ./mydata --list # only your data
fejkdata 'name: {/sv_SE.person.last}' # name: <a surname> — an inline template
fejkdata "{date(1990-01-01,2010-12-31,'2006-01-02')}" # 2003-11-27 — the argument in "…", the layout in '…'
fejkdata '{"format":"name: {x}","x":["bosse","lina"]}' # name: bosse or name: lina
```
@@ -32,11 +33,9 @@ a JSON object, array or string, or that carries a `{` token, is instead an
**inline template**: a format string or a JSON value compiled and rendered on the
spot. Its tokens reach the data by reference from the root —
`{/sv_SE.person.last}`, so shipped and `--data-path` categories are alike
available. An inline template sits in no folder, so the folder-relative `{.name}`
and `{..name}` are rejected naming the root spelling, and one reference alone —
`{/sv_SE.person}` — is the path written as a template, rejected naming the path, as is
a path written `/sv_SE.person`. A path never contains a brace or a quote, and a
bracket only as a selector after a name, so the two cannot collide (see [Decisions](#decisions)).
available. A path never contains a brace or a quote, and a bracket only as a selector
after a name, so the two spellings cannot collide; which spellings an inline template
rejects, and what each names instead, is under [Decisions](#decisions).
| Flag | |
|------|--|
@@ -114,6 +113,12 @@ the template itself, which composes the format into one string rather than
projecting columns — ask for more records with `--repeat`. A `repeat` on a column
is fine.
Columns are written in name order, whatever order the fields appear in. A category
whose fields are the parts of one value — `sv_SE.price`, `sv_SE.version`, `misc.uuid` —
projects those parts rather than the value, so for one column holding what `Fake`
renders, write `{"format":"","price":"{/sv_SE.price}"}` as a fieldless category asks
for.
A column carrying a newline keeps it inside the quoted CSV field or the SQL string
literal, so a row can span physical lines: read the stream with a CSV or SQL
parser rather than splitting it on newlines.
@@ -141,6 +146,72 @@ a value two fields share in its own category and reference that. A field hold, t
operand ties fields together within one column as always (see
[Correlated fields](#correlated-fields) and [Decisions](#decisions)).
## Data
The shipped set under [`data/`](data) — one folder per locale (`en_US`, `sv_SE`)
plus a locale-neutral `misc` folder — is embedded, so the CLI and the library
work with no data on disk. A directory is a namespace: each JSON file is a
category named after the file, each subdirectory a dot-path segment, so
`mydata/sv_SE/person.json` is `sv_SE.person` and replaces the shipped one.
Sources merge in order; matching folders combine, any other clash is won by the
last loaded. Names may not use `.`, `|`, `(`, `{`, `}`, `[`, `]`, `"` or `/`, nor be
`-`, which a struct tag reserves; dot-prefixed entries are skipped, so a data directory can also be a checkout.
Each locale carries `address`, `color`, `company`, `date`, `email`, `first-name`,
`ip`, `last-name`, `person`, `phone`, `price`, `sentence`, `sex`, `time`, `url`,
`username`, `version` and `word`, formatted per locale; `sv_SE` adds
`personnummer` and `samordningsnummer`, `en_US` adds `ssn` and `itin`. `misc`
carries `car`, `coordinate`, `country` (ISO 3166), `creditcard` (Luhn-valid),
`currency` (ISO 4217), `datetime` (RFC 3339), `emoji`, `httpstatus`, `language`
(ISO 639), `mac`, `mimetype`, `objectid`, `timezone` (IANA), `useragent` and
`uuid` (v4). Many carry sub-fields — `misc.currency.symbol`,
`misc.country.alpha2`, `misc.httpstatus.code` — which `--list` shows. `country`,
`currency`, `httpstatus`, `language` and `mimetype` are [tables](#table), so
`misc.country[SE].capital` and `misc.currency[Euro].symbol` select a row;
[`DATA-LICENSES.md`](DATA-LICENSES.md) names each table's source and licence.
`sex`, `first-name` and `last-name` are tables weighted by bearers, from SCB, the
SSA and the Census Bureau. `first-name` links to `sex`, so `sv_SE.sex[f].first-name`
draws a woman's name, and a name both sexes carry is a row under each, so
`en_US.sex[m].first-name[Taylor]` names the one a `first-name[Taylor]` alone cannot.
`person` reads one draw of the three, so its `first` and `sex` columns agree, and so
does a `personnummer` in the same render: its birth number, `sv_SE.birth-number`
under `sex`, is Skatteverket's test series, 238 for a woman and 239 for a man, which no
real person is ever given. `en_US.title` links to `sex` too, so a person's prefix
never contradicts it.
A person of a chosen sex is assembled from the tables — `sex[f].first-name` beside
`last-name` — while a shipped `personnummer` agrees with the sex its own render
*drew*, not with one a path selects. That test series is also small: a personnummer
is one of about 70,000 values, a day in 1930–2025 against the two birth numbers, so a
fixture past a few hundred rows repeats one and a `UNIQUE` column needs a category of
your own. `sv_SE.date` and `en_US.date` are uniform over 1970-01-01 to 2029-12-31,
`misc.datetime` over 2000-01-01 to 2029-12-31. What the two locales do not share:
`en_US.address` carries a `region` column the Swedish one has no use for, and
`sv_SE.title` has no `parent`, so it is selected as `sv_SE.title[dr]` rather than
inside a sex.
A `geo` folder holds one tree per country under its alpha-2 code: five
[linked tables](#linked-tables) named alike, and an `address` record over one
consistent draw of them, which the locale's `address` reads.
| Table | `geo.SE` | `geo.US` | Weight |
|-------|----------|----------|--------|
| `region` | län, by code or name | state, by USPS abbreviation or name; `code` is the FIPS code | population |
| `municipality` | kommun, by code or name | county, by FIPS code or name | population |
| `locality` | postort, by name | incorporated place of 25,000 people or more with a postal code of its own, by GEOID or name; Hawaii has none | tätort population, the kommun's where the postort names it, else 200; place population |
| `postal-code` | postnummer with street delivery, by code | ZCTA, by code | one; address ranges |
| `street` | gatunamn, the ten with most road segments per postort | street name, the ten with most address ranges per place | segments; address ranges |
`geo.SE.region[Skåne län].municipality` draws a kommun in Skåne,
`geo.SE.locality[Lund].street` a street in Lund, and
`geo.US.region[IL].locality[Springfield]` settles which Springfield. A region row
carries its `timezone`, the state's predominant zone, and a locality its `lat` and
`lon`. What ports across countries is the five table names, the `name` column,
selection by name, and the `address` record's columns `street`, `street-number`,
`postal-code` and `locality`; every other column is the country's own, `code` on a
Swedish region but `abbr` on a US one.
## Library
```sh
@@ -206,49 +277,6 @@ A `*Generator` is safe for concurrent use; a seeded sequence is reproducible onl
when drawn from one goroutine. Changing how a value is composed shifts the seeded
stream for that value and everything drawn after it.
## Data
The shipped set under [`data/`](data) — one folder per locale (`en_US`, `sv_SE`)
plus a locale-neutral `misc` folder — is embedded, so the CLI and the library
work with no data on disk. A directory is a namespace: each JSON file is a
category named after the file, each subdirectory a dot-path segment, so
`mydata/sv_SE/person.json` is `sv_SE.person` and replaces the shipped one.
Sources merge in order; matching folders combine, any other clash is won by the
last loaded. Names may not use `.`, `|`, `(`, `{`, `}`, `[`, `]`, `"` or `/`, nor be
`-`, which a struct tag reserves; dot-prefixed entries are skipped, so a data directory can also be a checkout.
Each locale carries `address`, `color`, `company`, `date`, `email`, `ip`,
`person`, `phone`, `price`, `sentence`, `ssn`, `time`, `url`, `username`,
`version` and `word`, formatted per locale. `misc` carries `car`, `coordinate`,
`country` (ISO 3166), `creditcard` (Luhn-valid), `currency` (ISO 4217), `emoji`,
`httpstatus`, `language` (ISO 639), `mac`, `mimetype`, `objectid`, `timezone`
(IANA), `useragent` and `uuid` (v4). Many carry sub-fields — `misc.currency.symbol`,
`misc.country.alpha2`, `misc.httpstatus.code` — which `--list` shows. `country`,
`currency`, `httpstatus`, `language` and `mimetype` are [tables](#table), so
`misc.country[SE].capital` and `misc.currency[Euro].symbol` select a row;
[`DATA-LICENSES.md`](DATA-LICENSES.md) names each table's source and licence.
A `geo` folder holds one tree per country under its alpha-2 code: five
[linked tables](#linked-tables) named alike, and an `address` record over one
consistent draw of them, which the locale's `address` reads.
| Table | `geo.SE` | `geo.US` | Weight |
|-------|----------|----------|--------|
| `region` | län, by code or name | state, by USPS abbreviation or name; `code` is the FIPS code | population |
| `municipality` | kommun, by code or name | county, by FIPS code or name | population |
| `locality` | postort, by name | incorporated place of 25,000 people or more with a postal code of its own, by GEOID or name; Hawaii has none | tätort population, the kommun's where the postort names it, else 200; place population |
| `postal-code` | postnummer with street delivery, by code | ZCTA, by code | one; address ranges |
| `street` | gatunamn, the ten with most road segments per postort | street name, the ten with most address ranges per place | segments; address ranges |
`geo.SE.region[Skåne län].municipality` draws a kommun in Skåne,
`geo.SE.locality[Lund].street` a street in Lund, and
`geo.US.region[IL].locality[Springfield]` settles which Springfield. A region row
carries its `timezone`, the state's predominant zone, and a locality its `lat` and
`lon`. What ports across countries is the five table names, the `name` column,
selection by name, and the `address` record's columns `street`, `street-number`,
`postal-code` and `locality`; every other column is the country's own, `code` on a
Swedish region but `abbr` on a US one.
## Data format
Every value is a **node**, nestable without limit:
@@ -407,7 +435,9 @@ cell token, and refuses a TSV no category names, a key that is empty or repeats,
weight that is not a positive number, and a key or name holding `[`, `]`, `{`, `}`,
`"` or `|`, which a selector cannot spell; the rows are indexed on the first draw that
selects one. A `name` needs a `key`, since a name naming several rows is reported by
their keys, and a name spelling another row's key is refused, since the key would
their keys, or a `parent`, inside whose row a name names one row, so `first-name[Kim]`
is settled by the `sex` selected before it and a name repeating inside one parent row
is refused; a name spelling another row's key is refused, since the key would
select first and the name never. The table's options are its own — `rows`, `key`,
`name`, `weight` and `parent` — so a column may be named `name`, as one usually is.
@@ -505,21 +535,34 @@ stays reproducible.
| `{ulid()}` | sample | ULID, 26 Crockford base32 chars |
| `{nanoid(n)}` | sample | URL-safe Nano ID, `n` chars |
| `{iban(CC)}` | sample | length- and mod-97-valid IBAN for BE, DE, DK, ES, FI, NO or SE |
| `{date(from,to,'layout')}` | sample | a second between two `YYYY-MM-DD` days, both included, in a quoted Go layout: `'2006-01-02'`, `'January 2, 2006'`, `'060102'`, `'2006-01-02T15:04:05Z'` |
| `{time('layout')}` | sample | a second within a day: `'15:04'`, `'3:04 PM'` |
| `{seq()}`, `{seq(name)}` | counter | next integer from 1 in this generator; `name` selects an independent counter |
| `{calc(expr)}`, `{calc(expr,dp)}` | computation | an arithmetic expression over sibling fields ([Computation](#computation)) |
| `{lowercase(x)}`, `{uppercase(x)}`, `{ascii(x)}` | transform | a field's value rewritten ([Transforms](#transforms)) |
A derivation reads what is to its left, so place it after its payload; the
buffer is per expansion, so a nested template keeps fixed parts out of the sum. A
Swedish personnummer is a Luhn checksum over the nine digits before it:
Swedish personnummer is a Luhn checksum over the nine digits before it, six of
them a birthdate:
```json
{ "format": "{century}{core}", "century": ["19", "20"],
"core": { "format": "{digits(2)}{mmdd}-{digits(3)}{luhn()}", "mmdd": ["0115", "0704", "1218"] } }
{ "format": "{date(1930-01-01,2010-12-31,'060102')}-{birth}{luhn()}", "birth": ["238", "239"] }
```
Renders e.g. `19811218-9876`. `{seq()}` spans `Fake` calls and `repeat`, resets
with a new generator, and is the natural primary key for the SQL example above.
Renders e.g. `811218-2389`. A layout is Go's: the reference time `Mon Jan 2
15:04:05 MST 2006` spelled as the output should look, quoted, since a layout may
carry the comma that separates arguments, with English names. Every second
between the two days is reachable, so a layout with a clock draws the time too, and
`from` may equal `to`, which is that one day. The instant is UTC, so a zone in the
layout prints `UTC` or `Z`.
The quotes delimit a layout outside a selector only, so `[O'Fallon]` in an
argument stays a name. Rejected at `New`: a bound that is no calendar date, or not
before the other; an unquoted layout, naming the single-quoted one; a layout naming no field, which is
text, as is one day in a layout with no clock; for `date` a layout naming no date
field, naming `time`; and for `time` a layout naming a date field, naming `date`. `{seq()}` spans `Fake`
calls and `repeat`, resets with a new generator, and is the natural primary key for
the SQL example above.
### Computation
@@ -573,9 +616,9 @@ without naming `sv_SE`:
Renders e.g. `Hej, Pat Smith!`. A reference path into a category is held like a
[correlated](#correlated-fields) path, but for the whole render — one `Fake`, or one
record — rather than one format: `{.person.femalefirst} {.person.last}` name one
record — rather than one format: `{.person.first} {.person.last}` name one
person, as do the same two references in sibling fields or a nested template, and
`{lowercase(.person.femalefirst)}` reads that same draw. Each `repeat` iteration is
`{lowercase(.person.first)}` reads that same draw. Each `repeat` iteration is
a render of its own, in no group, so it draws anew, and a [draw group](#draw-group) holds a
draw apart. A bare reference names no field and makes its own picks each time —
`{/misc.uuid} {/misc.uuid}` is two draws — while the reference paths inside what it
@@ -594,8 +637,8 @@ its groups by name; the unnamed group spans them all.
```json
{ "format": "{payer} pays {payee}; signed {signature}",
"payer": { "format": "{/sv_SE.person.femalefirst} {/sv_SE.person.last}", "drawGroup": "payer" },
"payee": "{/sv_SE.person.femalefirst} {/sv_SE.person.last}",
"payer": { "format": "{/sv_SE.person.first} {/sv_SE.person.last}", "drawGroup": "payer" },
"payee": "{/sv_SE.person.first} {/sv_SE.person.last}",
"signature": { "format": "{/sv_SE.person.last}", "drawGroup": "payer" } }
```
@@ -697,6 +740,15 @@ datatype and nullability, and each table's key, name, weight and parent columns;
changes it or `data/` adds its `CHANGELOG.md` entry, which CI checks. A removed,
renamed or retyped line is a major.
## Audience
App developers writing tests and fixtures, in Go and at a shell:
- a **bulk fixture author**, thousands of rows into CSV or SQL
- a **Go test author**, filling a struct with `FakeStruct`
- a **hand fixture author**, one value at a shell
- a **validator-facing author**, who needs a value a real checker accepts
## Goals
1. **Valid by construction** — every value passes the check its real consumer
@@ -728,8 +780,7 @@ renamed or retyped line is a major.
key, or prefixing options, would tax every template to guard against a
misspelt option.
- **`{a|b}` stays beside nested choices.** `[[…], […]]` picks the same way, but
its arms are anonymous; `{femalefirst|malefirst}` keeps `person.femalefirst`
addressable.
its arms are anonymous; `{female|male}` keeps `person.female` addressable.
- **Flags follow getopt_long.** `--name value` and `--name=value` both work; a
short flag's value attaches or follows (`-s42`, `-s 42`) and short flags bundle
(`-hn 3`), as every shell user expects. A single-dash long flag is rejected
@@ -747,7 +798,9 @@ renamed or retyped line is a major.
library's own advice reachable: the error for an object holding only a format
names `"…"`, and that spelling has to work where it is printed. An argument or struct
tag of one reference alone, `{/users}`, is refused naming the path `users`: both
render the same text, and only the path names a record. `IsTemplate` exports the
render the same text, and only the path names a record. A folder-relative `{.name}`
or `{..name}` is refused naming `{/name}`, since an inline template sits in no
folder. `IsTemplate` exports the
rule, so the CLI, struct tags and any other caller read one.
- **An inline template skips the cycle fence.** `New` proves the loaded tree
acyclic, an inline node is a finite tree of its own, and nothing in the tree can
@@ -1040,6 +1093,51 @@ renamed or retyped line is a major.
bigger neighbour ship no address; counting the land outside every place too would
drop a quarter of the places, whose codes straddle unincorporated land, for a
postal city the USPS mostly names the same way.
- **`--list` stays a plain list of paths.** It is what a script reads, so every line
has to be a path that `Fake` takes; a marker for the tables a `[selector]` follows,
or a legend above them, would make the output something to parse before use.
`--help` names the selector spelling instead, and the Table section teaches it.
- **A layout is always quoted.** A layout may carry the comma that separates
arguments, `'January 2, 2006'`, and one spelling for every layout beats a rule
about which ones need the quotes, so the bare spelling is refused naming the
quoted one. The layout is Go's reference time because the library renders with
it and a Go caller already knows it; its names are English, and a locale's own
month and weekday names are data.
- **A title is a table under `sex`.** A prefix drawn apart would put `Mr` on a record
whose `sex` column says `female`, which is the disagreement the record exists to
prevent; the tables this set already has are what a title needs, so `en_US.title`
links to `sex` as `first-name` does. Swedish has no everyday sexed honorific, so
`sv_SE.title` is a table as well but carries no `parent`. Its weight column is
`share`, not the `count` a name table carries, because the values are a curated
proportion rather than bearers anyone counted.
- **A table owns the spelling of a selector on it.** A reference reaches a table by a
path that carries no selector — `sv_SE.person.first` reads `first-name` through
`sex` — so the walk that resolved a name cannot say where a reader would type one.
The table's own location can, which is why it keeps its path, and why an ambiguity
error names `sv_SE.sex[f].first-name[Kim]` rather than the table's own name.
- **No builtin reads the clock, so a date is bounded by days, never by an age.**
An `age(min,max)` would make a seeded fixture change with the day it runs on,
which is what a seed exists to prevent; a birthdate for someone 20 to 60 is
`date(1966-01-01,2006-12-31,…)`, re-pinned as any fixture is.
- **A name column without a key resolves inside its parent.** A given name both
sexes carry is a row under each, so `name` cannot be the key; the parent's row
tells the two apart, `sex[f].first-name[Kim]`, the ambiguity error spells each
row inside its parent, and a name repeating inside one parent row is refused at
load, since nothing could then select it.
- **The Swedish ids draw Skatteverket's test series.** A Luhn-valid personnummer
over a random birth number may be a living person's; 238 and 239 after any date
are blocked from assignment, so the shipped `personnummer` and
`samordningsnummer` use those. They sit in a `birth-number` table under `sex`
rather than as a column of it: the render's shared draw of the family is what
makes the number and the name agree on sex, and `sex` stays one shape across
locales instead of collecting every sex-keyed id fact. A samordningsnummer's
day, the birthday plus 60, is drawn from 61 to 88, valid in every month, rather
than computed from the date drawn.
- **The US given names come from a mirror of the SSA file.** ssa.gov refuses a
client outside the US, so `names-us.py` reads a GitHub copy that ends at 2020,
which a count over the births since 1930 barely feels; `--names` takes the
official zip. The SSA's placeholder rows are top-1000 entries that name nobody, so
the import drops them by name rather than by a rank a regeneration would move.
- **`List` advertises direct descents only.** `region.municipality.locality` is
listed, and `region.locality` resolves too but is not: the set of every descent
through a chain of five tables is every subsequence of it, and the direct chain is
@@ -1091,15 +1189,19 @@ A shipped table built from a source is rebuilt by its script under
[`data-import/`](data-import), one command per dataset, fetching the source named in
[`DATA-LICENSES.md`](DATA-LICENSES.md). Downloads are cached under
`data-import/cache/`, so delete it to fetch afresh; `geo-us.py` fetches two
TIGER/Line files per county it ships, a few hundred megabytes, and `geo-se.py` needs
TIGER/Line files per county it ships, a few hundred megabytes, `geo-se.py` needs
a Trafikverket API key, free at [data.trafikverket.se](https://data.trafikverket.se/),
in `TRAFIKVERKET_API_KEY` or a `--key-file`:
in `TRAFIKVERKET_API_KEY` or a `--key-file`, and the Census host behind `geo-us.py`
and `names-us.py` rejects a client for a while after a burst, so `--surnames` takes
a copy of the surname file:
```sh
docker compose run --rm --user "$(id -u):$(id -g)" data-import data-import/country.py
docker compose run --rm --user "$(id -u):$(id -g)" data-import data-import/currency.py
docker compose run --rm --user "$(id -u):$(id -g)" data-import data-import/geo-us.py
docker compose run --rm --user "$(id -u):$(id -g)" -e TRAFIKVERKET_API_KEY data-import data-import/geo-se.py
docker compose run --rm --user "$(id -u):$(id -g)" data-import data-import/names-se.py
docker compose run --rm --user "$(id -u):$(id -g)" data-import data-import/names-us.py
```
To release, head `CHANGELOG.md` with the version's section in place of `Unreleased`
@@ -1125,6 +1227,9 @@ family.go a family of linked tables: the rows a render pins, and the fence
reference.go reference sigils, and binding references across the tree
graph.go the render graph: edges, cycles, the repeat bound, tree walks
builtins.go the {name()} function registry and its implementations
layout.go date and time layouts: the instants one is proved against, and the two samples
checksum.go the check characters a derivation appends, and the IBAN they sit inside
transform.go the builtins that rewrite an operand's value, and the ASCII folding
calc.go the {calc()} arithmetic evaluator: parser, eval, validation
datatype.go column datatypes: DataType, where datatype and null may sit, a column's datatype
value.go the value proof: what a typed column or calc operand holds, checked at load
+1 -1
View File
@@ -30,7 +30,7 @@ func BenchmarkPerson(b *testing.B) { benchPath(b, "data", "sv_SE.person") }
func BenchmarkAddress(b *testing.B) { benchPath(b, "data", "sv_SE.address") }
func BenchmarkWord(b *testing.B) { benchPath(b, "data", "sv_SE.word") }
func BenchmarkCreditcard(b *testing.B) { benchPath(b, "data", "misc.creditcard") }
func BenchmarkSSN(b *testing.B) { benchPath(b, "data", "sv_SE.ssn") }
func BenchmarkPersonnummer(b *testing.B) { benchPath(b, "data", "sv_SE.personnummer") }
func BenchmarkUUIDv7(b *testing.B) { benchPath(b, "data", "misc.uuid") }
func tmpData(b *testing.B, name, body string) string {
+3 -193
View File
@@ -7,7 +7,6 @@ import (
"math"
"strconv"
"strings"
"unicode"
)
// maxLen caps sample output lengths (hex, nanoid, base64, digits, upper, lower)
@@ -60,6 +59,8 @@ var builtins = map[string]builtin{
cc := a[0]
return func(s *session, _ string, _ []string) string { return iban(s, cc) }
}},
"date": {arity: -1, check: dateArgs, prep: datePrep},
"time": {arity: -1, check: timeArg, prep: timePrep},
"calc": {arity: -1, check: checkCalc, prep: calcPrep, operands: calcOperands},
"lowercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToLower), operands: transformOperand},
"uppercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToUpper), operands: transformOperand},
@@ -88,13 +89,11 @@ func derive(f func(emitted string) string) func([]string) callFn {
return func(_ *session, emitted string, _ []string) string { return f(emitted) }
}
}
func sample(f func(rng) string) func([]string) callFn {
return func([]string) callFn {
return func(s *session, _ string, _ []string) string { return f(s) }
}
}
func chars(alphabet string) func([]string) callFn {
return func(a []string) callFn {
n := atoi(a[0])
@@ -113,96 +112,6 @@ func formatFloat(v float64, dp int) string {
return s
}
// transforms are the builtins that rewrite one operand's value; they nest, so
// {lowercase(ascii(x))} folds then lowers.
var transforms = map[string]func(string) string{
"ascii": asciiFold,
"lowercase": strings.ToLower,
"uppercase": strings.ToUpper,
}
// unwrapTransform peels nested transform calls off an operand arg, returning the
// field it finally names and the transforms to apply, innermost last.
func unwrapTransform(arg string) (leaf string, chain []func(string) string, err error) {
for {
name, args, isCall := funcCall(arg)
if !isCall {
return arg, chain, nil
}
fn, isTransform := transforms[name]
if !isTransform {
return "", nil, fmt.Errorf("%s(%s) is not a transform, so it cannot be an operand", name, strings.Join(args, ","))
}
if len(args) != 1 {
return "", nil, fmt.Errorf("%s takes 1 arg, got %d", name, len(args))
}
chain = append(chain, fn)
arg = args[0]
}
}
func transformArg(fields map[string]node, a []string) error {
leaf, _, err := unwrapTransform(a[0])
if err != nil {
return err
}
if isRef(leaf) {
_, _, err := refShape(leaf)
return err
}
return checkArm(leaf, fields, false)
}
func transformOperand(a []string) []string {
leaf, _, err := unwrapTransform(a[0])
if err != nil {
return nil
}
return []string{leaf}
}
func transformPrep(outer func(string) string) func([]string) callFn {
return func(a []string) callFn {
_, chain, err := unwrapTransform(a[0])
if err != nil {
panic(fmt.Sprintf("fejkdata: transform arg %q reached prep unvalidated: %v", a[0], err))
}
return func(_ *session, _ string, operands []string) string {
v := operands[0]
for i := len(chain) - 1; i >= 0; i-- {
v = chain[i](v)
}
return outer(v)
}
}
}
// asciiFolds maps the Latin letters with diacritics or ligatures to ASCII.
var asciiFolds = map[rune]string{
'À': "A", 'Á': "A", 'Â': "A", 'Ã': "A", 'Ä': "A", 'Å': "A", 'Æ': "AE", 'Ç': "C",
'È': "E", 'É': "E", 'Ê': "E", 'Ë': "E", 'Ì': "I", 'Í': "I", 'Î': "I", 'Ï': "I",
'Ð': "D", 'Ñ': "N", 'Ò': "O", 'Ó': "O", 'Ô': "O", 'Õ': "O", 'Ö': "O", 'Ø': "O",
'Ù': "U", 'Ú': "U", 'Û': "U", 'Ü': "U", 'Ý': "Y", 'Þ': "Th", 'ß': "ss", 'Œ': "OE",
'à': "a", 'á': "a", 'â': "a", 'ã': "a", 'ä': "a", 'å': "a", 'æ': "ae", 'ç': "c",
'è': "e", 'é': "e", 'ê': "e", 'ë': "e", 'ì': "i", 'í': "i", 'î': "i", 'ï': "i",
'ð': "d", 'ñ': "n", 'ò': "o", 'ó': "o", 'ô': "o", 'õ': "o", 'ö': "o", 'ø': "o",
'ù': "u", 'ú': "u", 'û': "u", 'ü': "u", 'ý': "y", 'þ': "th", 'ÿ': "y", 'œ': "oe",
}
// asciiFold rewrites s to ASCII: folded Latin letters stay, any other non-ASCII
// rune is dropped.
func asciiFold(s string) string {
var b strings.Builder
for _, r := range s {
if r <= unicode.MaxASCII {
b.WriteRune(r)
} else {
b.WriteString(asciiFolds[r])
}
}
return b.String()
}
// atoi parses an arg a builtin's check already validated. It panics rather than
// returning zero, so a check that stops covering its own args is a stack trace and
// not a silently wrong length, range or decimal count.
@@ -222,7 +131,6 @@ func atof(s string) float64 {
}
return f
}
func randBytes(r rng, n int) []byte {
b := make([]byte, n)
for i := range b {
@@ -230,7 +138,6 @@ func randBytes(r rng, n int) []byte {
}
return b
}
func randChars(r rng, n int, alphabet string) string {
b := make([]byte, n)
for i := range b {
@@ -253,7 +160,6 @@ func plainInt(s string) (int, error) {
}
return n, nil
}
func posIntArg(_ map[string]node, a []string) error {
n, err := plainInt(a[0])
if errors.Is(err, strconv.ErrRange) {
@@ -270,7 +176,6 @@ func posIntArg(_ map[string]node, a []string) error {
}
return nil
}
func intRangeArgs(_ map[string]node, a []string) error {
lo, err := plainInt(a[0])
if err != nil {
@@ -291,7 +196,6 @@ func intRangeArgs(_ map[string]node, a []string) error {
}
return nil
}
func floatArgs(_ map[string]node, a []string) error {
lo, e1 := strconv.ParseFloat(a[0], 64)
hi, e2 := strconv.ParseFloat(a[1], 64)
@@ -319,10 +223,9 @@ func floatArgs(_ map[string]node, a []string) error {
}
return nil
}
func seqArg(_ map[string]node, a []string) error {
if len(a) > 1 {
return fmt.Errorf("seq takes at most one name, got %d args", len(a))
return fmt.Errorf("seq takes at most one name, got %d", len(a))
}
if len(a) == 1 && a[0] == "" {
return fmt.Errorf("seq name must not be empty")
@@ -373,96 +276,3 @@ func ulid(r rng) string {
}
return string(out)
}
// luhnCheck returns the Luhn check digit (0-9) over the digits of s; non-digit
// runes are skipped. Doubling runs from the rightmost digit, so the result is
// correct whatever the payload length.
func luhnCheck(s string) int {
sum, double := 0, true
for i := len(s) - 1; i >= 0; i-- {
c := s[i]
if c < '0' || c > '9' {
continue
}
d := int(c - '0')
if double {
if d *= 2; d > 9 {
d -= 9
}
}
double = !double
sum += d
}
return (10 - sum%10) % 10
}
// mod11Check returns the weighted mod-11 check character over the digits of s
// (weights 2..7 cycling from the right). A would-be value of 10 emits 'X', as in
// ISBN-10 / ISO 7064; non-digits are skipped.
func mod11Check(s string) string {
sum, w := 0, 2
for i := len(s) - 1; i >= 0; i-- {
c := s[i]
if c < '0' || c > '9' {
continue
}
sum += int(c-'0') * w
if w++; w > 7 {
w = 2
}
}
if chk := (11 - sum%11) % 11; chk != 10 {
return string(rune('0' + chk))
}
return "X"
}
// eanCheck returns the EAN-13 / UPC-A / ISBN-13 / GTIN check digit over the
// digits of s: weights 3 and 1 alternating from the rightmost digit, mod 10.
func eanCheck(s string) string {
sum, w := 0, 3
for i := len(s) - 1; i >= 0; i-- {
c := s[i]
if c < '0' || c > '9' {
continue
}
sum += int(c-'0') * w
w = 4 - w // 3 <-> 1
}
return string(rune('0' + (10-sum%10)%10))
}
// ibanLen maps a supported country code to the full IBAN length. The check digits
// sit between the country code and the BBAN, so — unlike luhn/ean — iban can't be
// a left-to-right derivation; it generates the whole value instead.
var ibanLen = map[string]int{"BE": 16, "DE": 22, "DK": 18, "ES": 24, "FI": 18, "NO": 15, "SE": 24}
func ibanArg(_ map[string]node, a []string) error {
if _, ok := ibanLen[a[0]]; !ok {
return fmt.Errorf("iban(%q): unsupported country code", a[0])
}
return nil
}
// iban generates a structurally valid IBAN for cc: a numeric BBAN of the right
// length, then mod-97 check digits. Real bank/branch structure isn't modelled —
// the result passes length and checksum validation, which is what fake data needs.
func iban(r rng, cc string) string {
bban := make([]byte, ibanLen[cc]-4)
for i := range bban {
bban[i] = byte('0' + r.IntN(10))
}
rem := 0
feed := func(d int) { rem = (rem*10 + d) % 97 }
for _, c := range bban {
feed(int(c - '0'))
}
for i := 0; i < len(cc); i++ { // letters A-Z -> 10..35, fed as two digits
v := int(cc[i]-'A') + 10
feed(v / 10)
feed(v % 10)
}
feed(0)
feed(0)
return fmt.Sprintf("%s%02d%s", cc, 98-rem, bban)
}
+135
View File
@@ -4,7 +4,9 @@ import (
"encoding/base64"
"regexp"
"strconv"
"strings"
"testing"
"time"
)
func TestBuiltinIDGenerators(t *testing.T) {
@@ -41,6 +43,7 @@ func TestBuiltinSamplesReproducible(t *testing.T) {
`"{uuid()}"`, `"{ulid()}"`,
`"{nanoid(12)}"`, `"{int(1,1000000)}"`,
`"{float(0,1,6)}"`, `"{base64(12)}"`, `"{iban(SE)}"`,
`"{date(2000-01-01,2020-12-31,'2006-01-02 15:04:05')}"`, `"{time('15:04:05')}"`,
} {
if a, b := mustRender(t, engine(7), tmpl), mustRender(t, engine(7), tmpl); a != b {
t.Fatalf("%s not reproducible: %q != %q", tmpl, a, b)
@@ -233,3 +236,135 @@ func TestClassBuiltinArgs(t *testing.T) {
}
}
}
// TestBuiltinDateAndTime pins the two clock-free samples: date draws a second in
// [from 00:00:00, to 23:59:59], both days reachable, and renders it in the quoted Go
// layout, commas and English names included; time draws a second within one day.
func TestBuiltinDateAndTime(t *testing.T) {
f := engine(1)
seen := map[string]bool{}
for i := 0; i < 500; i++ {
got := mustRender(t, f, `"{date(1990-01-01,1990-12-31,'2006-01-02')}"`)
if d, err := time.Parse("2006-01-02", got); err != nil || d.Year() != 1990 {
t.Fatalf("date = %q, want a 1990 calendar date (err %v)", got, err)
}
seen[got] = true
}
if len(seen) < 200 {
t.Fatalf("date drew %d distinct days of 365 in 500, want a uniform spread", len(seen))
}
lo, hi := false, false
for i := 0; i < 200; i++ {
got := mustRender(t, f, `"{date(2020-02-28,2020-02-29,'2006-01-02')}"`)
if got != "2020-02-28" && got != "2020-02-29" {
t.Fatalf("date(2020-02-28,2020-02-29) = %q, out of range", got)
}
lo, hi = lo || got == "2020-02-28", hi || got == "2020-02-29"
}
if !lo || !hi {
t.Fatalf("date never hit a bound: lo=%v hi=%v (bounds must be inclusive)", lo, hi)
}
if got := mustRender(t, f, `"{date(2020-07-04,2020-07-05,'January 2, 2006')}"`); got != "July 4, 2020" && got != "July 5, 2020" {
t.Fatalf("date with a comma in its layout = %q", got)
}
rfc := regexp.MustCompile(`^2021-\d\d-\d\dT\d\d:\d\d:\d\dZ$`)
clock := regexp.MustCompile(`^([01]\d|2[0-3]):[0-5]\d$`)
ampm := regexp.MustCompile(`^(1[0-2]|[1-9]):[0-5]\d (AM|PM)$`)
seconds := map[string]bool{}
for i := 0; i < 300; i++ {
got := mustRender(t, f, `"{date(2021-01-01,2021-12-31,'2006-01-02T15:04:05Z07:00')}"`)
if !rfc.MatchString(got) {
t.Fatalf("date in an RFC 3339 layout = %q, want %s", got, rfc)
}
seconds[got[17:19]] = true
if got := mustRender(t, f, `"{time('15:04')}"`); !clock.MatchString(got) {
t.Fatalf("time('15:04') = %q, want %s", got, clock)
}
if got := mustRender(t, f, `"{time('3:04 PM')}"`); !ampm.MatchString(got) {
t.Fatalf("time('3:04 PM') = %q, want %s", got, ampm)
}
got = mustRender(t, f, `"{date(1950-01-01,2000-12-31,'060102')}-238{luhn()}"`)
if d := digitsOnly(got); len(d) != 10 || !luhnValid(d) {
t.Fatalf("a personnummer over date() = %q, want ten Luhn-valid digits", got)
}
}
if len(seconds) < 30 {
t.Fatalf("date drew %d distinct seconds in 300, want the whole day, not midnight", len(seconds))
}
}
// TestBuiltinDateArgs pins the New-time checks: bounds are calendar dates in order,
// the layout is quoted, names a field, and for time names no date field.
func TestBuiltinDateArgs(t *testing.T) {
for tmpl, want := range map[string]string{
`"{date(1990-13-01,1990-12-31,'2006-01-02')}"`: "1990-13-01",
`"{date(1990-12-31,1990-01-01,'2006-01-02')}"`: "is after",
`"{date(1990-01-01,1990-01-01,'2006-01-02')}"`: "write it as text",
`"{date(1990-01-01,1990-12-31,2006-01-02)}"`: "'2006-01-02'",
`"{date(1990-01-01,1990-12-31,'January 2, 2006)}"`: "'",
`"{date(1990-01-01,1990-12-31,\"January 2, 2006\")}"`: "quoted: 'January 2, 2006'",
`"{date(1990-01-01,1990-12-31,'x')}"`: "text",
`"{date(1990-01-01,1990-12-31,'')}"`: "text",
`"{date(1990-01-01,1990-12-31)}"`: "3 arguments",
`"{date(1990-01-01,1990-12-31,January 2, 2006)}"`: "'January 2, 2006'",
`"{time(3:04 PM, Mon)}"`: "'3:04 PM, Mon'",
`"{time(15:04)}"`: "'15:04'",
`"{time('2006-01-02 15:04')}"`: "date(",
`"{time('x')}"`: "text",
`"{date(1990-01-01,1990-12-31,'15:04')}"`: "time('15:04')",
} {
_, err := compile(parse(t, tmpl))
if err == nil || !strings.Contains(err.Error(), want) {
t.Errorf("compile(%s) = %v, want an error mentioning %q", tmpl, err, want)
}
}
}
// TestBuiltinLayoutErrorsNameARunnableSpelling pins the two layout errors a user
// can follow: a double-quoted layout is named single-quoted, without its own
// quotes carried into the suggestion, whether it split on a comma or not.
func TestBuiltinLayoutErrorsNameARunnableSpelling(t *testing.T) {
_, err := compile(parse(t, `"{date(1990-01-01,1990-12-31,\"2006-01-02\")}"`))
if err == nil || !strings.Contains(err.Error(), "write '2006-01-02'") {
t.Fatalf("a double-quoted layout = %v, want it named single-quoted", err)
}
if strings.Contains(err.Error(), `'"`) {
t.Errorf("%v names a layout that renders its own quotes", err)
}
if !strings.Contains(err.Error(), "is double-quoted") {
t.Errorf("%v does not say which quotes were wrong", err)
}
_, err = compile(parse(t, `"{time(15:04)}"`))
if err == nil || !strings.Contains(err.Error(), "write '15:04'") {
t.Fatalf("a bare layout = %v, want it named quoted", err)
}
// The comma hint belongs to a layout that split, not to a call given extra args.
_, err = compile(parse(t, `"{time(0,12,'15:04')}"`))
if err == nil || !strings.Contains(err.Error(), "takes 1 argument, got 3") {
t.Fatalf("time with three args = %v, want the count named in the singular", err)
}
if strings.Contains(err.Error(), "holding a comma") {
t.Errorf("%v offers the comma hint though the layout is already quoted", err)
}
}
// TestBuiltinDateSpansOneDay pins that from == to is a day: with a clock layout it
// draws every second of it, and without one it could only emit one value.
func TestBuiltinDateSpansOneDay(t *testing.T) {
f := engine(1)
seen := map[string]bool{}
for i := 0; i < 500; i++ {
got := mustRender(t, f, `"{date(2026-01-01,2026-01-01,'2006-01-02 15:04:05')}"`)
if !strings.HasPrefix(got, "2026-01-01 ") {
t.Fatalf("date over one day = %q, out of range", got)
}
seen[got] = true
}
if len(seen) < 400 {
t.Fatalf("date over one day drew %d distinct seconds in 500, want the whole day", len(seen))
}
_, err := compile(parse(t, `"{date(2026-01-01,2026-01-01,'2006-01-02')}"`))
if err == nil || !strings.Contains(err.Error(), "write it as text") {
t.Fatalf("one day in a date-only layout = %v, want it named a constant", err)
}
}
+96
View File
@@ -0,0 +1,96 @@
package fejkdata
import "fmt"
// luhnCheck returns the Luhn check digit (0-9) over the digits of s; non-digit
// runes are skipped. Doubling runs from the rightmost digit, so the result is
// correct whatever the payload length.
func luhnCheck(s string) int {
sum, double := 0, true
for i := len(s) - 1; i >= 0; i-- {
c := s[i]
if c < '0' || c > '9' {
continue
}
d := int(c - '0')
if double {
if d *= 2; d > 9 {
d -= 9
}
}
double = !double
sum += d
}
return (10 - sum%10) % 10
}
// mod11Check returns the weighted mod-11 check character over the digits of s
// (weights 2..7 cycling from the right). A would-be value of 10 emits 'X', as in
// ISBN-10 / ISO 7064; non-digits are skipped.
func mod11Check(s string) string {
sum, w := 0, 2
for i := len(s) - 1; i >= 0; i-- {
c := s[i]
if c < '0' || c > '9' {
continue
}
sum += int(c-'0') * w
if w++; w > 7 {
w = 2
}
}
if chk := (11 - sum%11) % 11; chk != 10 {
return string(rune('0' + chk))
}
return "X"
}
// eanCheck returns the EAN-13 / UPC-A / ISBN-13 / GTIN check digit over the
// digits of s: weights 3 and 1 alternating from the rightmost digit, mod 10.
func eanCheck(s string) string {
sum, w := 0, 3
for i := len(s) - 1; i >= 0; i-- {
c := s[i]
if c < '0' || c > '9' {
continue
}
sum += int(c-'0') * w
w = 4 - w // 3 <-> 1
}
return string(rune('0' + (10-sum%10)%10))
}
// ibanLen maps a supported country code to the full IBAN length. The check digits
// sit between the country code and the BBAN, so — unlike luhn/ean — iban can't be
// a left-to-right derivation; it generates the whole value instead.
var ibanLen = map[string]int{"BE": 16, "DE": 22, "DK": 18, "ES": 24, "FI": 18, "NO": 15, "SE": 24}
func ibanArg(_ map[string]node, a []string) error {
if _, ok := ibanLen[a[0]]; !ok {
return fmt.Errorf("iban(%q): unsupported country code", a[0])
}
return nil
}
// iban generates a structurally valid IBAN for cc: a numeric BBAN of the right
// length, then mod-97 check digits. Real bank/branch structure isn't modelled —
// the result passes length and checksum validation, which is what fake data needs.
func iban(r rng, cc string) string {
bban := make([]byte, ibanLen[cc]-4)
for i := range bban {
bban[i] = byte('0' + r.IntN(10))
}
rem := 0
feed := func(d int) { rem = (rem*10 + d) % 97 }
for _, c := range bban {
feed(int(c - '0'))
}
for i := 0; i < len(cc); i++ { // letters A-Z -> 10..35, fed as two digits
v := int(cc[i]-'A') + 10
feed(v / 10)
feed(v % 10)
}
feed(0)
feed(0)
return fmt.Sprintf("%s%02d%s", cc, 98-rem, bban)
}
+18 -1
View File
@@ -28,6 +28,9 @@ const usage = `Usage: fejkdata [flags] <path|template>
<template> a format string or JSON value to render inline, e.g.
'name: {/sv_SE.person.last}' or '{"format":"{x}","x":["bosse","lina"]}'
A layout inside a template is single-quoted, so quote the whole argument with " to
keep it: "{date(1990-01-01,2010-12-31,'2006-01-02')}".
An argument containing a { token, or a JSON object, array or string, is a
template; any other argument is a path (a path never contains a brace or a quote,
and a bracket only as a [selector] after a table's name). Templates reach the
@@ -279,7 +282,7 @@ func (in invocation) check() (argKind, error) {
return argPath, nil
}
if len(in.paths) != 1 {
return argPath, fmt.Errorf("expected one path or template, got %d", len(in.paths))
return argPath, fmt.Errorf("expected one path or template, got %d%s", len(in.paths), templateSplit(in.paths))
}
return classify(in.paths[0])
}
@@ -475,11 +478,25 @@ func run(args []string, stdout, stderr io.Writer) int {
return misuse(stderr, te.error)
}
fmt.Fprintln(stderr, err)
if errors.Is(err, fejkdata.ErrNoColumns) {
fmt.Fprintln(stderr, "wrap that JSON in single quotes, which keep a shell from expanding its braces")
}
return 1
}
return 0
}
// templateSplit names the cure for a template a shell split on its spaces, which is
// what several arguments carrying a token mean.
func templateSplit(paths []string) string {
for _, p := range paths {
if strings.Contains(p, "{") {
return `; a template's spaces split the argument, so wrap the whole argument in "…", or in '…' where it carries double quotes`
}
}
return ""
}
func misuse(stderr io.Writer, err error) int {
// A library error already names the program, so the prefix is not doubled.
fmt.Fprintf(stderr, "fejkdata: %s\ntry 'fejkdata --help'\n", strings.TrimPrefix(err.Error(), "fejkdata: "))
+26
View File
@@ -223,6 +223,14 @@ func TestRunMisuse(t *testing.T) {
t.Errorf("run(%v) stdout = %q, want nothing", args, out)
}
}
// A template a shell split names both quotings, since which one cures it turns
// on whether the argument carries a double-quoted JSON.
_, _, errb := runOut("{date(1990-01-01,2010-12-31,", "January", "2,", "2006)}")
for _, want := range []string{`wrap the whole argument in "…"`, `or in '…'`} {
if !strings.Contains(errb, want) {
t.Errorf("a split template stderr = %q, want it to name %s", errb, want)
}
}
}
func TestRunMultipleDataPaths(t *testing.T) {
@@ -613,3 +621,21 @@ func TestRunNumbersFollowTheShell(t *testing.T) {
t.Errorf("run(--seed 007 --repeat +3) = %d, %q, stderr %q; want the same as --seed 7 --repeat 3 %q", code, got, errb, want)
}
}
// TestRunRecordOfAFieldlessPath pins the way out of asking a category of one value
// for columns: the record to write, and the quoting a shell needs to take it. The
// argument is well-formed and only the data cannot serve it, so the code is 1.
func TestRunRecordOfAFieldlessPath(t *testing.T) {
code, out, errb := runOut("--format", "csv", "sv_SE.personnummer")
if code != 1 {
t.Fatalf("run(--format csv sv_SE.personnummer) = %d, want 1", code)
}
for _, want := range []string{`{"format":"","personnummer":"{/sv_SE.personnummer}"}`, "single quotes"} {
if !strings.Contains(errb, want) {
t.Errorf("stderr = %q, want it to name %s", errb, want)
}
}
if out != "" {
t.Errorf("stdout = %q, want nothing", out)
}
}
+2 -1
View File
@@ -9,6 +9,7 @@ import io
import re
from pathlib import Path
import source
import tsv
SOURCE = "https://raw.githubusercontent.com/datasets/country-codes/main/data/country-codes.csv"
@@ -58,7 +59,7 @@ def main():
p.add_argument("--source", default=SOURCE)
p.add_argument("--out", default=str(OUT))
a = p.parse_args()
table = rows(tsv.fetch(a.source, a.cache, "country-codes.csv").decode("utf-8"))
table = rows(source.fetch(a.source, a.cache, "country-codes.csv").decode("utf-8"))
tsv.write(a.out, COLUMNS, sorted(table, key=lambda r: r["alpha2"]))
+3 -2
View File
@@ -11,6 +11,7 @@ import io
import xml.etree.ElementTree as ET
from pathlib import Path
import source
import tsv
SOURCE = "https://raw.githubusercontent.com/datasets/currency-codes/main/data/codes-all.csv"
@@ -58,8 +59,8 @@ def main():
p.add_argument("--symbols", nargs="+", default=SYMBOLS)
p.add_argument("--out", default=str(OUT))
a = p.parse_args()
symbol = symbols(tsv.fetch(s, a.cache, Path(s).name).decode("utf-8") for s in a.symbols)
table = rows(tsv.fetch(a.source, a.cache, "codes-all.csv").decode("utf-8"), symbol)
symbol = symbols(source.fetch(s, a.cache, Path(s).name).decode("utf-8") for s in a.symbols)
table = rows(source.fetch(a.source, a.cache, "codes-all.csv").decode("utf-8"), symbol)
tsv.write(a.out, COLUMNS, sorted(table, key=lambda r: r["code"]))
+6 -18
View File
@@ -13,11 +13,12 @@ import os
import re
import sys
import urllib.request
import xml.etree.ElementTree as ET
import zipfile
from pathlib import Path
import source
import tsv
import xlsx
CODES = "https://www.scb.se/contentassets/7a89e48960f741e08918e489ea36354a/kommunlankod-2026.xlsx"
POPULATION = "https://api.scb.se/OV0104/v1/doris/sv/ssd/START/BE/BE0101/BE0101A/BefolkningNy"
@@ -38,24 +39,11 @@ CACHE = Path(__file__).resolve().parent / "cache"
TIMEZONE = "Europe/Stockholm"
ONE_POSITION = {"Stockholm", "Göteborg", "Malmö"}
UNMATCHED_POPULATION = 200
XLSX_NS = {"m": "http://schemas.openxmlformats.org/spreadsheetml/2006/main"}
def xlsx_rows(data):
z = zipfile.ZipFile(io.BytesIO(data))
strings = ["".join(t.text or "" for t in si.iter("{%s}t" % XLSX_NS["m"])) for si in ET.fromstring(z.read("xl/sharedStrings.xml")).findall("m:si", XLSX_NS)]
sheet = ET.fromstring(z.read("xl/worksheets/sheet1.xml"))
for row in sheet.findall(".//m:row", XLSX_NS):
cells = []
for c in row.findall("m:c", XLSX_NS):
v = c.find("m:v", XLSX_NS)
cells.append("" if v is None else strings[int(v.text)] if c.get("t") == "s" else v.text)
yield cells
def scb_codes(cache):
regions, municipalities = {}, {}
for cells in xlsx_rows(tsv.fetch(CODES, cache, "kommunlankod.xlsx", magic=b"PK")):
for cells in xlsx.rows(source.fetch(CODES, cache, "kommunlankod.xlsx", magic=b"PK")):
if len(cells) < 2 or not re.fullmatch(r"\d{2}|\d{4}", cells[0]):
continue
(regions if len(cells[0]) == 2 else municipalities)[cells[0]] = cells[1].strip()
@@ -64,12 +52,12 @@ def scb_codes(cache):
def scb_population(cache):
body = json.dumps(POPULATION_QUERY).encode()
data = tsv.fetch(POPULATION, cache, "befolkning.json", data=body, headers={"Content-Type": "application/json"})
data = source.fetch(POPULATION, cache, "befolkning.json", data=body, headers={"Content-Type": "application/json"})
return {row["key"][0]: row["values"][0] for row in json.loads(data.decode("utf-8-sig"))["data"]}
def scb_tatorter(cache):
text = tsv.fetch(TATORTER, cache, "tatorter.csv").decode("utf-8")
text = source.fetch(TATORTER, cache, "tatorter.csv").decode("utf-8")
by_name = collections.defaultdict(list)
for r in csv.DictReader(io.StringIO(text)):
by_name[r["tatort"]].append((r["kommun"], int(r["bef"])))
@@ -77,7 +65,7 @@ def scb_tatorter(cache):
def geonames(cache):
z = zipfile.ZipFile(io.BytesIO(tsv.fetch(POSTAL_CODES, cache, "SE.zip", magic=b"PK")))
z = zipfile.ZipFile(io.BytesIO(source.fetch(POSTAL_CODES, cache, "SE.zip", magic=b"PK")))
rows = []
for line in z.read("SE.txt").decode("utf-8").splitlines():
f = line.split("\t")
+5 -4
View File
@@ -14,6 +14,7 @@ import sys
import zipfile
from pathlib import Path
import source
import tsv
GAZETTEER = "https://www2.census.gov/geo/docs/maps-data/data/gazetteer/2026_Gazetteer/2026_Gaz_{}_national.zip"
@@ -56,14 +57,14 @@ def text(data):
def gazetteer(cache, kind):
z = zipfile.ZipFile(io.BytesIO(tsv.fetch(GAZETTEER.format(kind), cache, f"gaz_{kind}.zip", magic=b"PK")))
z = zipfile.ZipFile(io.BytesIO(source.fetch(GAZETTEER.format(kind), cache, f"gaz_{kind}.zip", magic=b"PK")))
rows = text(z.read(z.namelist()[0])).splitlines()
header = [h.strip() for h in rows[0].split("|")]
return [dict(zip(header, (c.strip() for c in row.split("|")))) for row in rows[1:]]
def csv_rows(cache, url, name):
return list(csv.DictReader(io.StringIO(text(tsv.fetch(url, cache, name)))))
return list(csv.DictReader(io.StringIO(text(source.fetch(url, cache, name)))))
def dbf_rows(data, wanted):
@@ -89,7 +90,7 @@ def dbf_rows(data, wanted):
def tiger_zip(cache, kind, county):
return tsv.fetch(TIGER.format(kind.upper(), county, kind), cache, f"tl_{county}_{kind}.zip", magic=b"PK")
return source.fetch(TIGER.format(kind.upper(), county, kind), cache, f"tl_{county}_{kind}.zip", magic=b"PK")
def tiger(cache, kind, county, wanted):
@@ -130,7 +131,7 @@ def localities(cache, min_population, counties):
def postal_codes(cache, localities):
"""Each ZCTA whose largest part inside an incorporated place lies in a shipped place."""
parts = {}
for r in csv.DictReader(io.StringIO(text(tsv.fetch(ZCTA_PLACE, cache, "zcta-place.txt"))), delimiter="|"):
for r in csv.DictReader(io.StringIO(text(source.fetch(ZCTA_PLACE, cache, "zcta-place.txt"))), delimiter="|"):
if r["GEOID_ZCTA5_20"] and r["GEOID_PLACE_20"] and not r["NAMELSAD_PLACE_20"].endswith(" CDP"):
parts.setdefault(r["GEOID_ZCTA5_20"], []).append((int(r["AREALAND_PART"]), r["GEOID_PLACE_20"]))
largest = {zcta: max(p)[1] for zcta, p in parts.items()}
+54
View File
@@ -0,0 +1,54 @@
#!/usr/bin/env python3
"""Rebuild data/sv_SE/first-name.tsv and last-name.tsv from SCB's whole-population name counts (CC0, "Källa: SCB").
data-import/names-se.py [--source URL_OR_FILE] [--first N] [--last N] [--cache DIR] [--out DIR]
A tilltalsnamn goes under each sex that carries it, so a name both sexes carry has two rows.
"""
import argparse
import re
from pathlib import Path
import source
import tsv
import xlsx
SOURCE = "https://www.scb.se/contentassets/9fe7dbb460994c72b835163dbc491ef9/namn-med-minst-tva-barare-31-december-2022.xlsx"
OUT = Path(__file__).resolve().parent.parent / "data" / "sv_SE"
CACHE = Path(__file__).resolve().parent / "cache"
SHEETS = {"f": "Tilltalsnamn kvinnor", "m": "Tilltalsnamn män"}
SURNAMES = "Efternamn"
NAME = re.compile(r"^[^\W\d_]{2,}([ '-][^\W\d_]{2,})*$")
PARTICLES = {"af", "av", "da", "de", "del", "den", "der", "di", "dos", "du", "la", "le", "van", "von"}
def cased(name):
"""SCB's uppercase name as it is written: each part capitalised, a particle before another part lowercased."""
parts = name.lower().split(" ")
return " ".join(w if w in PARTICLES and len(parts) > 1 else w.title() for w in parts)
def counted(rows, top):
"""The top names of a sheet by bearers, initials and single letters dropped."""
names = [(cells[0], int(cells[1])) for cells in rows if len(cells) >= 2 and cells[1].isdigit() and NAME.match(cells[0])]
return sorted(names, key=lambda n: (-n[1], n[0]))[:top]
def main():
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
p.add_argument("--cache", default=str(CACHE))
p.add_argument("--first", type=int, default=2000, help="names per sex")
p.add_argument("--last", type=int, default=5000)
p.add_argument("--out", default=str(OUT))
p.add_argument("--source", default=SOURCE)
a = p.parse_args()
data = source.fetch(a.source, a.cache, "scb-namn-2022.xlsx", magic=b"PK")
first = [{"name": cased(n), "sex": sex, "count": c} for sex, sheet in SHEETS.items() for n, c in counted(xlsx.rows(data, sheet), a.first)]
last = [{"name": cased(n), "count": c} for n, c in counted(xlsx.rows(data, SURNAMES), a.last)]
out = Path(a.out)
tsv.write(out / "first-name.tsv", ["name", "sex", "count"], sorted(first, key=lambda r: (r["name"], r["sex"])))
tsv.write(out / "last-name.tsv", ["name", "count"], sorted(last, key=lambda r: r["name"]))
if __name__ == "__main__":
main()
+79
View File
@@ -0,0 +1,79 @@
#!/usr/bin/env python3
"""Rebuild data/en_US/first-name.tsv and last-name.tsv from SSA baby names and the Census 2010 surnames (public domain).
data-import/names-us.py [--names URL_OR_FILE] [--surnames URL_OR_FILE] [--from-year YEAR] [--first N] [--last N] [--cache DIR] [--out DIR]
A given name is counted over the births of --from-year and later, under each sex it was given to, so a
name both sexes carry has two rows. --names takes SSA's names.zip or a merged name,sex,count,year file;
--surnames takes the Census names.zip or its Names_2010Census.csv.
"""
import argparse
import collections
import csv
import io
import re
import zipfile
from pathlib import Path
import source
import tsv
NAMES = "https://raw.githubusercontent.com/hackerb9/ssa-baby-names/master/alldata.txt"
SURNAMES = "https://www2.census.gov/topics/genealogy/2010surnames/names.zip"
OUT = Path(__file__).resolve().parent.parent / "data" / "en_US"
CACHE = Path(__file__).resolve().parent / "cache"
MC = re.compile(r"^Mc([a-z])")
PLACEHOLDERS = {"Baby", "Female", "Infant", "Male", "Notnamed", "Unknown", "Unnamed"}
def csv_or_zip_rows(data, member):
"""The rows of a CSV, or of every member of a zip named like member, name,sex,count[,year]."""
if data.startswith(b"PK"):
z = zipfile.ZipFile(io.BytesIO(data))
for n in sorted(z.namelist()):
if re.fullmatch(member, n):
year = re.sub(r"\D", "", n)
for r in csv.reader(io.StringIO(z.read(n).decode("utf-8-sig"))):
yield r + [year] if year else r
return
yield from csv.reader(io.StringIO(data.decode("utf-8-sig")))
def given(data, from_year):
counts = collections.Counter()
for r in csv_or_zip_rows(data, r"yob\d{4}\.txt"):
if len(r) >= 4 and r[3].isdigit() and int(r[3]) >= from_year and r[0] not in PLACEHOLDERS:
counts[(r[0], r[1].lower())] += int(r[2])
return counts
def surname(name):
"""A Census uppercase surname as it is written: capitalised, and Mc before a capital."""
return MC.sub(lambda m: "Mc" + m.group(1).upper(), name.title())
def main():
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
p.add_argument("--cache", default=str(CACHE))
p.add_argument("--first", type=int, default=2000, help="names per sex")
p.add_argument("--from-year", type=int, default=1930)
p.add_argument("--last", type=int, default=5000)
p.add_argument("--names", default=NAMES)
p.add_argument("--out", default=str(OUT))
p.add_argument("--surnames", default=SURNAMES)
a = p.parse_args()
counts = given(source.fetch(a.names, a.cache, "ssa-names.txt"), a.from_year)
first = []
for sex in ("f", "m"):
top = sorted(((n, c) for (n, s), c in counts.items() if s == sex), key=lambda n: (-n[1], n[0]))[:a.first]
first += [{"name": n, "sex": sex, "count": c} for n, c in top]
rows = csv_or_zip_rows(source.fetch(a.surnames, a.cache, "census-surnames-2010.zip"), r"Names_2010Census\.csv")
last = [{"name": surname(r[0]), "count": int(r[2])} for r in rows if len(r) >= 3 and r[2].isdigit() and r[0].isalpha()]
last = sorted(last, key=lambda r: (-r["count"], r["name"]))[:a.last]
out = Path(a.out)
tsv.write(out / "first-name.tsv", ["name", "sex", "count"], sorted(first, key=lambda r: (r["name"], r["sex"])))
tsv.write(out / "last-name.tsv", ["name", "count"], sorted(last, key=lambda r: r["name"]))
if __name__ == "__main__":
main()
+29
View File
@@ -0,0 +1,29 @@
"""A source fetched once into the cache."""
import re
import sys
import time
import urllib.request
from pathlib import Path
def fetch(source, cache, name, magic=b"", data=None, headers=None):
"""The bytes of a URL, downloaded into cache/name once, or of a local file."""
if not re.match(r"^https?://", source):
return Path(source).read_bytes()
path = Path(cache) / name
if path.exists():
return path.read_bytes()
for attempt in range(1, 6):
req = urllib.request.Request(source, data=data, headers={"User-Agent": "fejkdata data-import", **(headers or {})})
try:
with urllib.request.urlopen(req, timeout=600) as r:
body = r.read()
except OSError:
body = b""
if body and body.startswith(magic) and b"Request Rejected" not in body[:512]:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_bytes(body)
return body
if attempt < 5:
time.sleep(10 * attempt)
sys.exit(f"{source}: no valid download in 5 attempts")
+1 -25
View File
@@ -1,33 +1,9 @@
"""A source fetched once into the cache, and a table written as the loader admits it."""
"""A table written as the loader admits it."""
import re
import sys
import time
import urllib.request
from pathlib import Path
def fetch(source, cache, name, magic=b"", data=None, headers=None):
"""The bytes of a URL, downloaded into cache/name once, or of a local file."""
if not re.match(r"^https?://", source):
return Path(source).read_bytes()
path = Path(cache) / name
for attempt in range(1, 6):
if path.exists():
return path.read_bytes()
req = urllib.request.Request(source, data=data, headers={"User-Agent": "fejkdata data-import", **(headers or {})})
try:
with urllib.request.urlopen(req, timeout=600) as r:
body = r.read()
except OSError:
body = b""
if body and body.startswith(magic) and b"Request Rejected" not in body[:512]:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_bytes(body)
elif attempt < 5:
time.sleep(10 * attempt)
sys.exit(f"{source}: no valid download in 5 attempts")
def write(path, columns, rows):
"""Write the rows as a TSV; every cell must be non-empty and free of tabs, newlines and braces."""
lines = ["\t".join(columns)]
+21
View File
@@ -0,0 +1,21 @@
"""The rows of an xlsx sheet."""
import io
import xml.etree.ElementTree as ET
import zipfile
NS = {"m": "http://schemas.openxmlformats.org/spreadsheetml/2006/main", "r": "http://schemas.openxmlformats.org/officeDocument/2006/relationships"}
def rows(data, sheet=None):
"""The rows of an xlsx sheet named sheet, the first sheet by default, each a list of cell texts."""
z = zipfile.ZipFile(io.BytesIO(data))
strings = ["".join(t.text or "" for t in si.iter("{%s}t" % NS["m"])) for si in ET.fromstring(z.read("xl/sharedStrings.xml")).findall("m:si", NS)]
rels = {r.get("Id"): r.get("Target") for r in ET.fromstring(z.read("xl/_rels/workbook.xml.rels"))}
sheets = {s.get("name"): rels[s.get("{%s}id" % NS["r"])] for s in ET.fromstring(z.read("xl/workbook.xml")).iter("{%s}sheet" % NS["m"])}
target = sheets[sheet] if sheet else next(iter(sheets.values()))
for row in ET.fromstring(z.read("xl/" + target)).findall(".//m:row", NS):
cells = []
for c in row.findall("m:c", NS):
v = c.find("m:v", NS)
cells.append("" if v is None else strings[int(v.text)] if c.get("t") == "s" else v.text)
yield cells
+2 -1
View File
@@ -43,7 +43,7 @@ func loadData(sources []dataSource) (map[string]node, error) {
}
info, err := os.Stat(src.path)
if err != nil {
return nil, err
return nil, fmt.Errorf("data path %s: %w", src.path, err)
}
if !info.IsDir() {
return nil, fmt.Errorf("%s is not a directory", src.path)
@@ -62,6 +62,7 @@ func loadData(sources []dataSource) (map[string]node, error) {
if len(root) == 0 {
return nil, fmt.Errorf("no .json data found")
}
setTablePaths(root)
if err := linkTables(root); err != nil {
return nil, err
}
+1 -13
View File
@@ -1,13 +1 @@
{
"format": "{month}/{day}/{year}",
"month": ["01", "02", "03", "04", "05", "06", "07", "08", "09", "10", "11", "12"],
"day": ["01", "02", "03", "04", "05", "06", "07", "08", "09", "10", "11", "12", "13", "14", "15", "16", "17", "18", "19", "20", "21", "22", "23", "24", "25", "26", "27", "28"],
"year": [
"197{digits(1)}",
"198{digits(1)}",
{ "format": "199{digits(1)}", "weight": 2 },
{ "format": "200{digits(1)}", "weight": 3 },
{ "format": "201{digits(1)}", "weight": 3 },
{ "format": "202{digits(1)}", "weight": 2 }
]
}
"{date(1970-01-01,2029-12-31,'01/02/2006')}"
+1
View File
@@ -0,0 +1 @@
{ "format": "{name}", "rows": "first-name.tsv", "name": "name", "parent": "sex", "weight": "count" }
File diff suppressed because it is too large Load Diff
+16
View File
@@ -0,0 +1,16 @@
{
"format": "{area}-{group}-{serial}",
"area": "9{digits(2)}",
"group": [
{ "format": "{int(50,65)}", "weight": 16 },
{ "format": "{int(70,88)}", "weight": 19 },
{ "format": "{int(90,92)}", "weight": 3 },
{ "format": "{int(94,99)}", "weight": 6 }
],
"serial": [
{ "format": "000{int(1,9)}", "weight": 9 },
{ "format": "00{int(10,99)}", "weight": 90 },
{ "format": "0{int(100,999)}", "weight": 900 },
{ "format": "{int(1000,9999)}", "weight": 9000 }
]
}
+1
View File
@@ -0,0 +1 @@
{ "format": "{name}", "rows": "last-name.tsv", "key": "name", "weight": "count" }
File diff suppressed because it is too large Load Diff
+5 -5
View File
@@ -1,7 +1,7 @@
{
"format": "{prefix}{femalefirst|malefirst} {last}",
"femalefirst": ["Abigail", "Addison", "Amelia", "Aria", "Aubrey", "Audrey", "Aurora", "Ava", "Bella", "Brooklyn", "Camila", "Caroline", "Charlotte", "Chloe", "Claire", "Eleanor", "Elizabeth", "Ella", "Ellie", "Emily", "Emma", "Evelyn", "Gianna", "Grace", "Hannah", "Harper", "Hazel", "Isabella", "Layla", "Leah", "Lillian", "Lily", "Lucy", "Luna", "Madison", "Maya", "Mia", "Mila", "Naomi", "Natalie", "Nora", "Olivia", "Paisley", "Penelope", "Riley", "Savannah", "Scarlett", "Sofia", "Sophia", "Stella", "Victoria", "Violet", "Zoe"],
"malefirst": ["Aiden", "Alexander", "Andrew", "Anthony", "Asher", "Benjamin", "Caleb", "Carter", "Charles", "Christopher", "Daniel", "David", "Dylan", "Elijah", "Ethan", "Ezra", "Gabriel", "Grayson", "Henry", "Isaac", "Jack", "Jackson", "Jacob", "James", "Jayden", "John", "Joseph", "Joshua", "Julian", "Levi", "Liam", "Lincoln", "Logan", "Lucas", "Luke", "Mason", "Mateo", "Matthew", "Michael", "Nathan", "Noah", "Oliver", "Owen", "Samuel", "Sebastian", "Theodore", "Thomas", "William", "Wyatt"],
"last": ["Adams", "Allen", "Anderson", "Bailey", "Baker", "Bell", "Bennett", "Brooks", "Brown", "Campbell", "Carter", "Clark", "Collins", "Cook", "Cooper", "Cox", "Davis", "Edwards", "Evans", "Flores", "Foster", "Garcia", "Gonzalez", "Gray", "Green", "Hall", "Harris", "Hayes", "Hernandez", "Hill", "Howard", "Hughes", "Jackson", "James", "Jenkins", "Johnson", "Jones", "Kelly", "King", "Lee", "Lewis", "Long", "Lopez", "Martin", "Martinez", "Miller", "Mitchell", "Moore", "Morgan", "Morris", "Murphy", "Nelson", "Parker", "Perez", "Perry", "Peterson", "Phillips", "Powell", "Price", "Ramirez", "Reed", "Richardson", "Rivera", "Roberts", "Robinson", "Rodriguez", "Rogers", "Ross", "Russell", "Sanchez", "Sanders", "Scott", "Smith", "Stewart", "Sullivan", "Taylor", "Thomas", "Thompson", "Torres", "Turner", "Walker", "Ward", "Watson", "White", "Williams", "Wilson", "Wood", "Wright", "Young"],
"prefix": ["", { "format": "{title} ", "title": ["Dr", "Miss", "Mr", "Mrs", "Ms", "Mx", "Prof"], "weight": 0.1 }]
"format": "{prefix}{first} {last}",
"first": "{.first-name.name}",
"last": "{.last-name.name}",
"prefix": [null, { "format": "{.title.name} ", "weight": 0.1 }],
"sex": "{.sex.name}"
}
+1
View File
@@ -0,0 +1 @@
{ "format": "{name}", "rows": "sex.tsv", "key": "code", "name": "name" }
+3
View File
@@ -0,0 +1,3 @@
code name
f female
m male
1 code name
2 f female
3 m male
+19 -1
View File
@@ -1 +1,19 @@
"{int(100,999)}-{digits(2)}-{digits(4)}"
{
"format": "{area}-{group}-{serial}",
"area": [
{ "format": "00{int(1,9)}", "weight": 9 },
{ "format": "0{int(10,99)}", "weight": 90 },
{ "format": "{int(100,665)}", "weight": 566 },
{ "format": "{int(667,899)}", "weight": 233 }
],
"group": [
{ "format": "0{int(1,9)}", "weight": 9 },
{ "format": "{int(10,99)}", "weight": 90 }
],
"serial": [
{ "format": "000{int(1,9)}", "weight": 9 },
{ "format": "00{int(10,99)}", "weight": 90 },
{ "format": "0{int(100,999)}", "weight": 900 },
{ "format": "{int(1000,9999)}", "weight": 9000 }
]
}
+1 -6
View File
@@ -1,6 +1 @@
{
"format": "{hour}:{minute} {ampm}",
"hour": ["1", "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12"],
"minute": { "format": "{t}{digits(1)}", "t": ["0", "1", "2", "3", "4", "5"] },
"ampm": ["AM", "PM"]
}
"{time('3:04 PM')}"
+1
View File
@@ -0,0 +1 @@
{ "format": "{name}", "rows": "title.tsv", "name": "name", "parent": "sex", "weight": "share" }
+11
View File
@@ -0,0 +1,11 @@
name sex share
Dr f 6
Dr m 6
Miss f 4
Mr m 60
Mrs f 22
Ms f 30
Mx f 1
Mx m 1
Prof f 3
Prof m 3
1 name sex share
2 Dr f 6
3 Dr m 6
4 Miss f 4
5 Mr m 60
6 Mrs f 22
7 Ms f 30
8 Mx f 1
9 Mx m 1
10 Prof f 3
11 Prof m 3
+1
View File
@@ -0,0 +1 @@
"{date(2000-01-01,2029-12-31,'2006-01-02T15:04:05Z')}"
+1
View File
@@ -0,0 +1 @@
{ "format": "{number}", "rows": "birth-number.tsv", "parent": "sex" }
+3
View File
@@ -0,0 +1,3 @@
sex number
f 238
m 239
1 sex number
2 f 238
3 m 239
+1 -13
View File
@@ -1,13 +1 @@
{
"format": "{year}-{month}-{day}",
"month": ["01", "02", "03", "04", "05", "06", "07", "08", "09", "10", "11", "12"],
"day": ["01", "02", "03", "04", "05", "06", "07", "08", "09", "10", "11", "12", "13", "14", "15", "16", "17", "18", "19", "20", "21", "22", "23", "24", "25", "26", "27", "28"],
"year": [
"197{digits(1)}",
"198{digits(1)}",
{ "format": "199{digits(1)}", "weight": 2 },
{ "format": "200{digits(1)}", "weight": 3 },
{ "format": "201{digits(1)}", "weight": 3 },
{ "format": "202{digits(1)}", "weight": 2 }
]
}
"{date(1970-01-01,2029-12-31,'2006-01-02')}"
+1
View File
@@ -0,0 +1 @@
{ "format": "{name}", "rows": "first-name.tsv", "name": "name", "parent": "sex", "weight": "count" }
File diff suppressed because it is too large Load Diff
+1
View File
@@ -0,0 +1 @@
{ "format": "{name}", "rows": "last-name.tsv", "key": "name", "weight": "count" }
File diff suppressed because it is too large Load Diff
+5 -30
View File
@@ -1,32 +1,7 @@
{
"format": "{prefix}{femalefirst|malefirst} {last}",
"femalefirst": ["Agnes", "Alice", "Alicia", "Alma", "Amanda", "Anna", "Astrid", "Cornelia", "Ebba", "Elin", "Ella", "Ellen", "Elsa", "Emilia", "Emma", "Ester", "Eva", "Frida", "Hanna", "Hedda", "Ida", "Ingrid", "Isabelle", "Julia", "Karin", "Klara", "Kristina", "Lena", "Linnéa", "Lova", "Maja", "Maria", "Moa", "Molly", "Märta", "Nellie", "Nora", "Olivia", "Saga", "Sara", "Selma", "Signe", "Siri", "Sofia", "Stina", "Tilde", "Tuva", "Wilma", "Ylva", "Åsa"],
"malefirst": ["Adam", "Albin", "Alexander", "Alfred", "Anders", "Anton", "Arvid", "Axel", "Bengt", "Bo", "Carl", "David", "Edvin", "Elias", "Emil", "Erik", "Filip", "Folke", "Fredrik", "Gustav", "Göran", "Hampus", "Hans", "Henrik", "Hugo", "Isak", "Johan", "Jonas", "Karl", "Kjell", "Lars", "Leo", "Linus", "Love", "Magnus", "Mats", "Mattias", "Mikael", "Nils", "Olle", "Oskar", "Otto", "Patrik", "Per", "Pontus", "Rasmus", "Samuel", "Sixten", "Stefan", "Sten", "Sven", "Theodor", "Tobias", "Viktor", "William", "Åke"],
"last": [
{
"format": "{first}sson",
"weight": 5,
"first": ["Ander", "Arvid", "Bengt", "Daniel", "David", "Erik", "Fredrik", "Gustaf", "Henrik", "Håkan", "Isak", "Jakob", "Johan", "Jon", "Jön", "Karl", "Lar", "Magnu", "Matt", "Mikael", "Mårten", "Nil", "Ol", "Per", "Petter", "Samuel", "Sven"]
},
{
"format": "{first}{last}",
"weight": 3.6,
"first": ["Ahl", "Berg", "Blom", "Ceder", "Dahl", "Ek", "Eng", "Falk", "Fors", "Gran", "Hag", "Hed", "Holm", "Lind", "Lund", "Mal", "Ny", "Rosen", "Sand", "Sjö", "Skog", "Söder", "Sten", "Strand", "Sund", "Wik", "Öst"],
"last": ["berg", "blad", "crona", "dahl", "ed", "fors", "gren", "holm", "in", "kvist", "löf", "lund", "man", "mark", "qvist", "roth", "stedt", "sten", "strand", "ström", "vall"]
},
{
"format": "{first}{last}",
"weight": 0.267,
"first": ["Hell", "Wall"],
"last": ["berg", "blad", "crona", "dahl", "ed", "fors", "gren", "holm", "in", "kvist", "man", "mark", "qvist", "roth", "stedt", "sten", "strand", "ström", "vall"]
},
{
"format": "{first}{last}",
"weight": 0.133,
"first": "Norr",
"last": ["berg", "blad", "crona", "dahl", "ed", "fors", "gren", "holm", "in", "kvist", "löf", "lund", "man", "mark", "qvist", "stedt", "sten", "strand", "ström", "vall"]
},
["Berg", "Blom", "Eismar", "Falk", "Holm", "Lind", "Norberg", "Strand", "Ström", "von Flemming", "Åberg", "Öberg"]
],
"prefix": ["", { "format": "{string} ", "string": ["dr", "prof"], "weight": 0.05 }]
"format": "{prefix}{first} {last}",
"first": "{.first-name.name}",
"last": "{.last-name.name}",
"prefix": [null, { "format": "{.title.name} ", "weight": 0.05 }],
"sex": "{.sex.name}"
}
+1
View File
@@ -0,0 +1 @@
"{date(1930-01-01,2025-12-31,'060102')}-{.birth-number.number}{luhn()}"
+1
View File
@@ -0,0 +1 @@
"{date(1930-01-01,2025-12-31,'0601')}{int(61,88)}-{.birth-number.number}{luhn()}"
+1
View File
@@ -0,0 +1 @@
{ "format": "{name}", "rows": "sex.tsv", "key": "code", "name": "name" }
+3
View File
@@ -0,0 +1,3 @@
code name
f kvinna
m man
1 code name
2 f kvinna
3 m man
-22
View File
@@ -1,22 +0,0 @@
{
"format": "{digits(2)}{mmdd}-{digits(3)}{luhn()}",
"mmdd": [
{
"format": "{m}{d}",
"weight": 7,
"m": ["01", "03", "05", "07", "08", "10", "12"],
"d": ["01", "02", "03", "04", "05", "06", "07", "08", "09", "10", "11", "12", "13", "14", "15", "16", "17", "18", "19", "20", "21", "22", "23", "24", "25", "26", "27", "28", "29", "30", "31"]
},
{
"format": "{m}{d}",
"weight": 4,
"m": ["04", "06", "09", "11"],
"d": ["01", "02", "03", "04", "05", "06", "07", "08", "09", "10", "11", "12", "13", "14", "15", "16", "17", "18", "19", "20", "21", "22", "23", "24", "25", "26", "27", "28", "29", "30"]
},
{
"format": "{m}{d}",
"m": "02",
"d": ["01", "02", "03", "04", "05", "06", "07", "08", "09", "10", "11", "12", "13", "14", "15", "16", "17", "18", "19", "20", "21", "22", "23", "24", "25", "26", "27", "28"]
}
]
}
+1 -17
View File
@@ -1,17 +1 @@
{
"format": "{hour}:{minute}{sec}",
"hour": [
{ "format": "0{digits(1)}", "weight": 4 },
{ "format": "1{digits(1)}", "weight": 4 },
{ "format": "2{t}", "t": ["0", "1", "2", "3"], "weight": 2 }
],
"minute": { "format": "{t}{digits(1)}", "t": ["0", "1", "2", "3", "4", "5"] },
"sec": [
"",
{
"format": ":{s}",
"s": { "format": "{t}{digits(1)}", "t": ["0", "1", "2", "3", "4", "5"] },
"weight": 0.4
}
]
}
"{time('15:04')}"
+1
View File
@@ -0,0 +1 @@
{ "format": "{name}", "rows": "title.tsv", "key": "name", "weight": "share" }
+3
View File
@@ -0,0 +1,3 @@
name share
dr 3
prof 1
1 name share
2 dr 3
3 prof 1
+139 -18
View File
@@ -10,7 +10,7 @@ import (
// TestShippedDataCategories asserts every shipped locale emits only well-formed
// values for each data category. en and sv patterns differ where the format is
// locale-specific (date, time, ssn, company, price, ...); the rest are shared.
// locale-specific (date, time, company, price, ...); the rest are shared.
func TestShippedDataCategories(t *testing.T) {
letters := regexp.MustCompile(`^[\pL'-]+$`)
semver := regexp.MustCompile(`^v?\d+\.\d+\.\d+(-(alpha|beta|rc)\.\d+)?$`)
@@ -33,10 +33,7 @@ func TestShippedDataCategories(t *testing.T) {
regexp.MustCompile(`^\d{4}-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])$`)},
{"time",
regexp.MustCompile(`^([1-9]|1[0-2]):[0-5]\d (AM|PM)$`),
regexp.MustCompile(`^([01]\d|2[0-3]):[0-5]\d(:[0-5]\d)?$`)},
{"ssn",
regexp.MustCompile(`^[1-9]\d{2}-\d{2}-\d{4}$`),
regexp.MustCompile(`^\d{2}(0[1-9]|1[0-2])(0[1-9]|[12]\d|3[01])-\d{4}$`)},
regexp.MustCompile(`^([01]\d|2[0-3]):[0-5]\d$`)},
{"version", semver, semver},
{"email", email, email},
{"ip", ip, ip},
@@ -72,7 +69,11 @@ func TestShippedMiscCategories(t *testing.T) {
v4 := regexp.MustCompile(`^[0-9a-f]{8}-[0-9a-f]{4}-4[0-9a-f]{3}-[89ab][0-9a-f]{3}-[0-9a-f]{12}$`)
mac := regexp.MustCompile(`^([0-9a-f]{2}:){5}[0-9a-f]{2}$`)
objectid := regexp.MustCompile(`^[0-9a-f]{24}$`)
datetime := regexp.MustCompile(`^\d{4}-\d\d-\d\dT\d\d:\d\d:\d\dZ$`)
for i := 0; i < 200; i++ {
if v := fake(t, f, "datetime"); !datetime.MatchString(v) {
t.Fatalf("misc datetime %q is not an RFC 3339 UTC instant", v)
}
if v := fake(t, f, "objectid"); !objectid.MatchString(v) {
t.Fatalf("misc objectid %q is not 24 hex chars", v)
}
@@ -147,30 +148,122 @@ func TestSwedishPersonNamesHaveNoTripleLetter(t *testing.T) {
}
}
// TestSwedishPersonnummer checks the two rules the shape regex can't: the date
// is a real calendar date (so month-length variants never emit e.g. Apr 31 or
// Feb 30) and the trailing digit is a valid Luhn checksum over the other nine.
// TestSwedishPersonnummer checks what the shape regex can't: the date is a real
// calendar date, the birth number is Skatteverket's test series (238 female, 239
// male), the trailing digit is a Luhn checksum over the other nine, and a
// samordningsnummer is the same with 60 added to the day.
func TestSwedishPersonnummer(t *testing.T) {
sv := newGenerator(t, "data", WithSeed(1))
sawLongMonthEnd := false
re := regexp.MustCompile(`^\d{6}-23[89]\d$`)
sawLongMonthEnd, birth := false, map[string]bool{}
for i := 0; i < 2000; i++ {
v := fake(t, sv, "sv_SE.ssn")
v := fake(t, sv, "sv_SE.personnummer")
d := digitsOnly(v)
if len(d) != 10 {
t.Fatalf("ssn %q has %d digits, want 10", v, len(d))
if !re.MatchString(v) {
t.Fatalf("personnummer %q, want YYMMDD-238C or YYMMDD-239C", v)
}
if _, err := time.Parse("060102", d[:6]); err != nil { // 2-digit year, real-date check
t.Fatalf("ssn %q is not a valid calendar date: %v", v, err)
if _, err := time.Parse("060102", d[:6]); err != nil {
t.Fatalf("personnummer %q is not a valid calendar date: %v", v, err)
}
if !luhnValid(d) {
t.Fatalf("ssn %q fails the Luhn check", v)
t.Fatalf("personnummer %q fails the Luhn check", v)
}
if d[4:6] == "31" {
sawLongMonthEnd = true
sawLongMonthEnd = sawLongMonthEnd || d[4:6] == "31"
birth[d[6:9]] = true
s := fake(t, sv, "sv_SE.samordningsnummer")
sd := digitsOnly(s)
day, err := strconv.Atoi(sd[4:6])
if !re.MatchString(s) || err != nil || day < 61 || day > 88 || !luhnValid(sd) {
t.Fatalf("samordningsnummer %q, want YYMM(61-88)-23[89]C, Luhn-valid", s)
}
if _, err := time.Parse("0601", sd[:4]); err != nil {
t.Fatalf("samordningsnummer %q is not a valid year and month: %v", s, err)
}
}
if !sawLongMonthEnd {
t.Fatal("never generated a 31st — 31-day months are not reaching their last day")
t.Fatal("never generated a 31st")
}
if !birth["238"] || !birth["239"] {
t.Fatalf("birth numbers drawn %v, want both 238 and 239", birth)
}
for i := 0; i < 200; i++ {
got := fakeTemplate(t, sv, `{/sv_SE.person.sex} {/sv_SE.personnummer}`)
sex, id, _ := strings.Cut(got, " ")
if want := map[string]string{"kvinna": "238", "man": "239"}[sex]; want == "" || !strings.Contains(id, "-"+want) {
t.Fatalf("%q: a person and a personnummer in one render disagree on sex", got)
}
}
}
// TestShippedUSTaxIds pins the SSA and IRS ranges: an SSN's area is 001-899 but
// 666, its group 01-99 and its serial 0001-9999; an ITIN is 9XX-GG-XXXX with GG
// in 50-65, 70-88, 90-92 or 94-99.
func TestShippedUSTaxIds(t *testing.T) {
f := newGenerator(t, "data", WithSeed(5))
ssn := regexp.MustCompile(`^(\d{3})-(\d{2})-(\d{4})$`)
itin := regexp.MustCompile(`^9\d{2}-(5\d|6[0-5]|7\d|8[0-8]|9[0-24-9])-\d{4}$`)
for i := 0; i < 2000; i++ {
v := fake(t, f, "en_US.ssn")
m := ssn.FindStringSubmatch(v)
if m == nil {
t.Fatalf("ssn %q, want AAA-GG-SSSS", v)
}
area, _ := strconv.Atoi(m[1])
group, _ := strconv.Atoi(m[2])
serial, _ := strconv.Atoi(m[3])
if area < 1 || area > 899 || area == 666 || group < 1 || serial < 1 {
t.Fatalf("ssn %q is in a range the SSA never assigns", v)
}
if v := fake(t, f, "en_US.itin"); !itin.MatchString(v) {
t.Fatalf("itin %q, want %s", v, itin)
}
}
}
// TestShippedPersonNames pins the name tables: a sex pins its first names, a name
// both sexes carry appears under both, and the record's sex agrees with its name.
func TestShippedPersonNames(t *testing.T) {
f := newGenerator(t, "data", WithSeed(9))
for path, want := range map[string]string{
"sv_SE.sex[f]": "kvinna",
"sv_SE.sex[man].code": "m",
"en_US.sex[f]": "female",
"sv_SE.sex[f].first-name[Anna]": "Anna",
"sv_SE.sex[m].first-name[Erik].sex": "m",
"en_US.sex[m].first-name[James]": "James",
"en_US.sex[f].first-name[Taylor].sex": "f",
"en_US.sex[m].first-name[Taylor].sex": "m",
"sv_SE.last-name[Andersson]": "Andersson",
"sv_SE.last-name[Andersson].count": "",
"en_US.last-name[Smith]": "Smith",
} {
got := fake(t, f, path)
if want == "" && !regexp.MustCompile(`^[1-9]\d*$`).MatchString(got) {
t.Errorf("Fake(%q) = %q, want a count", path, got)
} else if want != "" && got != want {
t.Errorf("Fake(%q) = %q, want %q", path, got, want)
}
}
if _, err := f.Fake("en_US.first-name[Taylor]"); err == nil || !strings.Contains(err.Error(), "sex[f].first-name[Taylor]") {
t.Fatalf("Fake(en_US.first-name[Taylor]) = %v, want both sexes' rows listed", err)
}
for _, locale := range []string{"sv_SE", "en_US"} {
seen := map[string]bool{}
for i := 0; i < 2000; i++ {
seen[fake(t, f, locale+".sex[f].first-name")] = true
first, sex := fake(t, f, locale+".person.first"), fake(t, f, locale+".person.sex")
if first == "" || sex == "" {
t.Fatalf("%s.person lacks a first name or a sex", locale)
}
}
if len(seen) < 300 || !seen["Anna"] && locale == "sv_SE" || !seen["Mary"] && locale == "en_US" {
t.Fatalf("%s female first names: %d distinct in 2000, want a weighted register", locale, len(seen))
}
got := fakeTemplate(t, f, `{/`+locale+`.sex.code}|{/`+locale+`.person.first}`)
code, first, _ := strings.Cut(got, "|")
if v := fake(t, f, locale+".sex["+code+"].first-name["+first+"]"); v != first {
t.Fatalf("%s: %q drawn as a %s name is not one", locale, first, code)
}
}
}
@@ -300,3 +393,31 @@ func TestShippedStreetNumberFormats(t *testing.T) {
}
}
}
// TestShippedUSTitleAgreesWithSex pins that one person's title and sex agree, and
// that the everyday titles are reachable.
func TestShippedUSTitleAgreesWithSex(t *testing.T) {
f := newGenerator(t, "data", WithSeed(4))
female := map[string]bool{"Miss": true, "Mrs": true, "Ms": true}
seen := map[string]bool{}
for i := 0; i < 3000; i++ {
got := fakeTemplate(t, f, `{/en_US.person.prefix}|{/en_US.person.sex}`)
title, sex, _ := strings.Cut(got, "|")
title = strings.TrimSpace(title)
seen[title] = true
if title == "Mr" && sex != "male" || female[title] && sex != "female" {
t.Fatalf("%q: the title contradicts the sex", got)
}
}
for _, want := range []string{"", "Mr", "Ms"} {
if !seen[want] {
t.Errorf("en_US.person.prefix never drew %q in 3000 draws", want)
}
}
swedish := map[string]bool{"dr": true, "prof": true}
for i := 0; i < 50; i++ {
if got := fake(t, f, "sv_SE.title"); !swedish[got] {
t.Fatalf("sv_SE.title = %q, want an unsexed Swedish honorific", got)
}
}
}
+1 -1
View File
@@ -97,7 +97,7 @@ func (d *draws) rowOf(t *table) int {
// pinRow pins row r of t where it agrees with the rows pinned before it.
func (d *draws) pinRow(t *table, r int) error {
if pr, ok := d.pinned(t); ok && pr != r {
return fmt.Errorf("%s and %s are two rows of %s", t.selectorSpelling(pr), t.selectorSpelling(r), t.category)
return fmt.Errorf("%s and %s are two rows of %s", t.selectorSpelling(pr), t.selectorSpelling(r), t.path)
}
for a := t.parentT; a != nil; a = a.parentT {
if pa, ok := d.pinned(a); ok && !t.under(r, a, pa) {
+11
View File
@@ -81,3 +81,14 @@ func TestDifferentSeedsDiffer(t *testing.T) {
}
t.Fatal("seeds 1 and 2 produced identical sequences")
}
// fakeTemplate renders an inline template against a loaded generator, so its
// references resolve.
func fakeTemplate(t *testing.T, f *Generator, s string) string {
t.Helper()
got, err := f.FakeTemplate(s)
if err != nil {
t.Fatalf("FakeTemplate(%s) = %v", s, err)
}
return got
}
+146
View File
@@ -0,0 +1,146 @@
package fejkdata
import (
"fmt"
"strings"
"time"
)
const dayLayout = "2006-01-02"
// The instants a layout is proved against: layoutProbe2 is alike in no field, while
// layoutDay differs from layoutProbe in its date fields alone and layoutClock in its
// clock fields alone, so formatting two of them tells which kind a layout names.
var (
layoutProbe = time.Date(2001, 2, 3, 4, 5, 6, 0, time.UTC)
layoutProbe2 = time.Date(2010, 11, 12, 13, 14, 15, 0, time.UTC)
layoutDay = time.Date(2010, 11, 12, 4, 5, 6, 0, time.UTC)
layoutClock = time.Date(2001, 2, 3, 13, 14, 15, 0, time.UTC)
)
func namesAField(layout string) bool {
return layoutProbe.Format(layout) != layoutProbe2.Format(layout)
}
func namesADateField(layout string) bool {
return layoutProbe.Format(layout) != layoutDay.Format(layout)
}
func namesAClockField(layout string) bool {
return layoutProbe.Format(layout) != layoutClock.Format(layout)
}
// quotedLayout reports whether an arg carries the single quotes a layout is written in.
func quotedLayout(a string) bool {
return len(a) >= 2 && a[0] == '\'' && a[len(a)-1] == '\''
}
// layoutArg is the Go layout a quoted arg holds, refused when unquoted or constant.
func layoutArg(a string) (string, error) {
if !quotedLayout(a) {
bare := strings.Trim(a, `'"`)
if strings.HasPrefix(a, `"`) || strings.HasSuffix(a, `"`) {
return "", fmt.Errorf("layout %s is double-quoted; write '%s'", a, bare)
}
return "", fmt.Errorf("layout %s is not quoted; write '%s'", a, bare)
}
layout := a[1 : len(a)-1]
if !namesAField(layout) {
return "", fmt.Errorf("layout '%s' names no field, so it is the constant %q; write it as text", layout, layout)
}
return layout, nil
}
// layoutOf is layoutArg for an arg a check already validated.
func layoutOf(a string) string {
layout, err := layoutArg(a)
if err != nil {
panic(fmt.Sprintf("fejkdata: builtin arg %q reached prep unvalidated: %v", a, err))
}
return layout
}
// layoutArity checks a call ending in a layout takes n args, naming the quoted
// layout where an unquoted one split into more.
func layoutArity(name string, n int, a []string) error {
if len(a) == n {
return nil
}
hint := ""
if len(a) > n && !holdsQuotedLayout(a[n-1:]) {
hint = fmt.Sprintf("; a layout holding a comma is quoted: '%s'", strings.Trim(strings.Join(a[n-1:], ", "), `'"`))
}
return fmt.Errorf("%s takes %d argument%s, got %d%s", name, n, plural(n), len(a), hint)
}
// holdsQuotedLayout reports whether the surplus args already carry a quoted layout,
// which no comma split apart.
func holdsQuotedLayout(a []string) bool {
for _, arg := range a {
if quotedLayout(arg) {
return true
}
}
return false
}
func dateArgs(_ map[string]node, a []string) error {
if err := layoutArity("date", 3, a); err != nil {
return err
}
from, err := time.Parse(dayLayout, a[0])
if err != nil {
return fmt.Errorf("date(from,to,layout): from %q is not a YYYY-MM-DD date", a[0])
}
to, err := time.Parse(dayLayout, a[1])
if err != nil {
return fmt.Errorf("date(from,to,layout): to %q is not a YYYY-MM-DD date", a[1])
}
if to.Before(from) {
return fmt.Errorf("date(from,to,layout): from %s is after to %s", a[0], a[1])
}
layout, err := layoutArg(a[2])
if err != nil {
return err
}
if !namesADateField(layout) {
return fmt.Errorf("date(from,to,layout): '%s' names no date field; write time('%s')", layout, layout)
}
if from.Equal(to) && !namesAClockField(layout) {
return fmt.Errorf("date(%s,%s,'%s') is the constant %q; write it as text", a[0], a[1], layout, from.Format(layout))
}
return nil
}
func timeArg(_ map[string]node, a []string) error {
if err := layoutArity("time", 1, a); err != nil {
return err
}
layout, err := layoutArg(a[0])
if err != nil {
return err
}
if namesADateField(layout) {
return fmt.Errorf("time(layout): '%s' names a date field; write date(from,to,layout)", layout)
}
return nil
}
// datePrep draws a second in [from 00:00:00, to 23:59:59] UTC; the span is counted
// in seconds, since a Duration overflows past 292 years.
func datePrep(a []string) callFn {
from, err := time.Parse(dayLayout, a[0])
if err != nil {
panic(fmt.Sprintf("fejkdata: builtin arg %q reached prep unvalidated: %v", a[0], err))
}
to, err := time.Parse(dayLayout, a[1])
if err != nil {
panic(fmt.Sprintf("fejkdata: builtin arg %q reached prep unvalidated: %v", a[1], err))
}
layout, start, span := layoutOf(a[2]), from.Unix(), int(to.Unix()-from.Unix())+86400
return func(s *session, _ string, _ []string) string {
return time.Unix(start+int64(s.IntN(span)), 0).UTC().Format(layout)
}
}
func timePrep(a []string) callFn {
layout := layoutOf(a[0])
return func(s *session, _ string, _ []string) string {
return time.Unix(int64(s.IntN(86400)), 0).UTC().Format(layout)
}
}
+9 -1
View File
@@ -140,6 +140,10 @@ func (f *Generator) FakeRecord(path string) (*Record, error) {
return nil, fmt.Errorf("fejkdata: %s descends into %q, a field; only a category-level template is a record", path, tail[0])
}
shape := f.recordShapeOf(n)
if errors.Is(shape.err, ErrNoColumns) {
ns := names(segments)
return nil, fmt.Errorf(`fejkdata: %s %w; render it as a column of one: {"format":"","%s":"{/%s}"}`, path, shape.err, ns[len(ns)-1], path)
}
if shape.err != nil {
return nil, fmt.Errorf("fejkdata: %s %w", path, shape.err)
}
@@ -227,6 +231,10 @@ func (f *Generator) FakeRecordTemplate(input string) (*Record, error) {
return t.Fake(), nil
}
// ErrNoColumns is the one record fence a path can answer, so its entry point names
// the record to write instead.
var ErrNoColumns = errors.New("has no fields, so no columns")
// recordOf is the fence both record entry points pass. The columns come back with
// the template, fixed for every draw the caller goes on to make.
func recordOf(n node) (*template, []Column, error) {
@@ -239,7 +247,7 @@ func recordOf(n node) (*template, []Column, error) {
}
names := recordColumns(t)
if len(names) == 0 {
return nil, nil, errors.New("has no fields, so no columns")
return nil, nil, ErrNoColumns
}
if !t.record {
return nil, nil, fmt.Errorf("carries repeat %d, which composes its format into one string; a record projects columns instead — drop the repeat and render the record again for more rows", t.repeat)
+15
View File
@@ -547,3 +547,18 @@ func TestRecordTemplateRejectsATopLevelRepeat(t *testing.T) {
t.Errorf("NewTemplate on the same input = %v, want the string view to still compile", err)
}
}
// TestRecordOfAFieldlessCategoryNamesTheWrapper pins the way out of the first
// command a fixture author types: a category of one value has no columns, and the
// error names the record that gives it one.
func TestRecordOfAFieldlessCategoryNamesTheWrapper(t *testing.T) {
f := newGenerator(t, "data", WithSeed(1))
want := `{"format":"","personnummer":"{/sv_SE.personnummer}"}`
_, err := f.FakeRecord("sv_SE.personnummer")
if err == nil || !strings.Contains(err.Error(), want) {
t.Fatalf("FakeRecord(sv_SE.personnummer) = %v, want it to name %s", err, want)
}
if _, err := f.FakeRecordTemplate(`{"format":"","personnummer":"{/sv_SE.personnummer}"}`); err != nil {
t.Fatalf("the named wrapper does not render: %v", err)
}
}
+1 -1
View File
@@ -340,7 +340,7 @@ func (p *valueProof) checkField(label string, ft reflect.Type, column node) erro
return fmt.Errorf("%s (%s): %s", label, ft, reason)
}
if !kind.holds(v) {
return fmt.Errorf("%s (%s): %q is not proven within %s; make it %s", label, ft, it.format, elem.Kind(), kind.wider)
return fmt.Errorf("%s (%s): %q is not proven within %s; narrow it to that range, or make the field %s", label, ft, it.format, elem.Kind(), kind.wider)
}
}
return nil
+4 -4
View File
@@ -245,16 +245,16 @@ func TestFakeStructErrors(t *testing.T) {
}{}, `"Ada" is not an integer`},
{&struct {
A int8 `fake:"{int(0,300)}"`
}{}, `"{int(0,300)}" is not proven within int8; make it int64`},
}{}, `"{int(0,300)}" is not proven within int8; narrow it to that range, or make the field int64`},
{&struct {
A uint `fake:"{int(-1,5)}"`
}{}, `"{int(-1,5)}" is not proven within uint; make it int64`},
}{}, `"{int(-1,5)}" is not proven within uint; narrow it to that range, or make the field int64`},
{&struct {
A float32 `fake:"[\"1\",\"1e39\"]"`
}{}, `"1e39" is not proven within float32; make it float64`},
}{}, `"1e39" is not proven within float32; narrow it to that range, or make the field float64`},
{&struct {
A int32 `fake:"{seq()}"`
}{}, `"{seq()}" is not proven within int32; make it int64`},
}{}, `"{seq()}" is not proven within int32; narrow it to that range, or make the field int64`},
{&struct {
A bool `fake:"{int(0,1)}"`
}{}, "prints an integer, not a boolean"},
+69 -11
View File
@@ -14,6 +14,7 @@ import (
// cell of the row a render pinned; a cell carrying tokens compiles to a string node.
type table struct {
category string
path string // the category's path from the data root, which a selector is written at
file string
format *template // fields are the column nodes
columns []string
@@ -222,15 +223,34 @@ func (t *table) bindOptions(o tableOptionValues) error {
}
*opt.into = i
}
if t.name >= 0 && t.key < 0 {
return fmt.Errorf("name selects a row as a key does, and lists the rows it matches by their keys, so it needs a key column; add key")
if t.name >= 0 && t.key < 0 && t.parent < 0 {
return fmt.Errorf("name selects a row as a key does, and lists the rows it matches by their keys, so it needs a key column, or a parent inside which each name is one row; add key")
}
if err := t.indexKeys(); err != nil {
return err
}
if err := t.proveNamesInsideParent(); err != nil {
return err
}
return t.sumWeights()
}
// proveNamesInsideParent proves a name without a key names one row inside its parent.
func (t *table) proveNamesInsideParent() error {
if t.key >= 0 || t.name < 0 {
return nil
}
inside := make(map[string]int, t.rows())
for r := 0; r < t.rows(); r++ {
k := t.cell(r, t.parent) + "\t" + t.cell(r, t.name)
if first, dup := inside[k]; dup {
return fmt.Errorf("%s line %d: name %q repeats line %d inside %s %q; a name selects one row inside its parent; drop one, or add a key column", t.file, r+2, t.cell(r, t.name), first+2, t.columns[t.parent], t.cell(r, t.parent))
}
inside[k] = r
}
return nil
}
// indexKeys proves every key names one row, and keeps the index a link is proved by.
func (t *table) indexKeys() error {
if t.key < 0 {
@@ -317,6 +337,23 @@ func (t *table) compileFormat(format string) error {
return t.format.compileFormat()
}
// setTablePaths gives every table the path a selector on it is written at, before a
// link or a draw can name one.
func setTablePaths(root map[string]node) {
var walk func(dir string, children map[string]node)
walk = func(dir string, children map[string]node) {
for name, child := range children {
switch n := child.(type) {
case *folder:
walk(join(dir, name), n.children)
case *table:
n.path = join(dir, name)
}
}
}
walk("", root)
}
// linkTables binds every table's parent to the table beside it, and proves the
// links: a parent has a key, every link cell is one, every parent row is linked
// to, no chain of parents closes, and no child is named like a parent's column.
@@ -488,8 +525,8 @@ func (t *table) under(r int, a *table, pr int) bool {
// find is the row a selector names: by key first, then by name, where a name
// naming several rows resolves inside the ancestors pinned in d.
func (t *table) find(sel string, d *draws) (int, error) {
if t.key < 0 {
return 0, fmt.Errorf("%s has no key column to select a row by", t.category)
if t.key < 0 && t.name < 0 {
return 0, fmt.Errorf("%s has no key or name column to select a row by", t.path)
}
if r, ok := t.byKey[sel]; ok {
return r, nil
@@ -502,20 +539,41 @@ func (t *table) find(sel string, d *draws) (int, error) {
case 1:
return rows[0], nil
case 0:
return 0, fmt.Errorf("no row of %s has key or name %q", t.category, sel)
return 0, fmt.Errorf("no row of %s has key or name %q", t.path, sel)
}
keys := make([]string, len(rows))
return 0, t.ambiguous(sel, rows)
}
// ambiguous names each row a selector could have meant: a keyed table by its keys,
// and one told apart only by its parent by the path that selects it.
func (t *table) ambiguous(sel string, rows []int) error {
spellings := make([]string, len(rows))
for i, r := range rows {
keys[i] = t.cell(r, t.key)
if t.key < 0 {
spellings[i] = t.selectorSpelling(r)
} else {
spellings[i] = t.cell(r, t.key)
}
}
listed := strings.Join(spellings, ", ")
if t.key < 0 {
return fmt.Errorf("%q names %d rows of %s; select it inside its %s, one of %s", sel, len(rows), t.path, t.parentT.path, listed)
}
inside := ""
if t.parentT != nil {
inside = fmt.Sprintf(", or select it inside its %s", t.parentT.category)
inside = fmt.Sprintf(", or select it inside its %s", t.parentT.path)
}
return 0, fmt.Errorf("%q names %d rows of %s; select one by key, one of %v%s", sel, len(rows), t.category, keys, inside)
return fmt.Errorf("%q names %d rows of %s; select one by key, one of %s%s", sel, len(rows), t.path, listed, inside)
}
// selectorSpelling is how a path writes a selected row, for messages.
// selectorSpelling is how a path writes a selected row, for messages: by key, or
// by name inside its parent's row.
func (t *table) selectorSpelling(r int) string {
return t.category + "[" + t.cell(r, t.key) + "]"
switch {
case t.key >= 0:
return t.path + "[" + t.cell(r, t.key) + "]"
case t.name >= 0:
return t.parentT.selectorSpelling(t.parentRow(r)) + "." + t.category + "[" + t.cell(r, t.name) + "]"
}
return fmt.Sprintf("%s line %d", t.file, r+2)
}
+66 -2
View File
@@ -172,7 +172,7 @@ func TestTableSelectsARowByKeyOrName(t *testing.T) {
}
for path, want := range map[string]string{
"region[99]": `"99"`,
"locality[Sandby]": "L7 L8",
"locality[Sandby]": "L7, L8",
"region[12].code[1]": "not a table",
"region[]": "empty",
"region[12": "]",
@@ -334,7 +334,7 @@ func TestTableSelectorInAReference(t *testing.T) {
}
for name, c := range map[string]struct{ json, want string }{
"unknown row": {`"{/region[99].name}"`, `"99"`},
"ambiguous name": {`"{/locality[Sandby].name}"`, "L7 L8"},
"ambiguous name": {`"{/locality[Sandby].name}"`, "L7, L8"},
"not inside": {`"{/region[12].municipality[0180].name}"`, "not inside"},
"selector on template": {`"{/x[1].a}"`, "not a table"},
} {
@@ -440,6 +440,7 @@ func TestTableFences(t *testing.T) {
"a bracket in a key": {map[string]string{"t.json": `{"format":"{a}","rows":"t.tsv","key":"a"}`, "t.tsv": "a\nx[1]\ny\n"}, `"["`},
"a brace in a name": {map[string]string{"t.json": `{"format":"{a}","rows":"t.tsv","key":"a","name":"n"}`, "t.tsv": "a\tn\nx\tx{1}\ny\ty\n"}, `"{"`},
"name without a key": {map[string]string{"t.json": `{"format":"{a}","rows":"t.tsv","name":"a"}`, "t.tsv": "a\nx\nx\n"}, "key"},
"a name repeating inside one parent row": {with(siblings(), map[string]string{"street.json": `{"format":"{name}","rows":"street.tsv","name":"name","parent":"locality"}`, "street.tsv": "name\tlocality\nAvenyn\tL1\nAvenyn\tL1\nStorgatan\tL2\nStorgatan\tL3\nStorgatan\tL4\nStorgatan\tL5\nStorgatan\tL6\nStorgatan\tL7\nStorgatan\tL8\n"}), `"Avenyn"`},
"a name that is another row's key": {map[string]string{"t.json": `{"format":"{a}","rows":"t.tsv","key":"a","name":"n"}`, "t.tsv": "a\tn\nx\ty\ny\tz\n"}, `"y"`},
"a cell reading its family": {with(geo(), map[string]string{"locality.tsv": "code\tname\tmunicipality\tnote\nL1\tStockholm\t0180\t{/municipality.code}\nL2\tSolna\t0184\t-\nL3\tMalmö\t1280\t-\nL4\tLund\t1281\t-\nL5\tGöteborg\t1480\t-\n"}), "family"},
"a format reading its family": {with(geo(), map[string]string{"locality.json": `{"format":"{name} {/region.name}","rows":"locality.tsv","key":"code","name":"name","parent":"municipality"}`}), "family"},
@@ -709,3 +710,66 @@ func TestShippedTables(t *testing.T) {
t.Fatalf("country draws %d distinct rows in 5000, want the full register", len(count))
}
}
// TestNamedTableWithoutAKeyResolvesInsideItsParent pins a table whose rows are
// told apart only by their parent: a name selects a row inside the parent pinned
// before it, an ambiguous one is listed by its parent's spelling, and a name
// repeating inside one parent row is a load error (see TestTableFences).
func TestNamedTableWithoutAKeyResolvesInsideItsParent(t *testing.T) {
files := siblings()
files["street.json"] = `{"format":"{name}","rows":"street.tsv","name":"name","parent":"locality","weight":"segments"}`
f := newGenerator(t, writeFiles(t, files), WithSeed(1))
for path, want := range map[string]string{
"street[Avenyn]": "Avenyn",
"street[Avenyn].locality": "L5",
"locality[L7].street[Sandbyvägen]": "Sandbyvägen",
"locality[L7].street[Sandbyvägen].segments": "2",
"municipality[0184].locality[Sandby].street[Sandbyvägen]": "Sandbyvägen",
"municipality[0184].locality[Sandby].street[Sandbyvägen].locality": "L8",
} {
if got := fake(t, f, path); got != want {
t.Errorf("Fake(%q) = %q, want %q", path, got, want)
}
}
for path, want := range map[string]string{
"street[Sandbyvägen]": "locality[L7].street[Sandbyvägen]",
"locality[L1].street[Avenyn]": "not inside",
"street[Kungsgatan]": `"Kungsgatan"`,
} {
if _, err := f.Fake(path); err == nil || !strings.Contains(err.Error(), want) {
t.Errorf("Fake(%q) = %v, want an error mentioning %s", path, err, want)
}
}
if got := fakeTemplate(t, f, `{/locality[L8].name}: {/locality[L8].street[Sandbyvägen].name}`); got != "Sandby: Sandbyvägen" {
t.Fatalf("a name selected inside a pinned parent = %q", got)
}
}
// TestAmbiguousNameNamesARunnablePath pins what a shell user reads: each row is
// named as the path they can type, from the root, and the rows are listed plainly.
func TestAmbiguousNameNamesARunnablePath(t *testing.T) {
files := map[string]string{}
for name, body := range siblings() {
files["se/"+name] = body
}
files["se/street.json"] = `{"format":"{name}","rows":"street.tsv","name":"name","parent":"locality","weight":"segments"}`
f := newGenerator(t, writeFiles(t, files), WithSeed(1))
_, err := f.Fake("se.street[Sandbyvägen]")
if err == nil {
t.Fatal("se.street[Sandbyvägen] resolved, though two rows carry that name")
}
for _, want := range []string{"se.locality[L7].street[Sandbyvägen]", "se.locality[L8].street[Sandbyvägen]"} {
if !strings.Contains(err.Error(), want) {
t.Errorf("%v does not name %s", err, want)
}
if got := fake(t, f, want); got != "Sandbyvägen" {
t.Errorf("Fake(%q) = %q, want the named path to run", want, got)
}
}
if strings.Contains(err.Error(), "[se.locality") {
t.Errorf("%v lists the rows as a Go slice; separate them with commas", err)
}
if _, err := f.Fake("se.locality[Sandby]"); err == nil || !strings.Contains(err.Error(), "one of L7, L8") {
t.Fatalf("se.locality[Sandby] = %v, want its keys listed plainly", err)
}
}
+25 -7
View File
@@ -88,17 +88,35 @@ func funcCall(body string) (name string, args []string, ok bool) {
return body[:lp], splitArgs(body[lp+1 : len(body)-1]), true
}
// splitArgs parses a function arg list: comma-separated outside a selector,
// trimmed; empty -> none.
// splitArgs parses a function arg list: comma-separated outside a selector or a
// quoted layout, trimmed; empty -> none.
func splitArgs(s string) []string {
if strings.TrimSpace(s) == "" {
return nil
}
args := splitOutside(s, ',')
for i := range args {
args[i] = strings.TrimSpace(args[i])
var args []string
depth, quoted, start := 0, false, 0
for i := 0; i < len(s); i++ {
switch c := s[i]; {
case c == '\'' && depth == 0:
quoted = !quoted
case quoted:
case c == '[':
depth++
case c == ']' && depth > 0:
depth--
case c == ',' && depth == 0:
args, start = append(args, strings.TrimSpace(s[start:i])), i+1
}
return args
}
return append(args, strings.TrimSpace(s[start:]))
}
func plural(n int) string {
if n == 1 {
return ""
}
return "s"
}
// checkFunc validates a function token at compile time: well-formed, naming a
@@ -114,7 +132,7 @@ func checkFunc(body string, fields map[string]node) error {
return fmt.Errorf("token {%s}: unknown function %q", body, name)
}
if b.arity >= 0 && len(args) != b.arity {
return fmt.Errorf("token {%s}: %s takes %d args, got %d", body, name, b.arity, len(args))
return fmt.Errorf("token {%s}: %s takes %d argument%s, got %d", body, name, b.arity, plural(b.arity), len(args))
}
if b.check != nil {
if err := b.check(fields, args); err != nil {
+16
View File
@@ -2,6 +2,7 @@ package fejkdata
import (
"encoding/json"
"reflect"
"regexp"
"strings"
"testing"
@@ -379,3 +380,18 @@ func TestArgErrorsNameTheSpelling(t *testing.T) {
}
}
}
// TestSplitArgsQuotesOutsideSelectors pins the arg grammar: a comma splits outside a
// selector and outside a quoted layout, and a quote inside a selector is a name's text.
func TestSplitArgsQuotesOutsideSelectors(t *testing.T) {
for in, want := range map[string][]string{
"a, b": {"a", "b"},
"1990-01-01,1990-12-31,'January 2, 2006'": {"1990-01-01", "1990-12-31", "'January 2, 2006'"},
"/geo.US.locality[O'Fallon].name, 2": {"/geo.US.locality[O'Fallon].name", "2"},
"'[a,b]'": {"'[a,b]'"},
} {
if got := splitArgs(in); !reflect.DeepEqual(got, want) {
t.Errorf("splitArgs(%q) = %q, want %q", in, got, want)
}
}
}
+68 -31
View File
@@ -8,20 +8,28 @@ en_US.color
en_US.company format "{base} {suffix}"
en_US.company.base string
en_US.company.suffix string
en_US.date format "{month}/{day}/{year}"
en_US.date.day string
en_US.date.month string
en_US.date.year string
en_US.date format "{date(1970-01-01,2029-12-31,'01/02/2006')}"
en_US.email format "{local}@{domain}"
en_US.email.domain string
en_US.email.local string
en_US.email.local.n
en_US.first-name format "{name}" name name weight count parent sex
en_US.first-name.count string
en_US.first-name.name string
en_US.first-name.sex string
en_US.ip
en_US.person format "{prefix}{femalefirst|malefirst} {last}"
en_US.person.femalefirst string
en_US.itin format "{area}-{group}-{serial}"
en_US.itin.area string
en_US.itin.group string
en_US.itin.serial string
en_US.last-name format "{name}" key name weight count
en_US.last-name.count string
en_US.last-name.name string
en_US.person format "{prefix}{first} {last}" reads en_US.first-name en_US.last-name en_US.sex en_US.title
en_US.person.first string
en_US.person.last string
en_US.person.malefirst string
en_US.person.prefix string
en_US.person.prefix string null
en_US.person.sex string
en_US.phone
en_US.phone.area
en_US.phone.exch
@@ -34,12 +42,26 @@ en_US.sentence.adj
en_US.sentence.noun
en_US.sentence.prep
en_US.sentence.verb
en_US.ssn format "{int(100,999)}-{digits(2)}-{digits(4)}"
en_US.time format "{hour}:{minute} {ampm}"
en_US.time.ampm string
en_US.time.hour string
en_US.time.minute string
en_US.time.minute.t
en_US.sex format "{name}" key code name name
en_US.sex.code string
en_US.sex.first-name
en_US.sex.first-name.count
en_US.sex.first-name.name
en_US.sex.first-name.sex
en_US.sex.name string
en_US.sex.title
en_US.sex.title.name
en_US.sex.title.sex
en_US.sex.title.share
en_US.ssn format "{area}-{group}-{serial}"
en_US.ssn.area string
en_US.ssn.group string
en_US.ssn.serial string
en_US.time format "{time('3:04 PM')}"
en_US.title format "{name}" name name weight share parent sex
en_US.title.name string
en_US.title.sex string
en_US.title.share string
en_US.url format "https://{host}{path}"
en_US.url.host string
en_US.url.path string
@@ -216,6 +238,7 @@ misc.currency.decimals string
misc.currency.name string
misc.currency.numeric string
misc.currency.symbol string
misc.datetime format "{date(2000-01-01,2029-12-31,'2006-01-02T15:04:05Z')}"
misc.emoji
misc.httpstatus format "{code} {reason}" key code name reason
misc.httpstatus.code string
@@ -237,24 +260,32 @@ sv_SE.address.locality string
sv_SE.address.postal-code string
sv_SE.address.street string
sv_SE.address.street-number string
sv_SE.birth-number format "{number}" parent sex
sv_SE.birth-number.number string
sv_SE.birth-number.sex string
sv_SE.color
sv_SE.company format "{base} {suffix}"
sv_SE.company.base string
sv_SE.company.suffix string
sv_SE.date format "{year}-{month}-{day}"
sv_SE.date.day string
sv_SE.date.month string
sv_SE.date.year string
sv_SE.date format "{date(1970-01-01,2029-12-31,'2006-01-02')}"
sv_SE.email format "{local}@{domain}"
sv_SE.email.domain string
sv_SE.email.local string
sv_SE.email.local.n
sv_SE.first-name format "{name}" name name weight count parent sex
sv_SE.first-name.count string
sv_SE.first-name.name string
sv_SE.first-name.sex string
sv_SE.ip
sv_SE.person format "{prefix}{femalefirst|malefirst} {last}"
sv_SE.person.femalefirst string
sv_SE.last-name format "{name}" key name weight count
sv_SE.last-name.count string
sv_SE.last-name.name string
sv_SE.person format "{prefix}{first} {last}" reads sv_SE.first-name sv_SE.last-name sv_SE.sex sv_SE.title
sv_SE.person.first string
sv_SE.person.last string
sv_SE.person.malefirst string
sv_SE.person.prefix string
sv_SE.person.prefix string null
sv_SE.person.sex string
sv_SE.personnummer format "{date(1930-01-01,2025-12-31,'060102')}-{.birth-number.number}{luhn()}" reads sv_SE.birth-number
sv_SE.phone
sv_SE.phone.a
sv_SE.phone.b
@@ -263,20 +294,26 @@ sv_SE.phone.prefix
sv_SE.price format "{amt}{ore} kr"
sv_SE.price.amt string
sv_SE.price.ore string
sv_SE.samordningsnummer format "{date(1930-01-01,2025-12-31,'0601')}{int(61,88)}-{.birth-number.number}{luhn()}" reads sv_SE.birth-number
sv_SE.sentence
sv_SE.sentence.adj
sv_SE.sentence.noun
sv_SE.sentence.prep
sv_SE.sentence.verb
sv_SE.ssn format "{digits(2)}{mmdd}-{digits(3)}{luhn()}"
sv_SE.ssn.mmdd string
sv_SE.ssn.mmdd.d
sv_SE.ssn.mmdd.m
sv_SE.time format "{hour}:{minute}{sec}"
sv_SE.time.hour string
sv_SE.time.minute string
sv_SE.time.minute.t
sv_SE.time.sec string
sv_SE.sex format "{name}" key code name name
sv_SE.sex.birth-number
sv_SE.sex.birth-number.number
sv_SE.sex.birth-number.sex
sv_SE.sex.code string
sv_SE.sex.first-name
sv_SE.sex.first-name.count
sv_SE.sex.first-name.name
sv_SE.sex.first-name.sex
sv_SE.sex.name string
sv_SE.time format "{time('15:04')}"
sv_SE.title format "{name}" key name weight share
sv_SE.title.name string
sv_SE.title.share string
sv_SE.url format "https://{host}{path}"
sv_SE.url.host string
sv_SE.url.path string
+11 -6
View File
@@ -42,7 +42,7 @@ value is composed; TSV says which values exist.
and a `data-import/` directory of Python scripts that rebuild each TSV from its
source, so a refresh is one command per dataset.
- Make every shipped category a record with its building blocks as columns
(`sv_SE.person` → `femalefirst`, `malefirst`; `misc.uuid` → `variant`).
(`sv_SE.person` → `first`, `last`, `sex`, done in step 3; `misc.uuid` → `variant`).
- Share the handle lists between `email.local` and `username` only if that is a
clean win — a reference between shipped categories is a major once tagged.
@@ -75,9 +75,8 @@ countries; the README maps each to the native term.
#### Builtins the data cannot express
- `{date(from,to,layout)}`: a date in `[from, to]` in a Go layout, and `{age(min,max)}`
as the birthdate spelling; `{time(layout)}`. Personnummer, SSN, birthdate, card
expiry, unix time and ISO datetime all build on it.
- `{date(from,to,'layout')}` and `{time('layout')}` shipped in step 3; `age()` is
rejected, README Decisions. Add a `unix` layout once something needs it.
- Derivations: `{isin()}` (Luhn over letters expanded to digits), `{cusip()}`,
`{aba()}` (3-7-1 weights), `{vin()}` (position 9 over the whole; a sample taking
the WMI, since the check sits mid-string).
@@ -98,7 +97,7 @@ ids. Shape: T = table, t = template, c = choice.
| Category | Shape | Source | Licence |
|---|---|---|---|
| `person` first (female, male, generic), middle, last, weighted | T | SCB 2022 whole-population xlsx, Skatteverket 2026 surnames | CC0, "Källa: SCB" |
| `person` title, gender, birthdate, age, blood type weighted | t | geblod.nu distribution | facts |
| `person` title, sex, birthdate, age, blood type weighted | t | geblod.nu distribution | facts |
| `personnummer`, `samordningsnummer` | t | Skatteverket test series: date + 238/239, Luhn | CC0 |
| `organisationsnummer` by form, `vat` | t | Bolagsverket group digits, Luhn, `SE…01` | facts |
| `company` name patterns, legal form weighted | t | Bolagsverket registrations 2025 | CC BY 2.5 SE |
@@ -165,7 +164,13 @@ address, phone, national id, company and date names each.
1. Table node, key and name selection, parent links, consistent draws, the
choice-of-rows fence, `DATA-LICENSES.md`, `data-import/` — done.
2. `geo/SE` and `geo/US`, and `address` in both locales on top of them — done.
3. Weighted person names and valid ids in both locales; `date()`.
3. Weighted person names and valid ids in both locales; `date()` — done. Left for
later: middle names, and drawing a shipped `personnummer` inside a *selected*
sex, both of which want a draw group that shares its family's pins; stop the
conflict error naming a rewrite that returns a different value where the read it
conflicts with sits inside another category. Give the national ids one
record shape in step 6, and report a struct column's draw conflict with the path
spelling a tag takes rather than the reference spelling.
4. `misc` conversions and the new `misc` tables.
5. The remaining locale categories: company, phone, finance, vehicle, words.
6. Records with building-block columns across the shipped set; shape re-pin.
+94
View File
@@ -0,0 +1,94 @@
package fejkdata
import (
"fmt"
"strings"
"unicode"
)
// transforms are the builtins that rewrite one operand's value; they nest, so
// {lowercase(ascii(x))} folds then lowers.
var transforms = map[string]func(string) string{
"ascii": asciiFold,
"lowercase": strings.ToLower,
"uppercase": strings.ToUpper,
}
// unwrapTransform peels nested transform calls off an operand arg, returning the
// field it finally names and the transforms to apply, innermost last.
func unwrapTransform(arg string) (leaf string, chain []func(string) string, err error) {
for {
name, args, isCall := funcCall(arg)
if !isCall {
return arg, chain, nil
}
fn, isTransform := transforms[name]
if !isTransform {
return "", nil, fmt.Errorf("%s(%s) is not a transform, so it cannot be an operand", name, strings.Join(args, ","))
}
if len(args) != 1 {
return "", nil, fmt.Errorf("%s takes 1 arg, got %d", name, len(args))
}
chain = append(chain, fn)
arg = args[0]
}
}
func transformArg(fields map[string]node, a []string) error {
leaf, _, err := unwrapTransform(a[0])
if err != nil {
return err
}
if isRef(leaf) {
_, _, err := refShape(leaf)
return err
}
return checkArm(leaf, fields, false)
}
func transformOperand(a []string) []string {
leaf, _, err := unwrapTransform(a[0])
if err != nil {
return nil
}
return []string{leaf}
}
func transformPrep(outer func(string) string) func([]string) callFn {
return func(a []string) callFn {
_, chain, err := unwrapTransform(a[0])
if err != nil {
panic(fmt.Sprintf("fejkdata: transform arg %q reached prep unvalidated: %v", a[0], err))
}
return func(_ *session, _ string, operands []string) string {
v := operands[0]
for i := len(chain) - 1; i >= 0; i-- {
v = chain[i](v)
}
return outer(v)
}
}
}
// asciiFolds maps the Latin letters with diacritics or ligatures to ASCII.
var asciiFolds = map[rune]string{
'À': "A", 'Á': "A", 'Â': "A", 'Ã': "A", 'Ä': "A", 'Å': "A", 'Æ': "AE", 'Ç': "C",
'È': "E", 'É': "E", 'Ê': "E", 'Ë': "E", 'Ì': "I", 'Í': "I", 'Î': "I", 'Ï': "I",
'Ð': "D", 'Ñ': "N", 'Ò': "O", 'Ó': "O", 'Ô': "O", 'Õ': "O", 'Ö': "O", 'Ø': "O",
'Ù': "U", 'Ú': "U", 'Û': "U", 'Ü': "U", 'Ý': "Y", 'Þ': "Th", 'ß': "ss", 'Œ': "OE",
'à': "a", 'á': "a", 'â': "a", 'ã': "a", 'ä': "a", 'å': "a", 'æ': "ae", 'ç': "c",
'è': "e", 'é': "e", 'ê': "e", 'ë': "e", 'ì': "i", 'í': "i", 'î': "i", 'ï': "i",
'ð': "d", 'ñ': "n", 'ò': "o", 'ó': "o", 'ô': "o", 'õ': "o", 'ö': "o", 'ø': "o",
'ù': "u", 'ú': "u", 'û': "u", 'ü': "u", 'ý': "y", 'þ': "th", 'ÿ': "y", 'œ': "oe",
}
// asciiFold rewrites s to ASCII: folded Latin letters stay, any other non-ASCII
// rune is dropped.
func asciiFold(s string) string {
var b strings.Builder
for _, r := range s {
if r <= unicode.MaxASCII {
b.WriteRune(r)
} else {
b.WriteString(asciiFolds[r])
}
}
return b.String()
}