Weighted person names, valid personal ids and date() in both locales #21

Merged
lilleman merged 38 commits from person-ids-date into main 2026-09-18 20:45:51 +02:00
19 changed files with 102 additions and 80 deletions
Showing only changes of commit 02eb95b94b - Show all commits
+3 -2
View File
@@ -43,5 +43,6 @@ replacement, and each removed path, column or flag.
`first`, `last`, `prefix` and `sex`, and `person.femalefirst` and `first`, `last`, `prefix` and `sex`, and `person.femalefirst` and
`person.malefirst` are no longer paths. `person.malefirst` are no longer paths.
- `sv_SE.personnummer` and `sv_SE.samordningsnummer`, Skatteverket's test series - `sv_SE.personnummer` and `sv_SE.samordningsnummer`, Skatteverket's test series
over the sex the render drew, in place of `sv_SE.ssn`; `en_US.ssn` in the ranges from a `sv_SE.birth-number` table under `sex`, in place of `sv_SE.ssn`; `en_US.ssn`
the SSA assigns, and `en_US.itin`. in the ranges the SSA assigns, and `en_US.itin`. `en_US.person.prefix` no longer
draws `Mr`, `Mrs`, `Ms` or `Miss`, which could contradict the record's `sex`.
+1 -1
View File
@@ -14,7 +14,7 @@ Every shipped dataset, its source, its licence and the attribution it asks for.
| `sv_SE/first-name.tsv`, `last-name.tsv` | [SCB](https://www.scb.se/) names with at least two bearers, 31 December 2022 | CC0 1.0 | "Källa: SCB" | `data-import/names-se.py` | | `sv_SE/first-name.tsv`, `last-name.tsv` | [SCB](https://www.scb.se/) names with at least two bearers, 31 December 2022 | CC0 1.0 | "Källa: SCB" | `data-import/names-se.py` |
| `en_US/first-name.tsv` | [SSA](https://www.ssa.gov/oact/babynames/) baby names, births 1930 to 2020, through [hackerb9/ssa-baby-names](https://github.com/hackerb9/ssa-baby-names) | public domain | none required | `data-import/names-us.py` | | `en_US/first-name.tsv` | [SSA](https://www.ssa.gov/oact/babynames/) baby names, births 1930 to 2020, through [hackerb9/ssa-baby-names](https://github.com/hackerb9/ssa-baby-names) | public domain | none required | `data-import/names-us.py` |
| `en_US/last-name.tsv` | Census Bureau surnames occurring 100 or more times, 2010 | public domain | none required | `data-import/names-us.py` | | `en_US/last-name.tsv` | Census Bureau surnames occurring 100 or more times, 2010 | public domain | none required | `data-import/names-us.py` |
| `sv_SE/sex.tsv`, `en_US/sex.tsv` | curated (Skatteverket's test birth numbers are facts) | — | — | — | | `sv_SE/sex.tsv`, `sv_SE/birth-number.tsv`, `en_US/sex.tsv` | curated (Skatteverket's test birth numbers are facts) | — | — | — |
| `misc/country.tsv` | [datasets/country-codes](https://github.com/datasets/country-codes) | [PDDL 1.0](https://opendatacommons.org/licenses/pddl/1-0/) | none required | `data-import/country.py` | | `misc/country.tsv` | [datasets/country-codes](https://github.com/datasets/country-codes) | [PDDL 1.0](https://opendatacommons.org/licenses/pddl/1-0/) | none required | `data-import/country.py` |
| `misc/currency.tsv` | [datasets/currency-codes](https://github.com/datasets/currency-codes); symbols from [Unicode CLDR](https://github.com/unicode-org/cldr) `en.xml` and `root.xml` | PDDL 1.0; [Unicode License v3](https://www.unicode.org/license.txt) | CLDR: "Copyright © 1991-2025 Unicode, Inc. Unicode and the Unicode Logo are registered trademarks of Unicode, Inc. in the United States and other countries." | `data-import/currency.py` | | `misc/currency.tsv` | [datasets/currency-codes](https://github.com/datasets/currency-codes); symbols from [Unicode CLDR](https://github.com/unicode-org/cldr) `en.xml` and `root.xml` | PDDL 1.0; [Unicode License v3](https://www.unicode.org/license.txt) | CLDR: "Copyright © 1991-2025 Unicode, Inc. Unicode and the Unicode Logo are registered trademarks of Unicode, Inc. in the United States and other countries." | `data-import/currency.py` |
| `misc/httpstatus.tsv` | curated (IANA HTTP status codes are facts) | — | — | — | | `misc/httpstatus.tsv` | curated (IANA HTTP status codes are facts) | — | — | — |
+11 -7
View File
@@ -235,8 +235,9 @@ SSA and the Census Bureau. `first-name` links to `sex`, so `sv_SE.sex[f].first-n
draws a woman's name, and a name both sexes carry is a row under each, so draws a woman's name, and a name both sexes carry is a row under each, so
`en_US.sex[m].first-name[Taylor]` names the one a `first-name[Taylor]` alone cannot. `en_US.sex[m].first-name[Taylor]` names the one a `first-name[Taylor]` alone cannot.
`person` reads one draw of the three, so its `first` and `sex` columns agree, and so `person` reads one draw of the three, so its `first` and `sex` columns agree, and so
does a `personnummer` in the same render: its birth number is Skatteverket's test does a `personnummer` in the same render: its birth number, `sv_SE.birth-number`
series, 238 for a woman and 239 for a man, which no real person is ever given. under `sex`, is Skatteverket's test series, 238 for a woman and 239 for a man, which no
real person is ever given.
A `geo` folder holds one tree per country under its alpha-2 code: five A `geo` folder holds one tree per country under its alpha-2 code: five
[linked tables](#linked-tables) named alike, and an `address` record over one [linked tables](#linked-tables) named alike, and an `address` record over one
@@ -536,8 +537,9 @@ Renders e.g. `811218-2389`. A layout is Go's: the reference time `Mon Jan 2
15:04:05 MST 2006` spelled as the output should look, quoted, since a layout may 15:04:05 MST 2006` spelled as the output should look, quoted, since a layout may
carry the comma that separates arguments, with English names. Every second carry the comma that separates arguments, with English names. Every second
between the two days is reachable, so a layout with a clock draws the time too. between the two days is reachable, so a layout with a clock draws the time too.
Rejected at `New`: a bound that is no calendar date, or not before the other; an The quotes delimit a layout outside a selector only, so `[O'Fallon]` in an
unquoted layout, naming the quoted one; a layout naming no field, which is text; argument stays a name. Rejected at `New`: a bound that is no calendar date, or not
before the other; an unquoted layout, naming the quoted one; a layout naming no field, which is text;
and for `time` a layout naming a date field, naming `date`. `{seq()}` spans `Fake` and for `time` a layout naming a date field, naming `date`. `{seq()}` spans `Fake`
calls and `repeat`, resets with a new generator, and is the natural primary key for calls and `repeat`, resets with a new generator, and is the natural primary key for
the SQL example above. the SQL example above.
@@ -1077,9 +1079,11 @@ renamed or retyped line is a major.
load, since nothing could then select it. load, since nothing could then select it.
- **The Swedish ids draw Skatteverket's test series.** A Luhn-valid personnummer - **The Swedish ids draw Skatteverket's test series.** A Luhn-valid personnummer
over a random birth number may be a living person's; 238 and 239 after any date over a random birth number may be a living person's; 238 and 239 after any date
are blocked from assignment, so the shipped `personnummer` and are blocked from assignment, so the shipped `personnummer` and
`samordningsnummer` use those, read from the `sex` table's `birth-number` `samordningsnummer` use those. They sit in a `birth-number` table under `sex`
column so the number and the name agree on sex. rather than as a column of it: the render's shared draw of the family is what
makes the number and the name agree on sex, and `sex` stays one shape across
locales instead of collecting every sex-keyed id fact.
- **The US given names come from a mirror of the SSA file.** ssa.gov refuses a - **The US given names come from a mirror of the SSA file.** ssa.gov refuses a
client outside the US, so `names-us.py` reads a GitHub copy that ends at 2020, client outside the US, so `names-us.py` reads a GitHub copy that ends at 2020,
which a count over the births since 1930 barely feels; `--names` takes the which a count over the births since 1930 barely feels; `--names` takes the
+2 -1
View File
@@ -9,6 +9,7 @@ import io
import re import re
from pathlib import Path from pathlib import Path
import source
import tsv import tsv
SOURCE = "https://raw.githubusercontent.com/datasets/country-codes/main/data/country-codes.csv" SOURCE = "https://raw.githubusercontent.com/datasets/country-codes/main/data/country-codes.csv"
@@ -58,7 +59,7 @@ def main():
p.add_argument("--source", default=SOURCE) p.add_argument("--source", default=SOURCE)
p.add_argument("--out", default=str(OUT)) p.add_argument("--out", default=str(OUT))
a = p.parse_args() a = p.parse_args()
table = rows(tsv.fetch(a.source, a.cache, "country-codes.csv").decode("utf-8")) table = rows(source.fetch(a.source, a.cache, "country-codes.csv").decode("utf-8"))
tsv.write(a.out, COLUMNS, sorted(table, key=lambda r: r["alpha2"])) tsv.write(a.out, COLUMNS, sorted(table, key=lambda r: r["alpha2"]))
+3 -2
View File
@@ -11,6 +11,7 @@ import io
import xml.etree.ElementTree as ET import xml.etree.ElementTree as ET
from pathlib import Path from pathlib import Path
import source
import tsv import tsv
SOURCE = "https://raw.githubusercontent.com/datasets/currency-codes/main/data/codes-all.csv" SOURCE = "https://raw.githubusercontent.com/datasets/currency-codes/main/data/codes-all.csv"
@@ -58,8 +59,8 @@ def main():
p.add_argument("--symbols", nargs="+", default=SYMBOLS) p.add_argument("--symbols", nargs="+", default=SYMBOLS)
p.add_argument("--out", default=str(OUT)) p.add_argument("--out", default=str(OUT))
a = p.parse_args() a = p.parse_args()
symbol = symbols(tsv.fetch(s, a.cache, Path(s).name).decode("utf-8") for s in a.symbols) symbol = symbols(source.fetch(s, a.cache, Path(s).name).decode("utf-8") for s in a.symbols)
table = rows(tsv.fetch(a.source, a.cache, "codes-all.csv").decode("utf-8"), symbol) table = rows(source.fetch(a.source, a.cache, "codes-all.csv").decode("utf-8"), symbol)
tsv.write(a.out, COLUMNS, sorted(table, key=lambda r: r["code"])) tsv.write(a.out, COLUMNS, sorted(table, key=lambda r: r["code"]))
+6 -4
View File
@@ -16,6 +16,8 @@ import urllib.request
import zipfile import zipfile
from pathlib import Path from pathlib import Path
import source
import xlsx
import tsv import tsv
CODES = "https://www.scb.se/contentassets/7a89e48960f741e08918e489ea36354a/kommunlankod-2026.xlsx" CODES = "https://www.scb.se/contentassets/7a89e48960f741e08918e489ea36354a/kommunlankod-2026.xlsx"
@@ -39,7 +41,7 @@ ONE_POSITION = {"Stockholm", "Göteborg", "Malmö"}
UNMATCHED_POPULATION = 200 UNMATCHED_POPULATION = 200
def scb_codes(cache): def scb_codes(cache):
regions, municipalities = {}, {} regions, municipalities = {}, {}
for cells in tsv.xlsx_rows(tsv.fetch(CODES, cache, "kommunlankod.xlsx", magic=b"PK")): for cells in xlsx.rows(source.fetch(CODES, cache, "kommunlankod.xlsx", magic=b"PK")):
if len(cells) < 2 or not re.fullmatch(r"\d{2}|\d{4}", cells[0]): if len(cells) < 2 or not re.fullmatch(r"\d{2}|\d{4}", cells[0]):
continue continue
(regions if len(cells[0]) == 2 else municipalities)[cells[0]] = cells[1].strip() (regions if len(cells[0]) == 2 else municipalities)[cells[0]] = cells[1].strip()
@@ -48,12 +50,12 @@ def scb_codes(cache):
def scb_population(cache): def scb_population(cache):
body = json.dumps(POPULATION_QUERY).encode() body = json.dumps(POPULATION_QUERY).encode()
data = tsv.fetch(POPULATION, cache, "befolkning.json", data=body, headers={"Content-Type": "application/json"}) data = source.fetch(POPULATION, cache, "befolkning.json", data=body, headers={"Content-Type": "application/json"})
return {row["key"][0]: row["values"][0] for row in json.loads(data.decode("utf-8-sig"))["data"]} return {row["key"][0]: row["values"][0] for row in json.loads(data.decode("utf-8-sig"))["data"]}
def scb_tatorter(cache): def scb_tatorter(cache):
text = tsv.fetch(TATORTER, cache, "tatorter.csv").decode("utf-8") text = source.fetch(TATORTER, cache, "tatorter.csv").decode("utf-8")
by_name = collections.defaultdict(list) by_name = collections.defaultdict(list)
for r in csv.DictReader(io.StringIO(text)): for r in csv.DictReader(io.StringIO(text)):
by_name[r["tatort"]].append((r["kommun"], int(r["bef"]))) by_name[r["tatort"]].append((r["kommun"], int(r["bef"])))
@@ -61,7 +63,7 @@ def scb_tatorter(cache):
def geonames(cache): def geonames(cache):
z = zipfile.ZipFile(io.BytesIO(tsv.fetch(POSTAL_CODES, cache, "SE.zip", magic=b"PK"))) z = zipfile.ZipFile(io.BytesIO(source.fetch(POSTAL_CODES, cache, "SE.zip", magic=b"PK")))
rows = [] rows = []
for line in z.read("SE.txt").decode("utf-8").splitlines(): for line in z.read("SE.txt").decode("utf-8").splitlines():
f = line.split("\t") f = line.split("\t")
+5 -4
View File
@@ -14,6 +14,7 @@ import sys
import zipfile import zipfile
from pathlib import Path from pathlib import Path
import source
import tsv import tsv
GAZETTEER = "https://www2.census.gov/geo/docs/maps-data/data/gazetteer/2026_Gazetteer/2026_Gaz_{}_national.zip" GAZETTEER = "https://www2.census.gov/geo/docs/maps-data/data/gazetteer/2026_Gazetteer/2026_Gaz_{}_national.zip"
@@ -56,14 +57,14 @@ def text(data):
def gazetteer(cache, kind): def gazetteer(cache, kind):
z = zipfile.ZipFile(io.BytesIO(tsv.fetch(GAZETTEER.format(kind), cache, f"gaz_{kind}.zip", magic=b"PK"))) z = zipfile.ZipFile(io.BytesIO(source.fetch(GAZETTEER.format(kind), cache, f"gaz_{kind}.zip", magic=b"PK")))
rows = text(z.read(z.namelist()[0])).splitlines() rows = text(z.read(z.namelist()[0])).splitlines()
header = [h.strip() for h in rows[0].split("|")] header = [h.strip() for h in rows[0].split("|")]
return [dict(zip(header, (c.strip() for c in row.split("|")))) for row in rows[1:]] return [dict(zip(header, (c.strip() for c in row.split("|")))) for row in rows[1:]]
def csv_rows(cache, url, name): def csv_rows(cache, url, name):
return list(csv.DictReader(io.StringIO(text(tsv.fetch(url, cache, name))))) return list(csv.DictReader(io.StringIO(text(source.fetch(url, cache, name)))))
def dbf_rows(data, wanted): def dbf_rows(data, wanted):
@@ -89,7 +90,7 @@ def dbf_rows(data, wanted):
def tiger_zip(cache, kind, county): def tiger_zip(cache, kind, county):
return tsv.fetch(TIGER.format(kind.upper(), county, kind), cache, f"tl_{county}_{kind}.zip", magic=b"PK") return source.fetch(TIGER.format(kind.upper(), county, kind), cache, f"tl_{county}_{kind}.zip", magic=b"PK")
def tiger(cache, kind, county, wanted): def tiger(cache, kind, county, wanted):
@@ -130,7 +131,7 @@ def localities(cache, min_population, counties):
def postal_codes(cache, localities): def postal_codes(cache, localities):
"""Each ZCTA whose largest part inside an incorporated place lies in a shipped place.""" """Each ZCTA whose largest part inside an incorporated place lies in a shipped place."""
parts = {} parts = {}
for r in csv.DictReader(io.StringIO(text(tsv.fetch(ZCTA_PLACE, cache, "zcta-place.txt"))), delimiter="|"): for r in csv.DictReader(io.StringIO(text(source.fetch(ZCTA_PLACE, cache, "zcta-place.txt"))), delimiter="|"):
if r["GEOID_ZCTA5_20"] and r["GEOID_PLACE_20"] and not r["NAMELSAD_PLACE_20"].endswith(" CDP"): if r["GEOID_ZCTA5_20"] and r["GEOID_PLACE_20"] and not r["NAMELSAD_PLACE_20"].endswith(" CDP"):
parts.setdefault(r["GEOID_ZCTA5_20"], []).append((int(r["AREALAND_PART"]), r["GEOID_PLACE_20"])) parts.setdefault(r["GEOID_ZCTA5_20"], []).append((int(r["AREALAND_PART"]), r["GEOID_PLACE_20"]))
largest = {zcta: max(p)[1] for zcta, p in parts.items()} largest = {zcta: max(p)[1] for zcta, p in parts.items()}
+5 -3
View File
@@ -9,6 +9,8 @@ import argparse
import re import re
from pathlib import Path from pathlib import Path
import source
import xlsx
import tsv import tsv
SOURCE = "https://www.scb.se/contentassets/9fe7dbb460994c72b835163dbc491ef9/namn-med-minst-tva-barare-31-december-2022.xlsx" SOURCE = "https://www.scb.se/contentassets/9fe7dbb460994c72b835163dbc491ef9/namn-med-minst-tva-barare-31-december-2022.xlsx"
@@ -40,9 +42,9 @@ def main():
p.add_argument("--out", default=str(OUT)) p.add_argument("--out", default=str(OUT))
p.add_argument("--source", default=SOURCE) p.add_argument("--source", default=SOURCE)
a = p.parse_args() a = p.parse_args()
data = tsv.fetch(a.source, a.cache, "scb-namn-2022.xlsx", magic=b"PK") data = source.fetch(a.source, a.cache, "scb-namn-2022.xlsx", magic=b"PK")
first = [{"name": cased(n), "sex": sex, "count": c} for sex, sheet in SHEETS.items() for n, c in counted(tsv.xlsx_rows(data, sheet), a.first)] first = [{"name": cased(n), "sex": sex, "count": c} for sex, sheet in SHEETS.items() for n, c in counted(xlsx.rows(data, sheet), a.first)]
last = [{"name": cased(n), "count": c} for n, c in counted(tsv.xlsx_rows(data, SURNAMES), a.last)] last = [{"name": cased(n), "count": c} for n, c in counted(xlsx.rows(data, SURNAMES), a.last)]
out = Path(a.out) out = Path(a.out)
tsv.write(out / "first-name.tsv", ["name", "sex", "count"], sorted(first, key=lambda r: (r["name"], r["sex"]))) tsv.write(out / "first-name.tsv", ["name", "sex", "count"], sorted(first, key=lambda r: (r["name"], r["sex"])))
tsv.write(out / "last-name.tsv", ["name", "count"], sorted(last, key=lambda r: r["name"])) tsv.write(out / "last-name.tsv", ["name", "count"], sorted(last, key=lambda r: r["name"]))
+5 -4
View File
@@ -15,6 +15,7 @@ import re
import zipfile import zipfile
from pathlib import Path from pathlib import Path
import source
import tsv import tsv
NAMES = "https://raw.githubusercontent.com/hackerb9/ssa-baby-names/master/alldata.txt" NAMES = "https://raw.githubusercontent.com/hackerb9/ssa-baby-names/master/alldata.txt"
@@ -24,7 +25,7 @@ CACHE = Path(__file__).resolve().parent / "cache"
MC = re.compile(r"^Mc([a-z])") MC = re.compile(r"^Mc([a-z])")
def csv_rows(data, member): def csv_or_zip_rows(data, member):
"""The rows of a CSV, or of every member of a zip named like member, name,sex,count[,year].""" """The rows of a CSV, or of every member of a zip named like member, name,sex,count[,year]."""
if data.startswith(b"PK"): if data.startswith(b"PK"):
z = zipfile.ZipFile(io.BytesIO(data)) z = zipfile.ZipFile(io.BytesIO(data))
@@ -39,7 +40,7 @@ def csv_rows(data, member):
def given(data, from_year): def given(data, from_year):
counts = collections.Counter() counts = collections.Counter()
for r in csv_rows(data, r"yob\d{4}\.txt"): for r in csv_or_zip_rows(data, r"yob\d{4}\.txt"):
if len(r) >= 4 and r[3].isdigit() and int(r[3]) >= from_year: if len(r) >= 4 and r[3].isdigit() and int(r[3]) >= from_year:
counts[(r[0], r[1].lower())] += int(r[2]) counts[(r[0], r[1].lower())] += int(r[2])
return counts return counts
@@ -60,12 +61,12 @@ def main():
p.add_argument("--out", default=str(OUT)) p.add_argument("--out", default=str(OUT))
p.add_argument("--surnames", default=SURNAMES) p.add_argument("--surnames", default=SURNAMES)
a = p.parse_args() a = p.parse_args()
counts = given(tsv.fetch(a.names, a.cache, "ssa-names.txt"), a.from_year) counts = given(source.fetch(a.names, a.cache, "ssa-names.txt"), a.from_year)
first = [] first = []
for sex in ("f", "m"): for sex in ("f", "m"):
top = sorted(((n, c) for (n, s), c in counts.items() if s == sex), key=lambda n: (-n[1], n[0]))[:a.first] top = sorted(((n, c) for (n, s), c in counts.items() if s == sex), key=lambda n: (-n[1], n[0]))[:a.first]
first += [{"name": n, "sex": sex, "count": c} for n, c in top] first += [{"name": n, "sex": sex, "count": c} for n, c in top]
rows = csv_rows(tsv.fetch(a.surnames, a.cache, "census-surnames-2010.zip"), r"Names_2010Census\.csv") rows = csv_or_zip_rows(source.fetch(a.surnames, a.cache, "census-surnames-2010.zip"), r"Names_2010Census\.csv")
last = [{"name": surname(r[0]), "count": int(r[2])} for r in rows if len(r) >= 3 and r[2].isdigit() and r[0].isalpha()] last = [{"name": surname(r[0]), "count": int(r[2])} for r in rows if len(r) >= 3 and r[2].isdigit() and r[0].isalpha()]
last = sorted(last, key=lambda r: (-r["count"], r["name"]))[:a.last] last = sorted(last, key=lambda r: (-r["count"], r["name"]))[:a.last]
out = Path(a.out) out = Path(a.out)
+28
View File
@@ -0,0 +1,28 @@
"""A source fetched once into the cache."""
import re
import sys
import time
import urllib.request
from pathlib import Path
def fetch(source, cache, name, magic=b"", data=None, headers=None):
"""The bytes of a URL, downloaded into cache/name once, or of a local file."""
if not re.match(r"^https?://", source):
return Path(source).read_bytes()
path = Path(cache) / name
for attempt in range(1, 6):
if path.exists():
return path.read_bytes()
req = urllib.request.Request(source, data=data, headers={"User-Agent": "fejkdata data-import", **(headers or {})})
try:
with urllib.request.urlopen(req, timeout=600) as r:
body = r.read()
except OSError:
body = b""
if body and body.startswith(magic) and b"Request Rejected" not in body[:512]:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_bytes(body)
elif attempt < 5:
time.sleep(10 * attempt)
sys.exit(f"{source}: no valid download in 5 attempts")
+1 -45
View File
@@ -1,52 +1,8 @@
"""A source fetched once into the cache, and a table written as the loader admits it.""" """A table written as the loader admits it."""
import io
import re import re
import sys import sys
import time
import urllib.request
import xml.etree.ElementTree as ET
import zipfile
from pathlib import Path from pathlib import Path
XLSX_NS = {"m": "http://schemas.openxmlformats.org/spreadsheetml/2006/main", "r": "http://schemas.openxmlformats.org/officeDocument/2006/relationships"}
def fetch(source, cache, name, magic=b"", data=None, headers=None):
"""The bytes of a URL, downloaded into cache/name once, or of a local file."""
if not re.match(r"^https?://", source):
return Path(source).read_bytes()
path = Path(cache) / name
for attempt in range(1, 6):
if path.exists():
return path.read_bytes()
req = urllib.request.Request(source, data=data, headers={"User-Agent": "fejkdata data-import", **(headers or {})})
try:
with urllib.request.urlopen(req, timeout=600) as r:
body = r.read()
except OSError:
body = b""
if body and body.startswith(magic) and b"Request Rejected" not in body[:512]:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_bytes(body)
elif attempt < 5:
time.sleep(10 * attempt)
sys.exit(f"{source}: no valid download in 5 attempts")
def xlsx_rows(data, sheet=None):
"""The rows of an xlsx sheet named sheet, the first sheet by default, each a list of cell texts."""
z = zipfile.ZipFile(io.BytesIO(data))
strings = ["".join(t.text or "" for t in si.iter("{%s}t" % XLSX_NS["m"])) for si in ET.fromstring(z.read("xl/sharedStrings.xml")).findall("m:si", XLSX_NS)]
rels = {r.get("Id"): r.get("Target") for r in ET.fromstring(z.read("xl/_rels/workbook.xml.rels"))}
sheets = {s.get("name"): rels[s.get("{%s}id" % XLSX_NS["r"])] for s in ET.fromstring(z.read("xl/workbook.xml")).iter("{%s}sheet" % XLSX_NS["m"])}
target = sheets[sheet] if sheet else next(iter(sheets.values()))
for row in ET.fromstring(z.read("xl/" + target)).findall(".//m:row", XLSX_NS):
cells = []
for c in row.findall("m:c", XLSX_NS):
v = c.find("m:v", XLSX_NS)
cells.append("" if v is None else strings[int(v.text)] if c.get("t") == "s" else v.text)
yield cells
def write(path, columns, rows): def write(path, columns, rows):
"""Write the rows as a TSV; every cell must be non-empty and free of tabs, newlines and braces.""" """Write the rows as a TSV; every cell must be non-empty and free of tabs, newlines and braces."""
+21
View File
@@ -0,0 +1,21 @@
"""The rows of an xlsx sheet."""
import io
import xml.etree.ElementTree as ET
import zipfile
NS = {"m": "http://schemas.openxmlformats.org/spreadsheetml/2006/main", "r": "http://schemas.openxmlformats.org/officeDocument/2006/relationships"}
def rows(data, sheet=None):
"""The rows of an xlsx sheet named sheet, the first sheet by default, each a list of cell texts."""
z = zipfile.ZipFile(io.BytesIO(data))
strings = ["".join(t.text or "" for t in si.iter("{%s}t" % NS["m"])) for si in ET.fromstring(z.read("xl/sharedStrings.xml")).findall("m:si", NS)]
rels = {r.get("Id"): r.get("Target") for r in ET.fromstring(z.read("xl/_rels/workbook.xml.rels"))}
sheets = {s.get("name"): rels[s.get("{%s}id" % NS["r"])] for s in ET.fromstring(z.read("xl/workbook.xml")).iter("{%s}sheet" % NS["m"])}
target = sheets[sheet] if sheet else next(iter(sheets.values()))
for row in ET.fromstring(z.read("xl/" + target)).findall(".//m:row", NS):
cells = []
for c in row.findall("m:c", NS):
v = c.find("m:v", NS)
cells.append("" if v is None else strings[int(v.text)] if c.get("t") == "s" else v.text)
yield cells
+1 -1
View File
@@ -2,6 +2,6 @@
"format": "{prefix}{first} {last}", "format": "{prefix}{first} {last}",
"first": "{.first-name.name}", "first": "{.first-name.name}",
"last": "{.last-name.name}", "last": "{.last-name.name}",
"prefix": ["", { "format": "{title} ", "title": ["Dr", "Miss", "Mr", "Mrs", "Ms", "Mx", "Prof"], "weight": 0.1 }], "prefix": ["", { "format": "{title} ", "title": ["Dr", "Mx", "Prof"], "weight": 0.1 }],
"sex": "{.sex.name}" "sex": "{.sex.name}"
} }
+1
View File
@@ -0,0 +1 @@
{ "format": "{number}", "rows": "birth-number.tsv", "parent": "sex" }
+3
View File
@@ -0,0 +1,3 @@
sex number
f 238
m 239
1 sex number
2 f 238
3 m 239
+1 -1
View File
@@ -1 +1 @@
"{date(1930-01-01,2025-12-31,'060102')}-{.sex.birth-number}{luhn()}" "{date(1930-01-01,2025-12-31,'060102')}-{.birth-number.number}{luhn()}"
+1 -1
View File
@@ -1 +1 @@
"{date(1930-01-01,2025-12-31,'0601')}{int(61,88)}-{.sex.birth-number}{luhn()}" "{date(1930-01-01,2025-12-31,'0601')}{int(61,88)}-{.birth-number.number}{luhn()}"
+3 -3
View File
@@ -1,3 +1,3 @@
code name birth-number code name
f kvinna 238 f kvinna
m man 239 m man
1 code name birth-number
2 f kvinna 238
3 m man 239
+1 -1
View File
@@ -97,7 +97,7 @@ ids. Shape: T = table, t = template, c = choice.
| Category | Shape | Source | Licence | | Category | Shape | Source | Licence |
|---|---|---|---| |---|---|---|---|
| `person` first (female, male, generic), middle, last, weighted | T | SCB 2022 whole-population xlsx, Skatteverket 2026 surnames | CC0, "Källa: SCB" | | `person` first (female, male, generic), middle, last, weighted | T | SCB 2022 whole-population xlsx, Skatteverket 2026 surnames | CC0, "Källa: SCB" |
| `person` title, gender, birthdate, age, blood type weighted | t | geblod.nu distribution | facts | | `person` title, sex, birthdate, age, blood type weighted | t | geblod.nu distribution | facts |
| `personnummer`, `samordningsnummer` | t | Skatteverket test series: date + 238/239, Luhn | CC0 | | `personnummer`, `samordningsnummer` | t | Skatteverket test series: date + 238/239, Luhn | CC0 |
| `organisationsnummer` by form, `vat` | t | Bolagsverket group digits, Luhn, `SE…01` | facts | | `organisationsnummer` by form, `vat` | t | Bolagsverket group digits, Luhn, `SE…01` | facts |
| `company` name patterns, legal form weighted | t | Bolagsverket registrations 2025 | CC BY 2.5 SE | | `company` name patterns, legal form weighted | t | Bolagsverket registrations 2025 | CC BY 2.5 SE |