Add the breadth and inertness goals, name the contributor in the audience, and move the handoff's contents into the repo #25

Merged
lilleman merged 1 commits from goals-and-repo-holds-all into main 2026-09-19 23:49:27 +02:00
12 changed files with 1957 additions and 14 deletions
Showing only changes of commit ad57bf0c54 - Show all commits
+3
View File
@@ -8,6 +8,9 @@
offers fails with "is it still open?", as does a merge whose required checks are
still pending, so read the checks before believing the style is the problem. Use
`tea api /repos/{owner}/{repo}/pulls/<n>/merge -f Do=fast-forward-only`.
- Gitea queues Actions runs that `/repos/{owner}/{repo}/actions/tasks` does not
list, so an empty task list says nothing about whether CI ran; read
`/repos/{owner}/{repo}/commits/<sha>/status`.
- Hard tabs. No comment by default; delete a restatement, a rationale, history, or a file preamble.
- One spelling per result: reject the other at `New`, and let the error name the spelling to use.
- A standing choice a reader would relitigate goes under Decisions in the README, not in a comment.
+16 -3
View File
@@ -770,6 +770,9 @@ App developers writing tests and fixtures, in Go and at a shell:
- a **hand fixture author**, one value at a shell
- a **validator-facing author**, who needs a value a real checker accepts
and a **contributor**, who reads [`todo.md`](todo.md), [`AGENTS.md`](AGENTS.md) and
the Development section below, and who ships a register the four above then draw from.
## Goals
1. **Valid by construction** — every value passes the check its real consumer
@@ -800,6 +803,16 @@ App developers writing tests and fixtures, in Go and at a shell:
step that replaces it. Only non-factual copy stays authored. A sourced table
holds the rows its source holds: none is added by hand, and one is dropped
only by a rule the script states.
11. **Breadth follows what most systems store** — a category is added in proportion
to how many real schemas hold it: names, addresses, phones, ids, money and
timestamps before anything domain-specific, and a catalogue serving one niche
waits behind everything serving many.
12. **Realism is the default, inertness is selectable** — where a value could reach
something real, a domain anyone may register or an account a bank could issue,
the realistic breadth ships *and* so does the subset that provably reaches
nothing, each on its own path. A fixture that looks nothing like production
tests nothing; the caller who needs a value that can touch nothing asks for it
by name.
## Decisions
@@ -1151,8 +1164,6 @@ App developers writing tests and fixtures, in Go and at a shell:
tells the two apart, `sex[f].first-name[Kim]`, the ambiguity error spells each
row inside its parent, and a name repeating inside one parent row is refused at
load, since nothing could then select it.
- **No pop-culture catalogues.** Every other faker ships film, band and character
names; fejkdata ships none.
- **`misc` is what every locale shares.** A category whose facts differ by country
belongs in that country's locale, read from the register that country's own
records use; `misc` takes only sources that are international. NHTSA vPIC and
@@ -1272,7 +1283,9 @@ REPIN=1 docker compose run --rm --user "$(id -u):$(id -g)" test
A shipped table built from a source is rebuilt by its script under
[`data-import/`](data-import), one command per dataset, fetching the source named in
[`DATA-LICENSES.md`](DATA-LICENSES.md). Downloads are cached under
`data-import/cache/`, so delete it to fetch afresh; `geo-us.py` fetches two
`data-import/cache/`, which is ignored by version control and grows past a
gigabyte, so pass `--cache DIR` to reuse a copy you already have and delete it to
fetch afresh; `geo-us.py` fetches two
TIGER/Line files per county it ships, a few hundred megabytes, `geo-se.py` needs
a Trafikverket API key, free at [data.trafikverket.se](https://data.trafikverket.se/),
in `TRAFIKVERKET_API_KEY` or a `--key-file`, and the Census host behind `geo-us.py`
@@ -0,0 +1,148 @@
# Part 1 — Names, finance, company: sources and rules
Researched 2026-09-17. "Embed" = ship as a data file in an MIT repo. Everything not marked UNVERIFIED was read from the cited page or file this session.
## 1. US person names and identifiers
### SSA baby names (given names by year and sex)
- Page: https://www.ssa.gov/oact/babynames/limits.html — national zip https://www.ssa.gov/oact/babynames/names.zip (7 MB), state zip https://www.ssa.gov/oact/babynames/state/namesbystate.zip (23 MB), territories 228 kB.
- Format (per SSA page + readme as mirrored by https://github.com/hackerb9/ssa-baby-names): one `yobYYYY.txt` per year 1880–2025, CSV `name,sex,count`, name 2–15 chars, sex `M`/`F`, sorted by sex then count desc. Names with < 5 occurrences in a year are excluded.
- Counts: ~100,364 distinct names, ~2.02 M rows across all years (hackerb9 snapshot; year of snapshot UNVERIFIED). Top 1000 names cover 71.5 % of 2025 births (SSA page).
- Licence: data.gov catalog entry https://catalog.data.gov/dataset/baby-names-from-social-security-card-applications-national-data lists licence **CC0 1.0** (`https://creativecommons.org/publicdomain/zero/1.0/`), publisher SSA. US federal work, no attribution required; embeddable.
- Note: ssa.gov returns 403 to non-browser user agents (curl/WebFetch); download with a browser UA or via the data.gov link.
### Census 2010 surnames
- Page: https://www.census.gov/topics/population/genealogy/data/2010_surnames.html
- Files: full set https://www2.census.gov/topics/genealogy/2010surnames/names.zip (CSV + XLSX, `Names_2010Census.csv`), top 1000 `Names_2010Census_Top1000.xlsx`, docs https://www2.census.gov/topics/genealogy/2010surnames/surnames.pdf.
- Fields (from surnames.pdf): `name, rank, count, prop100k, cum_prop100k, pctwhite, pctblack, pctapi, pctaian, pct2prace, pcthispanic`. `(S)` = suppressed percentage. Names are UPPERCASE.
- Count: 162,253 surnames occurring ≥ 100 times (covers 90.1 % of people with a recorded surname); the CSV also has a trailing `ALL OTHER NAMES` row (UNVERIFIED — www2.census.gov returned 403 to curl this session; download with a browser).
- Licence: US Census Bureau work → US federal government work, public domain in the US (17 U.S.C. §105). Census only asks for citation (https://www.census.gov/about/policies/citation.html: "U.S. Census Bureau, [Table], [Product], [Vintage], [URL], accessed on [date]"). Embeddable; cite source in the data file header.
### Name prefixes / suffixes
- No authoritative enumerated list exists. USPS Publication 28 §33 (https://pe.usps.com/text/pub28/28c3_011.htm) defines the fields: Name Prefix, First Name, Middle Name or Initial, Surname, **Suffix Title** = "maturity (e.g., JR, SR) and professional (e.g., PHD, DDS) suffixes"; example "MR. WALTER W. WITHERSPOON JR.".
- Recommendation: hand-curate. Prefixes `Mr., Mrs., Ms., Mx., Dr., Rev.`; suffixes `Jr., Sr., II, III, IV, PhD, MD, DDS, Esq.`. Rule of thumb: generational suffix (Jr/Sr/II–IV) only with a male name in traditional US usage; never combine Jr. with Sr.; at most one generational suffix.
### SSN
- Structure: `AAA-GG-SSSS` (area 3, group 2, serial 4). No check digit (Wikipedia).
- Randomization since **2011-06-25** (https://www.ssa.gov/employer/randomization.html): "Previously unassigned area numbers were introduced for assignment excluding area numbers 000, 666 and 900-999." Area no longer geographic.
- Invalid per SSA POMS RM 10201.035 (https://secure.ssa.gov/apps10/poms.nsf/lnx/0110201035): area `000`, `666`, `900–999`; group `00`; serial `0000`.
- Generator rule: area ∈ [001,899] \ {666}; group ∈ [01,99]; serial ∈ [0001,9999].
- Advertising SSNs (https://www.ssa.gov/history/ssn/misused.html): `078-05-1120` (Woolworth wallet card, 1938; > 40,000 people claimed it) and `219-09-9999` (1940 Social Security Board pamphlet). Both are structurally valid — exclude from output.
- `987-65-4320`–`987-65-4329` "reserved for advertising": repeated by searchbug.com, ssn-check.org, a Missouri state standard (https://oa.mo.gov/sites/default/files/CC-SocialSecurityNumberNamingStandardV050305.pdf, which even misprints it as 987-54-4329). **UNVERIFIED** — not found on any ssa.gov page or POMS this session; Wikipedia no longer carries it. Safe to use these ten as the library's "obviously fake" examples, but do not cite SSA for it.
### ITIN (IRS)
- IRS Publication 4757 (Rev. 6-2026), p. 5 (https://www.irs.gov/pub/irs-pdf/p4757.pdf): "All valid ITINs are nine-digit numbers in the same format as the SSN (9XX-8X-XXXX), beginning with a "9" and the 4th and 5th digits ranging from 50 to 65, 70 to 88, 90 to 92, and 94 to 99."
- Group `93` = ATIN (adoption TIN); `89` reserved (Intuit/TaxSlayer support pages, not IRS — UNVERIFIED).
- Generator rule: `9` + 2 digits + group ∈ {50–65, 70–88, 90–92, 94–99} + 4 digits (serial `0000` not documented as invalid for ITIN; avoid it anyway).
### EIN (IRS)
- Format `NN-NNNNNNN`. Page: https://www.irs.gov/businesses/small-businesses-self-employed/how-eins-are-assigned-and-valid-ein-prefixes ("The first two digits of your EIN show the IRS campus that assigned it or if you applied online").
- Valid prefixes (union of the table, 2026-09-17):
`01 02 03 04 05 06 10 11 12 13 14 15 16 20 21 22 23 24 25 26 27 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 71 72 73 74 75 76 77 80 81 82 83 84 85 86 87 88 90 91 92 93 94 95 98 99`
By campus: Andover 10,12 · Atlanta 60,67 · Austin 50,53 · Brookhaven 01–06,11,13,14,16,21–23,25,34,51,52,54–59,65 · Cincinnati 30,32,35–38,61 · Fresno 15,24 · Kansas City 40,44 · Memphis 94,95 · Ogden 80,90 · Philadelphia 33,39,41–43,46,48,62–64,66,68,71–77,85–88,91–93,98,99 · Internet 20,26,27,33,39,41,42,45–47,81–88,92,93,99 · SBA 31.
- Not valid: 00, 07, 08, 09, 17, 18, 19, 28, 29, 49, 69, 70, 78, 79, 89, 96, 97. No check digit.
## 2. Finance
### ABA routing number (RTN)
- Wikipedia https://en.wikipedia.org/wiki/ABA_routing_transit_number: 9 digits; check: `(3(d1+d4+d7) + 7(d2+d5+d8) + (d3+d6+d9)) mod 10 = 0`.
- First two digits: `00` US Government; `01–12` Federal Reserve districts (normal banks); `21–32` thrifts (historic, still valid); `61–72` non-bank payment processors/clearinghouses; `80` traveler's checks. Other prefixes unused (Wikipedia does not say so explicitly — inferred).
- Generator rule: prefix ∈ {01–12, 21–32, 61–72}, 6 random digits, compute 9th as `(10 − (3d1+7d2+d3+3d4+7d5+d6+3d7+7d8) mod 10) mod 10`.
- Open list of real RTNs: the Federal Reserve E-Payments Routing Directory (https://www.frbservices.org/resources/routing-number-directory/index.html) says "may not be sold, re-licensed, or otherwise used for commercial gain" → **not embeddable**. Generate synthetically.
### Payment card numbers (Wikipedia "Payment card number", https://en.wikipedia.org/wiki/Payment_card_number; all Luhn-checked)
| Network | IIN ranges | Length | CVV |
|---|---|---|---|
| Visa | 4 | 13, 16, 19 | 3 |
| Mastercard | 51–55, 2221–2720 | 16 | 3 |
| American Express | 34, 37 | 15 | 4 (front) |
| Discover | 6011, 644–649, 65, 622126–622925 (UnionPay co-brand) | 16–19 | 3 |
| JCB | 3528–3589 | 16–19 | 3 |
| Diners Club International | 30 (covers 300–305, 3095), 36, 38, 39 | 14–19 | 3 |
| Diners Club US & Canada | 55 (Mastercard co-brand) | 16 | 3 |
| UnionPay | 62 | 16–19 | 3 |
| Maestro | 5018, 5020, 5038, 5893, 6304, 6759, 6761, 6762, 6763 | 12–19 | 3 |
| Maestro UK | 6759, 676770, 676774 | 12–19 | 3 |
| Dankort 5019 · Mir 2200–2204 · RuPay 60,65,81,82,508 · Troy 65,9792 · UATP 1 (15) · Verve 506099–506198, 507865–507964, 650002–650027 (16/18/19) | | | |
- CVV lengths: https://en.wikipedia.org/wiki/Card_security_code ("three-digit ... Visa, Mastercard, and Discover"; "American Express is a four-digit code on the front").
- The requested "Maestro 50/56–58/6xxx" is a looser legacy rule; use Wikipedia's explicit IINs.
- Wikipedia text is CC BY-SA 4.0 — copy the *facts* (not prose) into your table; facts are not copyrightable.
### Published test cards
- Stripe: https://docs.stripe.com/testing — Visa 4242424242424242, Mastercard 5555555555554444, Amex 378282246310005, Discover 6011111111111117, Diners 3056930009020004 (14), JCB 3566002020360505, UnionPay 6200000000000005; any future expiry, any CVC (4 for Amex).
- PayPal: https://developer.paypal.com/tools/sandbox/card-testing/ — Visa 4005519200000004, 4012000033330026, 4012000077777777, 4012888888881881, 4217651111111119, 4500600000000061, 4772129056533503, 4915805038587737; Mastercard 2223000048400011; Amex 371449635398431, 376680816376961; Diners 36461510000039, 36461510000013; Maestro 6304000000000000, 5063516945005047; JCB 3636500000000260, 3636500000000989; CUP 6200680000000004, 6200680000000038.
- Reuse: neither page carries an open licence (Stripe docs are under Stripe's site terms). The numbers themselves are Luhn-valid facts, widely reproduced; embedding the numbers with a "from Stripe/PayPal test docs" note is low-risk but **UNVERIFIED** as an explicit grant. Alternative: generate Luhn-valid numbers from the IIN table above — no dependency on either vendor.
### BIC / SWIFT (ISO 9362)
- https://en.wikipedia.org/wiki/ISO_9362: 4 letters bank code + 2 letters ISO 3166-1 country + 2 alphanumeric location + optional 3 alphanumeric branch (8 or 11 chars; `XXX` = primary office). Location second char `0` = test BIC, `1` = passive participant, `2` = reverse billing.
- Generator rule: `[A-Z]{4}[A-Z]{2}[A-Z2-9][A-NP-Z0-9]([A-Z0-9]{3})?` — avoid `0`/`1` as the 8th char for a "live" BIC.
- SWIFT's BIC Directory is proprietary (SwiftRef licence, redistribution needs a separate "SwiftRef Redistribution License": https://www.swift.com/myswift/ordering/order-products-services/swiftref-redistribution-license) → **not embeddable**.
- Open alternative: GLEIF BIC-to-LEI mapping, https://www.gleif.org/en/lei-data/lei-mapping/download-bic-to-lei-relationship-files, files at https://mapping.gleif.org/api/v2/bic-lei/ (monthly zip; `LEI-BIC-20260828.zip` → `lei-bic-20260828T000000.csv`, columns `LEI,BIC`, 39,347 rows). Licence: "BIC/LEI Mapping Table License Agreement" (https://www.gleif.org/lei-data/lei-mapping/download-bic-to-lei-relationship-files/2017-12-21_annex-2_bic-to-lei-mapping-table-license-agreement_final.pdf) — a CC0-style grant "for any purpose whatsoever, including ... commercial" but **conditional on reproducing the notice** "SWIFT © and database rights [month year]. All rights reserved. This Mapping Table has been developed by SWIFT. Any use of the Mapping Table ... is subject to the BIC/LEI Mapping Table License Agreement ... For the latest BIC information and updates, always refer to www.swift.com/bic." Embeddable in an MIT repo with that notice in the data file; it is a list of real live BICs (bank names not included — join to GLEIF LEI data if names are wanted).
### ISIN (https://en.wikipedia.org/wiki/International_Securities_Identification_Number)
- 12 chars: 2-letter country + 9 alphanumeric NSIN + 1 check digit.
- Check: convert letters A=10…Z=35 (ASCII − 55) to produce a digit string, then Luhn over that string (double every second digit from the right, sum digits, check makes total ≡ 0 mod 10). Pitfall: work on the *expanded* digit string; a transposed letter pair can pass.
### CUSIP (https://en.wikipedia.org/wiki/CUSIP)
- 9 chars: 6 issuer + 2 issue + 1 check. Values: digits as-is, A=10…Z=35, `*`=36, `@`=37, `#`=38.
- Check: for i in 1..8, v = value; if i even, v ×= 2; sum += v div 10 + v mod 10; check = (10 − sum mod 10) mod 10.
- Real CUSIPs are proprietary (CGS/FactSet); generate synthetic ones only.
### IBAN (ISO 13616)
- Validation (https://en.wikipedia.org/wiki/International_Bank_Account_Number): move first 4 chars to end, letters → 10–35, integer mod 97 must equal 1; max 34 chars; check digits 02–98.
- SWIFT registry: https://www.swift.com/standards/data-standards/iban-international-bank-account-number, PDF release 101 https://www.swift.com/sites/default/files/files/iban-registry-v101.pdf (release 100 was Oct 2025); TXT also published (URL **UNVERIFIED** — swift.com returns 403/errors to automation this session). Terms: free of charge; no explicit licence text found (**UNVERIFIED**). Country count: 89 countries per Wikipedia (Dec 2024).
- Embeddable per-country length tables already extracted from the registry:
- php-iban `registry.txt` (https://github.com/globalcitizen/php-iban, **LGPL-3.0**): pipe-separated, 121 rows (116 official + unofficial), columns `country_code|country_name|domestic_example|bban_example|bban_format_swift|bban_format_regex|bban_length|iban_example|iban_format_swift|iban_format_regex|iban_length|bban_bankid_start_offset|...|country_sepa|swift_official|...|currency_iso4217|central_bank_url|central_bank_name|membership`. LGPL data in an MIT repo is awkward; use it as a *reference* to build your own table (formats are facts) rather than copying the file.
- ibankit-js (https://github.com/koblas/ibankit-js, **Apache-2.0**, registry v95) — same caveat.
- Best route: own table `{country, length, bban_format}` derived from the SWIFT PDF (facts), cite SWIFT.
### Currencies (ISO 4217)
- ISO says use is free: https://www.iso.org/iso-4217-currency-codes.html — "ISO allows free-of-charge use of its country, currency and language codes from ISO 3166, ISO 4217 and ISO 639". Official list from SIX: https://www.six-group.com/dam/download/financial-information/data-center/iso-currrency/lists/list-one.xml (+ `.xls`, `list-three` historic). Fields `CtryNm, CcyNm, Ccy, CcyNbr, CcyMnrUnts`; published 2026-01-01; 280 entity rows, 178 distinct codes. No symbols. No licence text inside the XML.
- Alternatives:
- datasets/currency-codes (https://github.com/datasets/currency-codes): **PDDL** (public domain); `data/codes-all.csv` fields `Entity, Currency, AlphabeticCode, NumericCode, MinorUnit, WithdrawalDate`; built from SIX list one + three; no symbols. Embeddable.
- npm `currency-codes` v2.2.0 (https://github.com/freeall/currency-codes): **MIT**; fields `code, number, digits, currency, countries[]`; generated from the SIX XML; no symbols. Embeddable.
- umpirsky/currency-list (https://github.com/umpirsky/currency-list): **MIT**; code → localised name only (311 entries in `data/en_US/currency.json`, includes historic); no symbols, no numeric codes, no decimals.
- Debian iso-codes (https://salsa.debian.org/iso-codes-team/iso-codes): **LGPL-2.1+**; `data/iso_4217.json` has `alpha_3, name, numeric` only (179 entries); no symbols/minor units. LGPL → avoid embedding.
- Symbols: none of the above carry them. Unicode CLDR (`common/main/en.xml` currency symbols, Unicode License, MIT-compatible with notice) is the usual source — **UNVERIFIED this session**.
### Crypto addresses
- Bitcoin (https://en.bitcoin.it/wiki/List_of_address_prefixes): P2PKH version 0x00 → leading `1`; P2SH 0x05 → leading `3`; Base58Check 25–34 chars (alphabet excludes `0OIl`; last 4 bytes = first 4 of double-SHA256 of version+payload). Testnet `m`/`n`, `2`, `tb1`.
- Bech32 (BIP-173, https://github.com/bitcoin/bips/blob/master/bip-0173.mediawiki, BSD-2-Clause): HRP `bc`, separator `1`, charset `qpzry9x8gf2tvdw0s3jn54khce6mua7l`, all-lowercase (or all-uppercase), max 90 chars; P2WPKH `bc1q…` = 42 chars, P2WSH `bc1q…` = 62 chars.
- Bech32m (BIP-350, https://github.com/bitcoin/bips/blob/master/bip-0350.mediawiki): witness v1+ (P2TR `bc1p…`, 62 chars), checksum constant `0x2bc830a3` instead of 1.
- Ethereum EIP-55 (https://eips.ethereum.org/EIPS/eip-55, CC0): `0x` + 40 hex; keccak256 of the lowercase hex (no 0x, as ASCII); for each hex letter, uppercase iff the corresponding hash nibble ≥ 8 (bit 4·i set). Digits unchanged.
## 3. Company
### NAICS 2022 (US Census)
- Page https://www.census.gov/naics/ ; files: 6-digit list https://www.census.gov/naics/2022NAICS/6-digit_2022_Codes.xlsx (82 kB; **1,012** six-digit codes, 111110…928120, verified by parsing), 2–6 digit https://www.census.gov/naics/2022NAICS/2-6%20digit_2022_Codes.xlsx, structure https://www.census.gov/naics/2022NAICS/2022_NAICS_Structure.xlsx, descriptions https://www.census.gov/naics/2022NAICS/2022_NAICS_Descriptions.xlsx, manual PDF https://www.census.gov/naics/reference_files_tools/2022_NAICS_Manual.pdf.
- Format: xlsx, column A code (numeric), column B title (trailing spaces in some titles — trim).
- Licence: no statement on the page or in the manual front matter; US federal government work → public domain in the US (17 U.S.C. §105). Embeddable; cite Census/OMB.
### SEC EDGAR company names
- `https://www.sec.gov/files/company_tickers.json`: object keyed `"0".."N"` of `{cik_str, ticker, title}`; **10,422** entries (2026-09-17). `company_tickers_exchange.json`: `{"fields":["cik","name","ticker","exchange"],"data":[...]}`.
- `https://www.sec.gov/Archives/edgar/cik-lookup-data.txt`: 40 MB, **1,059,372** lines, `NAME:CIK:` (`COMPANY NAME:0001234567:`), all filers incl. individuals — filter to the ticker list for company names.
- Access: declare `User-Agent: Company Name email@domain`, ≤ 10 req/s (https://www.sec.gov/os/webmaster-faq). Licence: "All Government-created content on sec.gov and EDGAR public filing content are free to access and reuse" (same FAQ). Embeddable; the ticker file (~800 kB) is the practical one.
### DUNS
- https://en.wikipedia.org/wiki/Data_Universal_Numbering_System: 9 digits, no significance; shown `NN-NNN-NNNN` or plain. Had a mod-10 check digit until ~Dec 2006; **dropped** since (expanded the pool by 800 M). So: any 9 digits are structurally valid today; no validation rule to enforce. Optional DUNS+4 suffix (4 alphanumerics) is user-assigned.
### Legal entity suffixes per country
- GLEIF ISO 20275 Entity Legal Forms code list v1.6 (2026-02-19): page https://www.gleif.org/en/lei-data/code-lists/iso-20275-entity-legal-forms-code-list ; CSV https://www.gleif.org/lei-data/code-lists/iso-20275-entity-legal-forms-code-list/2026-02-19-elf-code-list-v1.6.csv (758 kB; xlsx alongside).
- Columns: `ELF Code, Country of formation, Country Code (ISO 3166-1), Jurisdiction of formation, Country sub-division code (ISO 3166-2), Entity Legal Form name Local name, Language, Language Code (ISO 639-1), Entity Legal Form name Transliterated name, Abbreviations Local language, Abbreviations transliterated, Date created, ELF Status ACTV/INAC, Modification, Modification date, Reason`.
- Counts: 4,002 rows; 130 countries; 3,790 ACTV / 212 INAC; 1,518 rows have a local abbreviation (e.g. US `N.A.`, `FSA`; SE `AB`… multiple abbreviations separated by `;`). US has 737 rows (per-state forms). Abbreviations are the field you want for `Inc`, `LLC`, `GmbH`, `AB`, `SAS`, `Pty Ltd`; many rows have none, so filter on non-empty abbreviation and status ACTV.
- Licence: GLEIF Open Data page (https://www.gleif.org/en/about/open-data): "The data on GLEIF's website is provided under a Creative Commons (CC0) license"; LEI Data Terms of Use: "provided under the CC0 licence, see CC0 1.0 Universal". The ELF page itself names no licence — that CC0 explicitly covers this code list is **UNVERIFIED**, but GLEIF treats all its published data this way. Embeddable.
## Quick embed verdict
| Item | Embed? | Licence / attribution |
|---|---|---|
| SSA baby names | yes | CC0 (data.gov) |
| Census 2010 surnames | yes | US gov PD; cite Census |
| EIN prefixes, ITIN/SSN rules | yes (rules) | facts from IRS/SSA |
| NAICS 2022 | yes | US gov PD |
| SEC company_tickers.json | yes | free to reuse (SEC) |
| GLEIF ELF code list | yes | CC0 (GLEIF) |
| GLEIF BIC-LEI (real BICs) | yes, with mandatory SWIFT notice | custom permissive licence |
| ISO 4217 (SIX XML / datasets PDDL / npm MIT) | yes | free use (ISO), PDDL, MIT |
| Fed routing directory, SwiftRef BIC directory, CUSIP lists | no | restricted/proprietary — generate synthetically |
| php-iban / ibankit registries | copy facts only | LGPL-3.0 / Apache-2.0 |
| Stripe / PayPal test cards | numbers only, note source | no explicit grant (UNVERIFIED) |
@@ -0,0 +1,118 @@
# Open data sources: Internet, Vehicle, Airline (researched 2026-09-17)
Verdict key: **EMBED** = safe to vendor into an MIT repo; **EMBED+ATTR** = ok with a notice; **AVOID** = licence incompatible or unclear.
## 4. Internet
### IANA registries (all of them)
- Licence: IANA/IETF statement at https://www.iana.org/help/licensing-terms — "IANA and IETF intend that the Protocol Registries may be freely used by any party for any purpose", data dedicated under **CC0 1.0**. No attribution required. **EMBED**.
- All registries below share this licence; ship one `SOURCES`/NOTICE line pointing at that URL.
| Registry | CSV URL | Rows (2026-09-16) | Notes |
|---|---|---|---|
| Media types | `https://www.iana.org/assignments/media-types/<type>.csv`, type ∈ application, audio, font, haptics, image, message, model, multipart, text, video (also `example`) | application 1800, audio 165, font 6, haptics 3, image 88, message 27, model 42, multipart 17, text 105, video 97 (header excluded) ≈ **2350** | Fields: `Name,Template,Reference`. Template is the full `type/subtype`; some rows are DEPRECATED/OBSOLETED in Name (filter). Format: RFC 6838 §4.2 — `type "/" subtype`, restricted chars, max 127 chars each, case-insensitive. |
| HTTP status codes | https://www.iana.org/assignments/http-status-codes/http-status-codes-1.csv | 75 rows, **64 assigned** (306 and 418 are "(Unused)"; 104 is TEMPORARY) → 62 usable | Fields: `Value,Description,Reference`. Value is 3 digits 1xx–5xx. |
| IPv4 special-purpose | https://www.iana.org/assignments/iana-ipv4-special-registry/iana-ipv4-special-registry-1.csv | 26 blocks | Fields: `Address Block,Name,RFC,Allocation Date,Termination Date,Source,Destination,Forwardable,Globally Reachable,Reserved-by-Protocol`. |
| IPv6 special-purpose | https://www.iana.org/assignments/iana-ipv6-special-registry/iana-ipv6-special-registry-1.csv | 27 blocks | Same fields. Contains `2001:db8::/32` and `3fff::/20`. |
| Service names / ports | https://www.iana.org/assignments/service-names-port-numbers/service-names-port-numbers.csv | 14,534 rows; 11,731 with both name and a single port; **6,102 distinct named ports**; 6,008 tcp-named; 707 tcp names < 1024 | Fields: `Service Name,Port Number,Transport Protocol,Description,Assignee,Contact,Registration Date,Modification Date,Reference,Service Code,Unauthorized Use Reported,Assignment Notes`. Port Number may be empty, single, or a range `a-b`. Well-known = 0–1023, registered 1024–49151, dynamic 49152–65535 (RFC 6335). ~1.1 MB; embed only the named-port subset. |
| TLDs | https://data.iana.org/TLD/tlds-alpha-by-domain.txt | **1,401** TLDs (first line is a `# Version …` comment) | Upper-case ASCII; IDNs as `XN--…` punycode. Same IANA CC0 terms (UNVERIFIED that data.iana.org is explicitly covered by the licensing page; it is the same registry data). Lower-case on use. |
### MIME extension mapping: jshttp/mime-db
- https://github.com/jshttp/mime-db — **MIT** (GitHub API spdx MIT; npm `license: MIT`), version 1.54.0 on npm.
- `db.json` (CDN: https://cdn.jsdelivr.net/npm/mime-db/db.json): object keyed by lower-case type → `{source, extensions[], compressible, charset}`.
- Counts (1.54.0): **2,522 types**, **1,015 with extensions**; source = iana 2,136, apache 275, nginx 13, custom 98.
- Sources: IANA registry (CC0), Apache httpd `mime.types` (Apache-2.0), nginx `mime.types` (BSD-2). **EMBED+ATTR** (keep the MIT copyright notice). Data updates are not semver-breaking per README.
### IPv4 / IPv6 generation rules (RFCs, no data file needed)
- RFC 5737 documentation: `192.0.2.0/24` (TEST-NET-1), `198.51.100.0/24` (TEST-NET-2), `203.0.113.0/24` (TEST-NET-3). https://www.rfc-editor.org/rfc/rfc5737.html
- RFC 1918 private: `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16`. Also useful: `100.64.0.0/10` CGNAT (RFC 6598), `198.18.0.0/15` benchmarking (RFC 2544), `169.254.0.0/16` link-local.
- IPv6 documentation: `2001:db8::/32` (RFC 3849) and `3fff::/20` (RFC 9637, "expands on the existing 2001:db8::/32 … with the reservation of an additional, larger prefix"). ULA `fc00::/7` (use `fd00::/8`), link-local `fe80::/10`.
- Validation: generate inside those blocks; avoid network/broadcast (`.0`/`.255`) if the consumer expects host addresses. Format IPv6 per RFC 5952 (lower-case hex, `::` once, longest zero run).
### MAC OUI (IEEE MA-L)
- CSV: https://standards-oui.ieee.org/oui/oui.csv (3.8 MB, **40,161 rows**, fields `Registry,Assignment,Organization Name,Organization Address`; Assignment is 6 hex chars). Server returns HTTP 418 to non-browser user agents; fetch with a browser UA. Also MA-M `oui28.csv`, MA-S `oui36.csv`, CID `cid.csv`.
- Licence: no licence on the IEEE pages themselves (site footer "© IEEE – All rights reserved" is generic). Debian's `ieee-data` copyright file (https://metadata.ftp-master.debian.org/changelogs/main/i/ieee-data/unstable_copyright) records IEEE's 2014 statement: *"IEEE does not assert any copyright in the OUI Public Listing or attempt to restrict distribution of the listing in any way. The IEEE Registration Authority does, however, strongly encourage those who use the list to obtain it directly from IEEE…"* — same text in the FSF directory. Debian ships it in `main` on that basis. **EMBED+ATTR** (cite the statement); the org-address column is not needed — keep `Assignment,Organization Name` only (~1 MB → few hundred KB). UNVERIFIED: no current IEEE page restates it (regauth FAQ and IPR pages do not mention it).
- Alternative needing no data: locally administered unicast MAC — first octet's bit 1 (U/L) = 1, bit 0 (I/G) = 0 → second hex digit ∈ {2, 6, A, E}, e.g. `x2:xx:xx:xx:xx:xx`. Guaranteed never to collide with a vendor OUI. Notations: `01:23:45:67:89:ab`, `01-23-45-67-89-AB`, Cisco `0123.4567.89ab`.
### User-agent strings
| Source | Licence | Verdict |
|---|---|---|
| https://www.useragents.me/ | No licence stated anywhere on the page ("© 2022–2025 Useragents.me"); `/api` returns 404 now; JSON blobs embedded in page, weekly updates | **AVOID** (unlicensed = all rights reserved) |
| https://explore.whatismybrowser.com/useragents/explore/ + API legal https://developers.whatismybrowser.com/api/about/legal/ | Proprietary: "You may NOT share the unique database download URL, or the downloaded file itself … EXCEPT for making it publicly available on the internet" | **AVOID** |
| npm `user-agents` (intoli) https://github.com/intoli/user-agents | **BSD-2-Clause**, v2.1.185, daily-rebuilt dataset from Intoli traffic, weighted by frequency; fields `userAgent, platform, vendor, appName, deviceCategory, screenWidth/Height, viewportWidth/Height, connection{…}, oscpu, cpuClass, pluginsLength` | **EMBED+ATTR** (keep BSD notice). Count not stated on the page (UNVERIFIED; dataset is a few thousand rows). |
| npm `top-user-agents` (microlinkhq) https://github.com/microlinkhq/top-user-agents | **MIT** (GitHub API). Top 100 UA strings from microlink.io traffic, weekly; JSON: https://cdn.jsdelivr.net/gh/microlinkhq/top-user-agents@master/src/index.json (+ desktop.json, mobile.json) | **EMBED+ATTR** — smallest clean option |
- None of the above is CC0. Alternative: generate UAs from templates (Chrome/Firefox/Safari/Edge strings are fixed-form; only version numbers vary), no data needed.
### Programming languages: GitHub Linguist
- https://raw.githubusercontent.com/github-linguist/linguist/main/lib/linguist/languages.yml — **MIT** (repo LICENSE, "Copyright (c) 2017 GitHub, Inc."). **EMBED+ATTR**.
- **835** top-level entries; **562** with `type: programming` (others: data, markup, prose). Per-entry fields: `type, color, extensions, aliases, tm_scope, ace_mode, language_id, group, interpreters, …`.
### Operating systems
- No open dataset found beyond awesome-lists. OS names/versions are facts — hand-author a small table (Windows 10/11, Windows Server, macOS versions+names, Ubuntu/Debian/Fedora/Arch/Alpine, iOS, Android, FreeBSD…). `user-agents` (BSD-2) carries `platform`/`oscpu` values if a data-driven list is wanted.
### Database engines / collations
- Engine names: hand-author (facts).
- MySQL manual (https://dev.mysql.com/doc/refman/8.4/en/preface.html): Oracle terms — "you may not … copy, reproduce … distribute … any part" except distributing the docs with the software. **AVOID copying tables from MySQL docs.** Instead dump `SHOW COLLATION` from a `mysql:8.4.x` container (server output is factual, not the manual) — ~286 collations / 41 charsets in 8.x (UNVERIFIED count). PostgreSQL docs are under the PostgreSQL licence (permissive, keep notice) — `pg_collation` dump or docs both fine. MariaDB KB is CC BY-SA — avoid.
## 5. Vehicle
### NHTSA vPIC
- API base `https://vpic.nhtsa.dot.gov/api/vehicles/`, `?format=json|xml|csv`. Endpoints verified: `GetAllMakes` (**12,363 makes**, fields `Make_ID,Make_Name`; 612 KB JSON, mostly small trailer/custom builders — filter with `GetMakesForVehicleType/car` → 195 makes, `/truck` → 207), `GetModelsForMake/honda` (362 models, `Make_ID,Make_Name,Model_ID,Model_Name`), `DecodeWMI/1HG`, `GetWMIsForManufacturer/honda` (45 WMIs, fields `WMI,Name,Country,VehicleType,Id,…`), `GetVehicleVariableList`, `GetVehicleVariableValuesList/<display name>` (must be the display name, URL-encoded, e.g. `Fuel%20Type%20-%20Primary`; the camel-case name returns 0 rows).
- Value lists (counts): **Fuel Type - Primary 14**, **Body Class 71**, **Transmission Style 12**, **Drive Type 23**, **Vehicle Type 9**. Fields `Id,Name`.
- Bulk: https://vpic.nhtsa.dot.gov/downloads/ — `vPICList_lite_2026_08.bak.zip` (MS SQL Server 2019+ backup, 190 MB, updated 2026-08-14), also `.plain.zip` / `.custom.zip`. Page says standalone DB "limited to VIN decoding"; makes/models/variables still via API. Contains `Wmi`, `Make`, `Model`, `WMIYearValidChars` tables (UNVERIFIED table names).
- Licence: FAQ https://vpic.nhtsa.dot.gov/api/home/index/faq — "No, NHTSA is a government agency and the services provided on the API are free for use by the public as an offering as a part of our Open Data initiatives"; no rate-limit number stated, batch jobs asked to run at night. Work of the US federal government → public domain under 17 USC §105; data.gov entry lists licence as "unknown". **EMBED** (note source; no attribution legally required).
### VIN (ISO 3779 / 49 CFR 565, North America)
- 17 chars from `0-9 A-H J-N P R-Z` (**I, O, Q never**). Positions: 1–3 WMI, 4–8 VDS, 9 check digit, 10 model year, 11 plant, 12–17 serial (last 4 must be digits for US, positions 12–14 alphanumeric — 49 CFR 565.15).
- Check digit: transliterate `A=1 B=2 C=3 D=4 E=5 F=6 G=7 H=8 J=1 K=2 L=3 M=4 N=5 P=7 R=9 S=2 T=3 U=4 V=5 W=6 X=7 Y=8 Z=9`, digits as-is; weights `8 7 6 5 4 3 2 10 0 9 8 7 6 5 4 3 2`; sum mod 11 → digit, remainder 10 → `X`.
- Model year (pos 10): excludes `I O Q U Z 0`; `A=2010 … H=2017 J=2018 K=2019 L=2020 M=2021 N=2022 P=2023 R=2024 S=2025 T=2026 V=2027 W=2028 X=2029 Y=2030`, `1=2031…9=2039` (30-year cycle; `1–9` also = 2001–2009). Source: https://en.wikipedia.org/wiki/Vehicle_identification_number (rules are facts; do not copy prose).
### WMI list
- SAE J1044 / SAE WMI database is paid (https://www.sae.org/standards/j1044_202501-world-manufacturer-identifier) — **AVOID**.
- Wikipedia "List of WMIs" — CC BY-SA — **AVOID**.
- Use vPIC: `GetWMIsForManufacturer/<name or id>` per manufacturer, or the `Wmi` table in the standalone `.bak`. Public domain. Country prefix rule (first char): `1,4,5` USA, `2` Canada, `3` Mexico, `J` Japan, `K` Korea, `S` UK, `W` Germany, `Y` Sweden/Finland, `Z` Italy, `L` China (ISO 3780 facts).
### US licence plate formats
| Source | Licence | Verdict |
|---|---|---|
| openalpr `runtime_data/postprocess/us.patterns` (https://github.com/openalpr/openalpr) | **AGPL-3.0** | **AVOID** (291 lines, `@`=letter `#`=digit — useful only as a reference to check hand-authored patterns) |
| Wikipedia "United States license plate designs and serial formats" | CC BY-SA | **AVOID copying**; formats themselves are facts |
| jonnii/platekit | MIT, but SVG rendering, no serial patterns | not useful |
- No MIT/CC0 per-state pattern list found. Recommendation: hand-author a per-state table (facts); current standard passenger formats as listed by Wikipedia (verify against DMV pages before use): AL `0AXXXXX`/`00AXXXX`, AK `ABC 123`, AZ `XXX 1XX`, AR `ABC 12D`, CA `1ABC123` (digit-3 letters-3 digits), CO `ABC-D12`, CT `AB·12345`, DE `123456`, DC `AB-1234`, FL `ABC D12`, GA `ABC1234`, HI `ABC 123`, ID `A 1234U`, IL `AB 12345`, IN `123ABC`, IA `ABC 123`, KS `1234ABC`, KY `ABC123`, LA `123 ABC`, ME `123·ABC`, MD `1AB2345`, MA `1ABC 23`, MI `ABC 1234`, MN `ABC-123`, MS `ABC 123`, MO `AB1 C2D`, MT `0-AB1234`, NE `ABC 123`, NV `123·A45`, NH `123 4567`, NJ `D12-ABC`, NM `123-ABC`, NY `ABC-1234`, NC `ABC-1234`, ND `123 ABC`, OH `ABC 1234`, OK `ABC-123`, OR `123 ABC`, PA `ABC1234`, RI `1AB 234`, SC `123ABC`, SD `0A1 234`, TN `ABC 1234`, TX `ABC-1234`, UT `A12 3BC`, VT `ABC 123`, VA `ABC-1234`, WA `ABC1234`, WV `X1A 2345`, WI `ABC-1234`, WY `1A-123A`. Most states omit I/O/Q from serials (state-specific, UNVERIFIED per state).
## 6. Airline
### OurAirports
- https://ourairports.com/data/ — "All data is released to the Public Domain"; mirror repo https://github.com/davidmegginson/ourairports-data is **Unlicense** (GitHub API). Nightly updates. **EMBED**.
- Raw URLs: `https://davidmegginson.github.io/ourairports-data/{airports,airport-frequencies,runways,navaids,countries,regions}.csv`.
| File | Rows | Fields |
|---|---|---|
| airports.csv (12.7 MB) | **86,089**; by type: small_airport 42,734, heliport 23,216, closed 13,524, medium 4,106, seaplane_base 1,273, large 1,174, balloonport 62 | `id,ident,type,name,latitude_deg,longitude_deg,elevation_ft,continent,iso_country,iso_region,municipality,scheduled_service,icao_code,iata_code,gps_code,local_code,home_link,wikipedia_link,keywords` |
| — with `iata_code` | **9,055** (large 1,171, medium 3,398, small 4,233, seaplane 153, heliport 100); `scheduled_service=yes` 4,335; US with IATA 2,036 (873 large/medium) | Suggested subset: `iata_code != '' AND type IN (large_airport, medium_airport)` → 4,569 rows |
| countries.csv | **249** | `id,code,name,continent,wikipedia_link,keywords` |
| regions.csv | **3,987** | `id,code,local_code,name,continent,iso_country,wikipedia_link,keywords` (code = ISO 3166-2, e.g. `US-KS`) |
- No airline file. Derived: datasets/airport-codes on datahub (PDDL) — no need, use upstream.
### OpenFlights
- https://openflights.org/data.php — airports/airlines/routes/planes under **ODbL 1.0** (+ DbCL); airline/plane data partly from Wikipedia (GFDL/CC BY-SA). 5,888 airlines (2012). **AVOID** (share-alike). npm `airline-codes` (npow, ISC) is a straight OpenFlights sync → inherits ODbL → **AVOID**.
### Airline IATA/ICAO codes — open options
| Source | Licence | Count | Verdict |
|---|---|---|---|
| Wikidata SPARQL (https://query.wikidata.org/sparql) | **CC0** (https://www.wikidata.org/wiki/Wikidata:Licensing) | **2,928** items with IATA code (P229), **3,644** with ICAO code (P230) | **EMBED**. Query: `SELECT ?item ?itemLabel ?iata ?icao ?callsign WHERE { ?item wdt:P229 ?iata . OPTIONAL { ?item wdt:P230 ?icao } OPTIONAL { ?item wdt:P432 ?callsign } SERVICE wikibase:label { bd:serviceParam wikibase:language "en" } }`. Includes defunct airlines — filter `MINUS { ?item wdt:P576 ?dissolved }` and/or P31 = Q46970 (airline). |
| OpenTravelData `optd_airlines.csv` (https://github.com/opentraveldata/opentraveldata) | **CC BY 4.0** | 1,620 rows; `^`-separated; fields `pk^env_id^validity_from^validity_to^3char_code^2char_code^num_code^name^name2^alliance_code^alliance_status^type^wiki_link^flt_freq^alt_names^bases^key^version^parent_pk_list^successor_pk_list` | **EMBED+ATTR** (attribution notice required) |
- Format rules: IATA designator = 2 alphanumeric chars (`[A-Z0-9]{2}`, at least one letter in practice; a third optional char exists in the standard but is unused); "controlled duplicates" share a code between non-overlapping regional carriers. ICAO designator = 3 letters `[A-Z]{3}`, unique. Accounting/prefix code = 3 digits (ticket number prefix, e.g. `016` United).
### Aircraft types
| Source | Licence | Verdict |
|---|---|---|
| ICAO Doc 8643 (https://www.icao.int/operational-safety/doc-8643-aircraft-type-designators) | ICAO copyright notice: "None of the materials … may be used, reproduced or transmitted … without permission in writing from ICAO"; API Data Service is paid (25 free trial calls) | **AVOID** — including GitHub mirrors such as ColtJD45/icao-aircraft-designator-list (MIT-labelled, 7,388 rows, fields `manufacturer,model,type_designator,description,engine_type,engine_count,wtc`) because the underlying data is scraped ICAO content (UNVERIFIED provenance) |
| OpenSky aircraft database (https://opensky-network.org/data/aircraft; CSV https://s3.opensky-network.org/data-samples/metadata/aircraftDatabase.csv) | "The aircraft database is unlicensed and does not fall under our terms of use … offered as is"; built from registries, openflights.org and ICAO Doc 8643; citation requested; page says it is no longer up to date | **AVOID** ("unlicensed" ≠ open; upstream includes ODbL/ICAO). Fields for reference: `icao24,registration,manufacturericao,manufacturername,model,typecode,serialnumber,linenumber,icaoaircrafttype,operator,operatorcallsign,operatoricao,operatoriata,owner,…,engines,…,categoryDescription` |
| Wikidata | **CC0** | **550** items with ICAO aircraft type designator (P8305); also IATA aircraft code P9040 (UNVERIFIED property id), manufacturer P176 | **EMBED**. Query: `SELECT ?item ?itemLabel ?icao ?mfrLabel WHERE { ?item wdt:P8305 ?icao . OPTIONAL { ?item wdt:P176 ?mfr } SERVICE wikibase:label { bd:serviceParam wikibase:language "en" } }` |
- Format: ICAO type designator 2–4 alphanumerics (`[A-Z0-9]{2,4}`, e.g. `B738`, `A20N`, `E190`); IATA aircraft code 3 alphanumerics (`738`, `32N`).
### Airport / flight / seat format rules
- IATA airport code `[A-Z]{3}`; ICAO location indicator `[A-Z0-9]{4}` (US `K…`, Canada `C…`, UK `EG…`, Sweden `ES…`). Use OurAirports rows so codes are real.
- Flight number: 2-char IATA designator + 1–4 digits, no leading zeros in display (`QF9`, `AA1234`); systems cap at 0001–9999. Conventions (not rules): 1–999 mainline, 3000–5999 regional affiliates, ≥6000 codeshares, 8xxx charters, ≥9000 ferry/positioning. ICAO form: 3-letter designator + 1–4 alphanumerics (`AFR1`). Source: https://en.wikipedia.org/wiki/Flight_number.
- Seat: `row + letter`. Rows typically 1–~60 (A380 up to ~90); many airlines skip 13 (and 14/17 on some carriers). Letters `A–K` **skipping I** (looks like 1/l); narrow-body 3-3 = `ABC DEF` (A/F windows, C/D aisles); wide-body 3-4-3 = `ABC DEFG HJK`, 2-4-2 = `AB DEFG JK`. Regex: `^[1-9][0-9]?[A-HJK]$`.
+149
View File
@@ -0,0 +1,149 @@
# Part 3 — Science/nature, Text, Dates: open data sources for an MIT fake-data library
Researched 2026-09-17. "OK to embed" = redistributable inside an MIT repo with the stated notice. Anything not fetched first-hand is marked **UNVERIFIED**.
## 7. Science / nature
### Periodic table
| Source | Licence | Embed in MIT repo | Fields | Rows | Format |
|---|---|---|---|---|---|
| PubChem periodic table — CSV https://pubchem.ncbi.nlm.nih.gov/rest/pug/periodictable/CSV, JSON https://pubchem.ncbi.nlm.nih.gov/rest/pug/periodictable/JSON (page: https://pubchem.ncbi.nlm.nih.gov/periodic-table/) | US-government work, public domain. NLM policy (https://www.ncbi.nlm.nih.gov/home/about/policies/): "Information that is created by or for the US government on this site is within the public domain … may be freely distributed and copied. However, it is requested that in any subsequent use of this work, NLM be given appropriate acknowledgment." | Yes; add an acknowledgment line ("Element data: PubChem/NLM"). | `AtomicNumber, Symbol, Name, AtomicMass, CPKHexColor, ElectronConfiguration, Electronegativity, AtomicRadius, IonizationEnergy, ElectronAffinity, OxidationStates, StandardState, MeltingPoint, BoilingPoint, Density, GroupBlock, YearDiscovered` (17) | 118 (Z=1..118, verified) | CSV; JSON is `{Table:{Columns:{Column:[...]},Row:[{Cell:[...]}]}}` |
| NIST periodic table https://www.nist.gov/pml/periodic-table-elements | Public domain (17 USC 105); NIST asks for citation (https://www.nist.gov/open/copyright-fair-use-and-licensing-statements-srd-data-software-and-technical-series-publications) | Yes, but | PDF only (no CSV/JSON) — not usable as a data source | 118 | PDF |
| Bowserinator/Periodic-Table-JSON https://github.com/Bowserinator/Periodic-Table-JSON | **CC BY-SA 3.0** (LICENSE.md, verified; data derived from Wikipedia) | **No** — ShareAlike is incompatible with MIT redistribution | 33 keys incl. `name, symbol, number, atomic_mass, category, phase, period, group, block, shells, electron_configuration, summary, cpk-hex, …` | **119** (includes hypothetical Ununennium) | JSON, CSV |
| faker-js `src/locales/en/science/chemical_element.ts` https://github.com/faker-js/faker (branch `next`) | MIT (LICENSE verified) | Yes, keep MIT notice | `{symbol, name, atomicNumber}` | 118 | TS |
Validation rules: `AtomicNumber` 1..118 unique; `Symbol` 1–2 letters, first uppercase, unique; `Name` unique. PubChem leaves numeric fields empty for Z≥104 where unknown; `ElectronConfiguration` for Og is `"[Rn]7s2 7p6 5f14 6d10 (predicted)"` and `StandardState` is `"Expected to be a Gas"` — strip "(predicted)"/"Expected to be a" if you want an enum. `AtomicMass` for unstable elements is the mass number of the most stable isotope as a bare decimal (e.g. `295.216`), no brackets. `CPKHexColor` is 6 hex digits without `#`, empty for some.
### SI units
- Source: BIPM SI Brochure, 9th ed. (2019; current revision v4.01, 2026) https://www.bipm.org/en/publications/si-brochure ; PDF https://www.bipm.org/documents/20126/41483022/SI-Brochure-9-EN.pdf
- Licence: **CC BY 4.0** — stated on the brochure page and on BIPM's copyright page (https://www.bipm.org/en/copyright): content may be adapted, copied, used in commercial products with "appropriate acknowledgement of the BIPM and its source"; BIPM name/logo must not imply endorsement. Verified via the HTML pages; the PDF's own front-matter licence line is **UNVERIFIED** (could not text-extract the PDF here).
- Embed: Yes — the unit *names/symbols* are facts (not copyrightable); attribute "SI Brochure, BIPM, CC BY 4.0" anyway.
- Data (cross-checked against https://en.wikipedia.org/wiki/International_System_of_Units):
- 7 base units: second s (time), metre m (length), kilogram kg (mass), ampere A (electric current), kelvin K (thermodynamic temperature), mole mol (amount of substance), candela cd (luminous intensity).
- 22 coherent derived units with special names: radian rad, steradian sr, hertz Hz, newton N, pascal Pa, joule J, watt W, coulomb C, volt V, farad F, ohm Ω, siemens S, weber Wb, tesla T, henry H, degree Celsius °C, lumen lm, lux lx, becquerel Bq, gray Gy, sievert Sv, katal kat.
- 24 prefixes: quetta Q 10^30, ronna R 10^27, yotta Y 10^24, zetta Z 10^21, exa E 10^18, peta P 10^15, tera T 10^12, giga G 10^9, mega M 10^6, kilo k 10^3, hecto h 10^2, deca da 10^1, deci d 10^-1, centi c 10^-2, milli m 10^-3, micro µ 10^-6, nano n 10^-9, pico p 10^-12, femto f 10^-15, atto a 10^-18, zepto z 10^-21, yocto y 10^-24, ronto r 10^-27, quecto q 10^-30.
- Format rules for valid output: unit symbols are case-sensitive (`s` vs `S`, `m` vs `M`); prefix and unit symbol are joined without a space (`kJ`, `µs`); a space separates number and unit (`12 kg`), except the degree/minute/second of plane angle; `kg` already carries a prefix — never `µkg`; micro is Greek mu U+03BC (U+00B5 micro sign is the legacy compatibility character); ohm is U+03A9 (U+2126 OHM SIGN is deprecated); `°C` not `° C`.
- MIT alternative: faker-js `en/science/unit.ts` — 29 `{name, symbol}` objects, MIT.
### Animals by class
- **Wikidata** (https://www.wikidata.org/wiki/Wikidata:Licensing): all structured data **CC0** — no attribution required. Embed: Yes.
- Feasibility, live SPARQL (https://query.wikidata.org/sparql), pattern `?t wdt:P31 wd:Q16521; wdt:P105 wd:Q7432; wdt:P171* wd:<class>; wdt:P1843 ?cn FILTER(LANG(?cn)='en')` (taxon, rank species, parent-taxon chain, English common name):
- Mammalia Q7377: **5,759** species with ≥1 English common name.
- Aves Q5113: **12,432**.
- Reptilia Q10811: **16,369** — suspiciously large; Wikidata's parent chain likely pulls birds/Sauropsida in. Use Squamata Q122422 + Testudines Q223044 + Crocodilia Q1387 (or check the chain) instead.
- Insecta Q1390 and Plantae Q756: **timed out (502/504) every attempt** — trees too big for the 60 s public endpoint. Options: paginate by order (`P171` one level at a time), use the weekly JSON dump, or take GBIF instead.
- Sample rows (mammals): `Notoryctes typhlops` → "Central Desert Marsupial Mole", "Itjaritjari", "Marsupial Mole", "Southern marsupial mole", "Southern Marsupial Mole" — i.e. **several P1843 values per taxon, inconsistent casing, and indigenous-language names tagged `en`**. Validation: pick one name per taxon (prefer the one matching the item's English `rdfs:label`), normalise case, drop names that equal the scientific name.
- **GBIF Backbone Taxonomy** https://www.gbif.org/dataset/d7dddbf4-2cf0-4f39-9b2a-bb099caae36c — licence **CC BY 4.0** (verified via https://api.gbif.org/v1/dataset/d7dddbf4-2cf0-4f39-9b2a-bb099caae36c, `license: http://creativecommons.org/licenses/by/4.0/legalcode`); required citation text: "GBIF Secretariat (2023). GBIF Backbone Taxonomy. Checklist dataset https://doi.org/10.15468/39omei". Embed: Yes with that citation. Vernacular names come via `/v1/species/{key}/vernacularNames`, each row from a *different* source checklist with its own licence (some CC BY-NC) — record and check per row. GBIF's site terms page (https://www.gbif.org/terms) was unreachable (timeout) — **UNVERIFIED** beyond the API field.
- **faker-js** `src/locales/en/animal/*` (MIT; counts = quoted string lines on branch `next`): bird 821, snake 533, dog 493, cow 467, horse 342, rodent 161, insect 129, fish 95, cat 55, cetacean 52, rabbit 49, type 44 (generic: "bat, bear, bee, …, zebra"), pet_name 42, crocodilia 24, bear 8, lion 7. These are breeds/species mixed with common names, no class field; fine as MIT drop-ins. (A WebFetch summariser claimed 1,036 dogs; the regex count is 493 — treat the exact count as UNVERIFIED.)
### Plants
- Same Wikidata approach (CC0) — but the Plantae query timed out; paginate by order or family (`?t wdt:P171 ?family . ?family wdt:P105 wd:Q35409`) or use the dump.
- GBIF Plantae (kingdomKey 6) vernacular names, CC BY 4.0, same per-row-licence caveat.
- faker-js has **no** plant list (en locale dirs: airline, animal, app, book, color, commerce, company, database, date, finance, food, hacker, internet, location, lorem, medical, music, person, phone_number, science, team, vehicle, word).
- USDA PLANTS database (public domain, has common names) — **UNVERIFIED**, not fetched.
### Planets / moons
- NASA NSSDCA Planetary Fact Sheet https://nssdc.gsfc.nasa.gov/planetary/factsheet/ — 10 columns (Mercury, Venus, Earth, Moon, Mars, Jupiter, Saturn, Uranus, Neptune, Pluto) × 20 rows (mass 10^24 kg, diameter km, density, gravity, escape velocity, rotation period, day length, distance from Sun, perihelion, aphelion, orbital period, orbital velocity, inclination, eccentricity, obliquity, mean temp °C, surface pressure, number of moons, rings Y/N, magnetic field Y/N). HTML table only. NASA content "generally not subject to copyright in the United States" (https://www.nasa.gov/nasa-brand-center/images-and-media/); NASA asks to be acknowledged; the NASA insignia must not be used. Embed: Yes with "Source: NASA NSSDCA".
- Wikidata (CC0): planets = `?p wdt:P31/wdt:P279* wd:Q634; wdt:P397 wd:Q525` → Mercury Q308, Venus Q313, Earth Q2, Mars Q111, Jupiter Q319, Saturn Q193, Uranus Q324, Neptune Q332 (plus hypothetical Theia Q1053432 — filter out). Moons = `?m wdt:P397 <planet>; wdt:P31/wdt:P279* wd:Q2537` (natural satellite; note **Q2199 is "dwarf planet"**, not moon): Saturn 165, Jupiter 85, Uranus 29, Neptune 16, Earth 5, Mars 2. Caveats: Wikidata lags IAU counts (Saturn has 274 recognised moons as of 2025 — UNVERIFIED figure); Earth's 5 includes quasi-satellites; provisional designations (`S/2004 S 3`) are labels too — filter `^S/\d{4}` if you want proper names only.
### Blood type distribution
- **US** (Stanford Blood Center https://stanfordbloodcenter.org/donate-blood/blood-donation-facts/blood-types/, verified; cites AABB Technical Manual 18th ed.): O+ 37.4 %, A+ 35.7 %, B+ 8.5 %, AB+ 3.4 %, O− 6.6 %, A− 6.3 %, B− 1.5 %, AB− 0.6 % (sums to 100.0). Facts — freely embeddable; cite Stanford/AABB.
- American Red Cross https://www.redcrossblood.org/donate-blood/blood-types.html — **UNVERIFIED**: site returns HTTP 403 (Akamai) to both WebFetch and the headless browser. Search snippets say the Red Cross gives type O by ethnicity (Latino 57 %, African American 51 %, Caucasian 45 %) rather than an 8-way national table.
- **Global**: no authoritative figure found. Wikipedia "Blood type distribution by country" gives a population-weighted row (O+ 38.4, A+ 27.3, B+ 8.1, AB+ 2.0, O− 13.1, A− 8.1, B− 2.0, AB− 0.01) that its own editors flag "unreliable source"; WorldAtlas gives O+ 42, A+ 31, B+ 15, AB+ 5, O− 3, A− 2.5, B− 1, AB− 0.5 with no source. Recommend shipping only the US table, or per-country tables from Wikipedia (facts; Wikipedia prose is CC BY-SA but the numbers are cited to national blood services — cite those).
- Validation: ABO ∈ {A, B, AB, O}; Rh ∈ {+, −}; weights must sum to 100 ± rounding.
## 8. Text
### English word lists with part of speech
| Source | Licence | Embed | Counts | Format |
|---|---|---|---|---|
| Princeton WordNet 3.0 / 3.1 https://wordnet.princeton.edu/license-and-commercial-use ; download https://wordnetcode.princeton.edu/wn3.1.dict.tar.gz | "WordNet License" (SPDX `WordNet`), BSD-style: permission to use/copy/modify/distribute "for any purpose and without fee or royalty … provided that … the following copyright notice and statements, including the disclaimer … appear on ALL copies … including modifications". Notice text: "WordNet 3.0 Copyright 2006 by Princeton University. All rights reserved." + AS-IS disclaimer; no use of Princeton's name in advertising. | Yes — ship the LICENSE text alongside the data | WordNet 3.0 (wnstats(7WN), verified): unique strings noun 117,798 / verb 11,529 / adj 21,479 / adv 4,481 (total 155,287); synsets 82,115 / 13,767 / 18,156 / 3,621 | WNDB: `index.noun|verb|adj|adv` (one lemma per line: `lemma pos synset_cnt p_cnt [ptr…] sense_cnt tagsense_cnt synset_offset…`), `data.<pos>` (synsets with glosses). Multi-word lemmas use `_` for spaces; lower-case; entries can contain digits/apostrophes/hyphens. |
| Open English WordNet 2025 https://en-word.net/ ; https://github.com/globalwordnet/english-wordnet | **CC BY 4.0**; cite "Open English WordNet … derived from Princeton WordNet" | Yes, with attribution | 135,969 words, 107,519 synsets (README, verified) | `english-wordnet-2025.zip` (WNDB, 9.2 MB), `english-wordnet-2025.xml.gz` (WN-LMF, 10.8 MB), `english-wordnet-2025-json.zip` (9.5 MB), `.ttl.gz` (16.9 MB) |
| SCOWL v2 / ESDB https://github.com/en-wl/wordlist (branch v2) ; classic http://wordlist.aspell.net/ | Custom MIT-like (Copyright 2000-2026 Kevin Atkinson: use/copy/modify/distribute/sell "provided that the above copyright notice appears in all copies and that both the above copyright notice and this notice appear in supporting documentation"). Sources 12dicts + ENABLE2K are public domain; lists above size 80 add the UKACD notice; **POS assignments partly from WordNet, so the WordNet notice "MIGHT apply"** (their words). Australian-English parts carry an extra notice. | Yes, ship the Copyright file | Sizes 35 (small), 50 (medium), 60 (spell-check default), 70 (large), 80 (incl. game words), 85 (archaic). No per-size counts published in the v2 README. Classic SCOWL sizes 10–95 — counts **UNVERIFIED** (README not reachable). | v2: `scowl.db` (SQLite), `scowl.txt`, Python `libscowl`. ESDB rows carry POS, spelling (A/B/Z/C/D) and region codes. |
| 12dicts http://wordlist.aspell.net/12dicts-readme/ | Public domain ("I explicitly release them to the public domain, but request acknowledgment"), **except** `2of12inf` and the `2+2+3*` lists, which depend on AGID and inherit its terms | Yes (PD lists); acknowledge Alan Beale | 3esl ≈22k, 6of12 ≈32k, 2of12 ≈41k, 2of12inf ≈82k, 3of6game ≈65k, 5d+2a ≈68k, 3of6all ≈83k, 2+2+3lem ≈84k, 2of5core ≈4.7k, 6phrase ≈22k, neol2016 ≈600 | Plain text, one word per line; marker suffixes `+ ! ^ &` (not POS). **No POS tags** in any current 12dicts list. |
| Moby Part-of-Speech II https://www.gutenberg.org/ebooks/3203 ; file https://www.gutenberg.org/files/3203/files/mobypos.txt | Public domain ("Public Domain material by grant from the author, January, 2001", in the package README) | Yes, no notice needed | **233,356** lines (verified) | One entry per line: `word\CODES`, CRLF line ends (e.g. `A-line\NA`, `a tempo\Avh`). Codes in priority order: `N` noun, `p` plural, `h` noun phrase, `V` verb (usu. participle), `t` transitive verb, `i` intransitive verb, `A` adjective, `v` adverb, `C` conjunction, `P` preposition, `!` interjection, `r` pronoun, `D` definite article, `I` indefinite article, `o` nominative. Many entries are phrases, proper nouns, or 1990s-era; non-ASCII entries use a legacy 8-bit encoding (e.g. `a bon march\v` lost its `é`) — restrict to ASCII `[a-z]+` for a clean list. |
| faker-js `en/word/*` (MIT) | MIT | Yes | adjective 1000, noun 1000, verb 1000, adverb 325, preposition 109, conjunction 51, interjection 46 | TS string arrays |
### Word frequency lists
| Source | Licence | Verdict |
|---|---|---|
| Google Books Ngram data v2/v3 https://storage.googleapis.com/books/ngrams/books/datasetsv3.html | **CC BY 3.0** ("This compilation is licensed under a Creative Commons Attribution 3.0 Unported License") | Embeddable with attribution. Format `ngram TAB year TAB match_count TAB volume_count`; 1-gram files split by leading letter, GBs each — aggregate offline to a top-N list. |
| wordfreq https://github.com/rspeer/wordfreq | Code **Apache 2.0** (not MIT). Data: "may be redistributed under a Creative Commons Attribution-ShareAlike 4.0 license" (mix of Google Books, Wikipedia, OpenSubtitles, SUBTLEX, Leeds, ParaCrawl, Twitter); maintainer notes the CSV export "does not follow the CC-By-SA license". Project in sunset mode (data snapshot ≈2021). | **Avoid** embedding the data (ShareAlike). |
| Peter Norvig count_1w.txt https://norvig.com/ngrams/ | Page says "Code … under the MIT license"; the **data** comes from the Google Web 1T corpus distributed by LDC under LDC terms — no data licence stated. | **UNVERIFIED / avoid**. 333,333 words, `word TAB count`, lowercase. |
| hermitdave/FrequencyWords https://github.com/hermitdave/FrequencyWords | "MIT License for code. CC-by-sa-4.0 for content." | **Avoid** (ShareAlike). |
### Lorem ipsum
- Canonical paragraph (Letraset, 1966; public domain — garbled Cicero): "Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum."
- Source: Cicero, *De finibus bonorum et malorum* 1.10.32–33 (45 BC), public domain; the "dolorem ipsum" fragment begins "Neque porro quisquam est qui dolorem ipsum quia dolor sit amet, consectetur, adipisci velit…". Reference: https://en.wikipedia.org/wiki/Lorem_ipsum
- MIT word pool alternative: faker-js `en/lorem/word.ts` — 999 words.
### Quote collections
- Project Gutenberg #27889 https://www.gutenberg.org/ebooks/27889 — Bartlett, *Familiar Quotations*, **9th edition** (title page "NINTH EDITION", copyright 1875/1882/1891/1903) — not the 1919 10th. "Public domain in the USA." Plain text 3.1 MB, HTML zip 50 MB, EPUB.
- Project Gutenberg #16732 https://www.gutenberg.org/ebooks/16732 — an early (Hurst & Co. reprint) edition, no edition number given. PD in USA.
- 10th edition (1914/1919, ed. Nathan Haskell Dole) exists on archive.org (https://archive.org/details/bartlettsfamilia0000john_o4b4) — not found on Gutenberg; **UNVERIFIED** availability as clean text.
- Either Bartlett text needs parsing (quote / author / source headings); no structured file. Wikiquote is CC BY-SA — avoid.
- Validation: strip Gutenberg header/footer (`*** START OF THE PROJECT GUTENBERG EBOOK` … `*** END …`) before extraction; Gutenberg's trademark terms only bind if you keep the "Project Gutenberg" name on the output — strip it.
### Colour names
| Source | Licence | Embed | Count / format |
|---|---|---|---|
| CSS Color Module Level 4 §named colors https://www.w3.org/TR/css-color-4/#named-colors | W3C Document License (2023) https://www.w3.org/copyright/document-license-2023/ — copy/distribute allowed for any purpose; derivatives allowed only "to facilitate implementation of the technical specifications"; notice required: "Copyright © 2023 W3C®. This software or document includes material copied from or derived from [title and URI]." | Yes — the name→hex table is factual data and is copied into every browser; include the W3C notice line with the spec title + URI. | **148** names (verified: aliceblue … yellowgreen, incl. `rebeccapurple #663399`); all names lowercase ASCII, hex 6 lowercase digits. |
| XKCD colour survey https://xkcd.com/color/rgb.txt | **CC0** (file header: `# License: https://creativecommons.org/publicdomain/zero/1.0/`) | Yes, no notice needed | **949** lines (verified). Format `name TAB #rrggbb TAB` (note the trailing tab). Max name length 26; all hex valid lowercase; no duplicate names. Contains crude names ("poo", "baby poop green", "diarrhea", "vomit") — apply a blocklist. |
| faker-js `en/color/human.ts` | MIT | Yes | 31 names |
| Pantone | Proprietary | **Avoid** | — |
## 9. Dates
### IANA tz database
- Source https://www.iana.org/time-zones ; files https://data.iana.org/time-zones/tzdb/ (current version **2026d**, verified). LICENSE: "Unless specified below, all files in the tz code and data (including this LICENSE file) are in the public domain" (only `date.c`, `newstrftime.3`, `strftime.c` are BSD-3). Embed: Yes, no notice needed.
- `zone1970.tab` — **312** data rows; tab-separated; UTF-8; `#` comments. Columns: (1) comma-separated ISO 3166-1 alpha-2 codes of countries overlapping the zone, most-populous first; (2) coordinates in ISO 6709 `±DDMM±DDDMM` or `±DDMMSS±DDDMMSS` (lat then lon, no separator); (3) TZ name (`Europe/Zurich`); (4) comment, present iff the country has several zones. One row per timezone where civil time has agreed since 1970.
- `zone.tab` — **418** rows, older one-country-per-row table (column 1 is a single country code); kept for backward compatibility. Prefer `zone1970.tab`.
- `backward` — **252** `Link TARGET LINK-NAME` lines mapping old/merged names (`US/Eastern`, `Asia/Calcutta`, `Europe/Kiev`) to current ones; also a few `Zone` entries. Use it to (a) accept aliases as input and (b) never emit them as output.
- UTC offsets: do **not** embed — offsets depend on DST rules that change several times a year. Derive at runtime: `new Intl.DateTimeFormat('en-US', { timeZone, timeZoneName: 'longOffset' }).formatToParts(date)` → part `timeZoneName` = `GMT-05:00` (or `GMT` for zero) — strip `GMT`, treat bare `GMT` as `+00:00`. Validation of a generated zone name: `Intl.supportedValuesOf('timeZone')` (canonical names only; whether links like `US/Eastern` appear depends on the ICU build) or `try { new Intl.DateTimeFormat(undefined, { timeZone }) } catch { invalid }`, which accepts links too.
### ISO 8601 layouts (https://en.wikipedia.org/wiki/ISO_8601)
| Layout | Extended | Basic | Example |
|---|---|---|---|
| Calendar date | `YYYY-MM-DD` | `YYYYMMDD` | 2009-01-06 |
| Ordinal date | `YYYY-DDD` | `YYYYDDD` | 1981-095 |
| Week date | `YYYY-Www-D` | `YYYYWwwD` | 2009-W01-1 (weeks start Monday, week 1 contains 4 Jan) |
| Time | `hh:mm:ss[.fff]` | `hhmmss` | 13:47:30 (24 h; `T` prefix optional standalone) |
| UTC / offset | `Z` or `±hh:mm` | `±hhmm`, `±hh` | 2007-04-05T14:30Z, 2024-06-01T09:00-05:00 |
| Date-time | date `T` time offset | | 2007-04-05T14:30:00+02:00 |
| Duration | `PnYnMnDTnHnMnS` / `PnW` | | P3Y6M4DT12H30M5S |
| Interval | start`/`end, start`/`duration, duration`/`end | | 2007-03-01T13:00:00Z/2008-05-11T15:30:00Z |
| Recurring | `Rn/`interval | | R5/2008-03-01T13:00:00Z/P1Y2M10DT2H30M |
RFC 3339 profile (what JSON/APIs usually mean): extended forms only, `T`/`Z` may be lowercase, space allowed instead of `T`, `-00:00` = unknown offset, no durations/intervals/week/ordinal dates. Generate RFC 3339 by default.
### Unicode CLDR (locale month/weekday names, date patterns)
- Licence: **Unicode License v3** (SPDX `Unicode-3.0`), OSI-approved 2023-11-17 (https://opensource.org/license/unicode-3-0). Text https://www.unicode.org/license.txt — MIT-style: use/copy/modify/sell "provided that either (a) this copyright and permission notice appear with all copies of the Data Files or Software, or (b) this copyright and permission notice appear in associated Documentation"; no use of the Unicode name in advertising. Embed: Yes; put the notice in LICENSE/NOTICE.
- npm: `cldr-core`, `cldr-dates-full`, `cldr-numbers-full` (peer of dates), etc.; current **48.2.0** (CLDR 48, verified via npm registry, `license: Unicode-3.0`); repo https://github.com/unicode-org/cldr-json .
- Path is **`cldr-dates-full/main/<locale>/ca-gregorian.json`** (not `dates/gregorian.json`), rooted at `main.<locale>.dates.calendars.gregorian`. Keys (verified for `en`): `months`, `days`, `quarters`, `dayPeriods`, `eras`, `dateFormats`, `dateSkeletons`, `timeFormats`, `timeSkeletons`, `dateTimeFormats`, `dateTimeFormats-atTime`, `dateTimeFormats-relative`.
- `months.{format|stand-alone}.{abbreviated|narrow|wide}` keyed `"1".."12"`; `days.{format|stand-alone}.{abbreviated|narrow|short|wide}` keyed `sun..sat`.
- `dateFormats`: en = full `EEEE, MMMM d, y`, long `MMMM d, y`, medium `MMM d, y`, short `M/d/yy`.
- `timeFormats`: en = full `h:mm:ss a zzzz`, long `h:mm:ss a z`, medium `h:mm:ss a`, short `h:mm a` — **the `a` is preceded by U+202F NARROW NO-BREAK SPACE**; `*-alt-ascii` variants use a plain space. Pick one spelling (ASCII) and ignore the other.
- `dateTimeFormats.{full|long|medium|short}` = glue like `{1}, {0}` ({1} date, {0} time); `availableFormats` = skeleton→pattern map (`yMMMd` → `MMM d, y`), plus `intervalFormats`, `appendItems`.
- `dayPeriods.format.abbreviated`: `am: AM`, `pm: PM`, plus `midnight`, `noon`, `morning1`… variants.
- Pattern letters follow LDML (https://unicode.org/reports/tr35/tr35-dates.html#Date_Field_Symbol_Table): `y` year, `M`/`L` month (format/stand-alone), `d` day, `E`/`c` weekday, `h` 1–12, `H` 0–23, `a` AM/PM, `z`/`zzzz` zone name, `x`/`X` ISO offset; literal text in single quotes. Same patterns drive `Intl.DateTimeFormat` internally, so runtime `Intl` can replace embedding for month/weekday names if bundle size matters.
## Summary of licence verdicts
- Embed freely (PD/CC0): PubChem periodic table, NIST, NASA fact sheet, Wikidata, Moby POS, 12dicts (PD lists), Bartlett via Gutenberg, XKCD colours, IANA tz, Lorem ipsum/Cicero.
- Embed with notice: WordNet 3.x (WordNet License), Open English WordNet (CC BY 4.0), SCOWL/ESDB (custom MIT-like), GBIF backbone (CC BY 4.0 + citation), BIPM SI (CC BY 4.0), CSS named colours (W3C notice), CLDR (Unicode-3.0), Google Books ngrams (CC BY 3.0), faker-js lists (MIT).
- Avoid: Bowserinator periodic table (CC BY-SA 3.0), wordfreq data (CC BY-SA 4.0), hermitdave (CC BY-SA 4.0), Norvig count_1w (LDC-derived, unstated), Wikiquote (CC BY-SA), Pantone.
- UNVERIFIED: American Red Cross figures (site blocks fetches), any global blood-type table, BIPM PDF front-matter licence line, classic SCOWL per-size counts, USDA PLANTS, IAU moon counts, exact faker dog count, Bartlett 10th-edition text availability.
+139
View File
@@ -0,0 +1,139 @@
# Part 4 — ISO lists, commerce, books/music/food
Researched 2026-09-17. "Embed?" = can the data ship inside an MIT-licensed repo. Anything marked UNVERIFIED was not confirmed against the primary source.
## 10. ISO lists
### ISO 3166-1 countries
| Source | Licence | Embed in MIT repo? | Rows | Fields | Format |
|---|---|---|---|---|---|
| [datasets/country-codes](https://github.com/datasets/country-codes) | PDDL (README: "Public Domain Dedication and License"; GitHub API detects no licence file) | Yes, no attribution required. README caveat: ISO itself says its list is "for internal use and non-commercial purposes free of charge" — the repo argues no rights subsist in a list of facts | 249 | ISO3166-1-Alpha-2/-Alpha-3/-numeric, official_name_en, UNTERM names (ar/zh/en/fr/ru/es short+formal), Dial, TLD, Capital, Continent, Languages, ISO4217 currency code/name/numeric/minor_unit, Region/Sub-region/Intermediate names+codes (M49), FIPS, IOC, FIFA, MARC, ITU, WMO, GAUL, EDGAR, Geoname ID, wikidata_id, CLDR display name, is_independent, LDC/LLDC/SIDS flags | CSV `data/country-codes.csv` |
| [annexare/Countries](https://github.com/annexare/Countries) (npm `countries-list` 3.4.1, 2026-07-14) | MIT | Yes; keep MIT notice | 252 countries, 185 languages (incl. non-ISO entries like AC, XK) | per country: name, native, phone[] (calling codes), continent, capital, currency[], languages[], optional alias[], partOf; separate ISO 4217 currency table (name, native, symbol, numeric, decimals); languages: name, native | TS source, exported JSON/CSV/SQL |
| [Debian iso-codes](https://salsa.debian.org/iso-codes-team/iso-codes) | LGPL-2.1-or-later (Debian copyright file) | Not safely as embedded data in an MIT repo — LGPL applies to the data files; you would have to ship it as a separately licensed component with LGPL notice. Avoid unless accepting that | 3166-1: 249 | alpha_2, alpha_3, numeric, name, official_name, common_name, flag | JSON `data/iso_3166-1.json` (raw URL needs `?inline=false`) |
| [mledoze/countries](https://github.com/mledoze/countries) | ODbL-1.0 | No — share-alike + attribution; derived databases must be ODbL. Avoid | ~250 | common/official names, cca2/cca3/ccn3/cioc, tld, currencies, idd (calling codes), capital, region/subregion, languages, translations, latlng, borders, area, demonyms, flag emoji, UN member | JSON/CSV/XML/YAML |
| [restcountries](https://gitlab.com/restcountries/restcountries) | MPL-2.0 (source repo; hosted API v5 is proprietary/API-key) | File-level copyleft: a copied data file stays MPL and must carry the notice; OK to ship next to MIT code but not to relicense. Prefer PDDL/MIT sources | 250+ | name, tld, cca2, ccn3, cca3, cioc, fifa, independent, status, unMember, currencies, idd, capital, capitalInfo, altSpellings, region, subregion, continents, languages, translations, latlng, landlocked, borders, area, flag (emoji), demonyms, flags, coatOfArms, population, maps, gini, car, postalCode, startOfWeek, timezones | JSON `src/main/resources/countriesV3.1.json` |
| [lukes/ISO-3166-Countries-with-Regional-Codes](https://github.com/lukes/ISO-3166-Countries-with-Regional-Codes) | CC BY-SA 4.0 (LICENSE.md) | No — share-alike. Avoid | 249 | name, alpha-2, alpha-3, country-code, iso_3166-2, region, sub-region, intermediate-region + codes | CSV/JSON/XML (all, slim-2, slim-3) |
| [stefangabos/world_countries](https://github.com/stefangabos/world_countries) | Data: CC BY-SA 4.0 (README, GitHub licence detection); npm `world_countries_lists` package.json says LGPL-3.0-or-later — contradictory | No — share-alike either way. Avoid | 249 world / 193 UN; subdivisions 5046 | id (numeric), alpha2, alpha3, name (37 languages); subdivisions: country, code, name, name_en, type, parent | JSON/CSV/PHP/SQL/XML |
| [i18n-iso-countries](https://www.npmjs.com/package/i18n-iso-countries) 7.14.0 | MIT | Yes | 78 language files | alpha2, alpha3, numeric, localized names | JSON per language |
| Wikipedia ISO 3166-1 | CC BY-SA 4.0 | No (share-alike) | — | — | — |
| [ISO OBP](https://www.iso.org/obp) | Proprietary; ISO grants free use only for "internal use and non-commercial purposes" (quoted in datasets/country-codes README) | No | — | — | — |
- Flag emoji: derive, no data needed — for alpha-2 `XY`, emit U+1F1E6 + (X − 'A') and U+1F1E6 + (Y − 'A') (regional indicator symbols).
- Recommendation: `datasets/country-codes` (PDDL) covers every requested field (name, alpha2, alpha3, numeric, calling code, TLD, capital, currency, languages) in one CSV; flag emoji derived. `annexare` (MIT) is the fallback for native names/calling codes if PDDL's ISO caveat worries you; it lacks TLD.
- Validation rules: alpha-2 `^[A-Z]{2}$`, alpha-3 `^[A-Z]{3}$`, numeric `^\d{3}$` (zero-padded, keep as string); calling code 1–3 digits, NANP countries share `1`; TLD `^\.[a-z]{2}$` (ccTLD = lowercase alpha-2, exceptions: GB uses `.uk`).
### ISO 639 languages
| Source | Licence | Embed? | Rows | Fields | Format |
|---|---|---|---|---|---|
| [langs](https://www.npmjs.com/package/langs) 2.0.0 (2017, unmaintained) | MIT | Yes | 184 | name, local, 1 (639-1), 2, 2T, 2B, 3 | `data.js` array |
| [iso-639-1](https://www.npmjs.com/package/iso-639-1) 3.1.6 (2026-07-02) | MIT | Yes | 183 | code, name, nativeName | `src/data.js` |
| annexare/Countries languages | MIT | Yes | 185 | code (639-1), name, native | TS/JSON |
| Debian iso-codes 639-2 / 639-3 / 639-5 | LGPL-2.1+ | See above (LGPL) | 639-2: 487; 639-3: 7923 | 639-2: alpha_2, alpha_3, bibliographic, name, common_name; 639-3: + inverted_name, scope, type | JSON |
| [SIL ISO 639-3 tables](https://iso639-3.sil.org/code_tables/download_tables) | SIL terms: may incorporate into software (commercial or not) with attribution to iso639-3.sil.org, must not modify identifiers, and product "does not provide a means to redistribute the code set" | Risky — a git repo with the table IS a means to redistribute; avoid embedding the full table. Codes themselves are facts | 7927 | Id, Part2b, Part2t, Part1, Scope, Language_Type, Ref_Name, Comment | tab-delimited `iso-639-3.tab` |
| [Library of Congress ISO 639-2](https://www.loc.gov/standards/iso639-2/ascii_8bits.html) | US federal government work → public domain in the US | Yes | ~487 (matches Debian 639-2 count) | pipe-delimited: alpha3-B, alpha3-T, alpha2, English name, French name — UNVERIFIED: loc.gov is behind a Cloudflare bot check (WebFetch, curl and Playwright all blocked); field layout from memory | `ISO-639-2_utf-8.txt` |
- Recommendation: `langs` or `iso-639-1` (MIT) for 639-1 names; both are small and derived from public code tables. Anything 639-3-sized is LGPL or SIL-restricted.
- Rules: 639-1 `^[a-z]{2}$`, 639-2/3 `^[a-z]{3}$`; 639-2 has B/T pairs (e.g. `fre`/`fra`, `ger`/`deu`) — pick one consistently (T codes = 639-3).
### ISO 3166-2 subdivisions
| Source | Licence | Embed? | Rows | Fields |
|---|---|---|---|---|
| Debian iso-codes 3166-2 | LGPL-2.1+ | LGPL caveat | 5046 | code, name, type, parent |
| [olahol/iso-3166-2.json](https://github.com/olahol/iso-3166-2.json) | package.json says ISC; no LICENSE file; GitHub detects none; data scraped from eQuest xls (provenance unclear); last push 2021 | No — unclear licence and provenance | 237 countries / 3807 divisions | `{alpha2: {name, divisions: {code: name}}}` |
| [esosedi/3166](https://github.com/esosedi/3166) (npm `iso3166-2-db` 2.3.11) | MIT | Yes; data mixes GeoNames (CC BY 4.0), Wikipedia (CC BY-SA), OSM (ODbL) — the MIT label may not survive scrutiny for the Wikipedia/OSM-derived parts. UNVERIFIED how much is derived from each | "all countries and regions" (count not stated) | ISO 3166-2 + FIPS codes, names in 12 languages, admin level, GeoNames/OSM/Wikipedia/WOF ids |
| stefangabos subdivisions | CC BY-SA 4.0 | No | 5046 | country, code, name, name_en, type, parent |
| [US Census state.txt](https://www2.census.gov/geo/docs/reference/state.txt) | US government work → public domain | Yes | 57 (50 states + DC + 6 territories) | STATE (FIPS), STUSAB, STATE_NAME, STATENS |
- Recommendation: for an en_US library, ship only the US list (Census/USPS, public domain: 50 states + DC, optionally territories) — ISO 3166-2 code = `US-` + USPS abbreviation. Skip the ~5000-row world list; every open compilation is LGPL, CC BY-SA or unclear.
- Rule: `^[A-Z]{2}-[A-Z0-9]{1,3}$`.
## 11. Commerce
### Word lists (product name / adjective / material)
- [@faker-js/faker](https://github.com/faker-js/faker) 10.6.0 — MIT (LICENSE covers code and locale data; copyright Faker contributors 2022-2025 and Marak Squires 2011-2020; acknowledges Ruby Faker and Perl Data::Faker, both MIT). No separate data licence.
- Provenance policy in [CONTRIBUTING.md](https://github.com/faker-js/faker/blob/next/CONTRIBUTING.md): "Faker must not contain copyrighted materials"; facts (finite known lists) are OK; compilations must not be copied from a single source (Wikipedia, "most popular" articles, government sites) because "a compilation of facts can be copyrighted" — contributors are told to draw on multiple sources and their own judgement.
- Verdict: reusing faker's `en` commerce lists (`product_name.adjective/material/product`, `department`) is fine under MIT — keep the MIT notice with the copied data. The lists are short curated words (no external dataset). Nothing in the repo or docs restricts locale-data reuse beyond MIT.
### Garment and shoe sizes
- Letter sizes XS–XXL: EN 13402-3 letter codes with chest/bust ranges (Wikipedia [EN 13402](https://en.wikipedia.org/wiki/EN_13402), CC BY-SA — but the table is a handful of facts; the standard itself is paywalled). Men chest / women bust cm: XXS 70–78 / 66–74, XS 78–86 / 74–82, S 86–94 / 82–90, M 94–102 / 90–98, L 102–110 / 98–107, XL 110–118 / 107–119, XXL 118–129 / 119–131, 3XL 129–141 / 131–143.
- US numeric women's sizes (0–20, even) come from ASTM D5585 (paywalled); no open dataset. Just generate the even-number series.
- Shoe sizes: no open dataset found (search hits are ML demos and retailer charts). Use the formulas from Wikipedia [Shoe size](https://en.wikipedia.org/wiki/Shoe_size) (facts, not copyrightable): EU (Paris point) ≈ 1.5 × foot cm + 2 (i.e. 3/20 × mm + 2); UK adult ≈ 3 × foot in − 23; US men ≈ 3 × foot in − 22 = UK + 1; US women (common) = US men + 1.5 (FIA scale: 3 × foot in − 21); Mondopoint = foot length mm. ISO/TS 19407:2023 has the official conversion tables (paywalled). Generate from a foot length (men 24–31 cm, women 21–27 cm) in half-size steps so all systems agree.
### Check-digit rules (all verified against Wikipedia articles; algorithms are facts)
| Code | Rule |
|---|---|
| EAN-13 / UPC-A / GTIN-8/12/13/14 | Weights 3,1,3,1… starting from the rightmost data digit (the one next to the check digit); check = (10 − sum mod 10) mod 10. UPC-A = EAN-13 with a leading 0 (GS1 prefix of a 12-digit GTIN is `0` + first two digits). GTIN-14 first digit is an indicator 1–8 (packaging level) or 9 (variable measure); 0 is not valid there. |
| UPC-A number system (first digit) | 0,1,6,7,8,9 regular; 2 variable-weight; 3 drugs (NDC); 4 in-store/local; 5 coupons. |
| ISBN-10 | Σ(weights 10..2 × first 9 digits) + check ≡ 0 mod 11; check = (11 − sum mod 11) mod 11, 10 → `X`. |
| ISBN-13 | EAN-13 with prefix 978 or 979. 979 groups in use: 979-8 USA, 979-10 France, 979-11 Korea, 979-12 Italy. 978 groups 0 and 1 = English-language. For an en_US generator: `978-0…`, `978-1…`, `979-8…`. The official range file (`https://www.isbn-international.org/export_rangemessage.xml`) is XML only and the site says "You must not republish material from our site … without our written permission" — do not embed it; group/registrant length only matters for hyphenation, not validity. |
| ISSN | Format `NNNN-NNNC`; weights 8..2 over first 7 digits, sum mod 11; check = 0 if remainder 0 else 11 − remainder; 10 → `X`. EAN-13 form: `977` + 7 ISSN digits (no check) + 2 variant digits + EAN check. |
| IMEI | 15 digits: TAC 8 (RBI 2 + 6) + serial 6 + Luhn check (ISO/IEC 7812; from the right, double every second digit, sum digits, total ≡ 0 mod 10). SVN never enters the check. IMEISV = 14 digits + 2-digit SVN, no check. RBI `00` = test IMEI (Wikipedia [Reporting Body Identifier](https://en.wikipedia.org/wiki/Reporting_Body_Identifier)); live RBIs: 01 CTIA/PTCRB (US), 35 TÜV SÜD/BABT (UK), 86 TAF (China), 99 GHA. For fake data use TAC `00xxxxxx` so it can never match a real handset — UNVERIFIED: third-party pages claim the GSMA TS.06 test-TAC shape is `00 44` + 4 digits; the TS.06 PDF could not be text-extracted here. |
| ASIN | 10 chars `[A-Z0-9]{10}`; non-books start `B0` + 8 alphanumerics (`^B0[A-Z0-9]{8}$`); books reuse their ISBN-10 verbatim. No check digit (Wikipedia). |
| SKU | Not standardised (Wikipedia). Convention only: 8–12 uppercase alphanumerics, e.g. `[A-Z]{3}-\d{4}-[A-Z]{2}`; avoid leading 0 and O/I if you want scanner-safe. |
### GS1 prefixes
- Official table: [gs1.org/standards/id-keys/company-prefix](https://www.gs1.org/standards/id-keys/company-prefix) (rendered via Playwright; WebFetch gets 403). Copyright notice ([gs1.org/terms-use](https://www.gs1.org/terms-use)): reproduction allowed only "in unaltered form … for your personal, non-commercial use or use within your organisation"; "You are not permitted to re-transmit, distribute or commercialise the information or material without seeking prior written approval from GS1." → Do not copy the GS1 page. The prefix-to-country mapping is factual and also on Wikipedia ([List of GS1 country codes](https://en.wikipedia.org/wiki/List_of_GS1_country_codes), CC BY-SA 4.0 — avoid verbatim copy). Safe path: embed only the US/special ranges you need, written from the facts below.
- Ranges relevant to an en_US generator (from the GS1 page): 001–019, 030–039, 060–139 GS1 US (UPC-A compatible); 020–029 restricted circulation within a region; 040–049 restricted within a company; 050–059 GS1 US reserved; 200–299 restricted circulation (region); 952 "used for demonstrations and examples of the GS1 system"; 977 ISSN; 978–979 ISBN; 980 refund receipts; 981–983 coupons (common currency); 990–999 coupons. Prefixes do not identify country of origin. Selected others: 300–379 France, 400–440 Germany, 450–459 & 490–499 Japan, 500–509 UK, 690–699 China, 730–739 Sweden, 750 Mexico, 754–755 Canada, 760–769 Switzerland, 800–839 Italy, 840–849 Spain, 870–879 Netherlands, 880–881 South Korea, 890 India, 930–939 Australia, 940–949 New Zealand.
- Generator rule: for fake UPC/EAN use prefix `952` (GS1 demo) or a `2xx` restricted-circulation prefix — never a real company prefix; for US-looking UPC-A use number system 0–1/6–8 with a random 5-digit manufacturer code and accept that it may collide with real products.
## 12. Books, music, food
### Project Gutenberg catalog
- URL: `https://www.gutenberg.org/cache/epub/feeds/pg_catalog.csv` (20 MB) / `pg_catalog.csv.gz` (5.3 MB), regenerated weekly; also RDF (`rdf-files.tar.bz2`, 121 MB) and MARC.
- Fields: `Text#, Type, Issued, Title, Language, Authors, Subjects, LoCC, Bookshelves` (comma CSV, quoted; Authors as `Last, First, birth-death`, `;`-separated).
- Count (2026-09-17): 79 381 rows; Type: Text 78 130, Sound 1 114, Dataset 89, Image 33, other 12; Language: en 62 860, fr 4 190, fi 3 684, de 2 425, it 1 110.
- Licence: each RDF record carries `<cc:Work><cc:license rdf:resource="https://creativecommons.org/publicdomain/zero/1.0/"/>` → catalog metadata is CC0 1.0 (verified on `cache/epub/1/pg1.rdf`). The CSV itself has no header notice; treat as CC0 by the same source. Terms of use: do not hammer their servers (bulk feeds are the sanctioned path), "Project Gutenberg" is a trademark — royalties for commercial use of the *name*; do not brand the data file with it beyond a source citation.
- Embed: yes (CC0). For a fake-data library ship a filtered subset (Type=Text, Language=en, title + first author, maybe 5–10k rows) rather than 20 MB.
### Music genres and instruments
| Source | Licence | Embed? | Count | Notes |
|---|---|---|---|---|
| [MusicBrainz genre list](https://musicbrainz.org/genres) | Genre is a core entity (schema doc lists 13 core entities incl. Genre) → core data dump `mbdump.tar.bz2` is CC0. The genre→entity *associations* come via user tags (CC BY-NC-SA 3.0) — do not use those | Yes, names only | 2 202 (`/ws/2/genre/all?fmt=json`, `genre-count`) | WS API needs a descriptive User-Agent; list is a JSON array of `{id, name, disambiguation}` |
| [Discogs data dumps](https://data.discogs.com/) | CC0 | Yes | genre/style vocab is embedded in release XML (monthly dumps, GBs) — no standalone style list; UNVERIFIED count (~15 genres, ~600 styles from memory) | Impractical to extract; prefer MusicBrainz |
| ID3v1 genres | Names 0–79 from the 1999 ID3v1 spec, 80–125 Winamp, 126–191 Winamp 5.6 (2010) — a de-facto spec table, treated as public domain facts | Yes | 192 (0–191) | Wikipedia [List of ID3v1 genres](https://en.wikipedia.org/wiki/List_of_ID3v1_genres) (CC BY-SA page; the list itself is a spec enumeration). id3.org returned HTTP 500 during research — UNVERIFIED against the original spec page |
| Wikidata instruments | CC0 | Yes | 9 194 items `P31/P279* Q34379` with English labels; 2 372 with a Hornbostel–Sachs number (P1762) | Filter by P1762 for a clean orchestral/folk instrument list |
| Wikidata music genres | CC0 | Yes | 6 619 (`Q188451`) | Noisier than MusicBrainz |
### USDA FoodData Central
- Downloads: [fdc.nal.usda.gov/download-datasets](https://fdc.nal.usda.gov/download-datasets/) (April 2026 release):
- Foundation Foods: JSON 459 KB zip / 6.5 MB; CSV 3.7 MB zip / 32 MB — 394 foods (API `totalHits`, dataType=Foundation).
- SR Legacy (final, April 2018): JSON 12.3 MB zip / 205 MB; CSV 6.7 MB zip / 54 MB — 7 793 foods.
- FNDDS 2021-2023: CSV 200 MB zip / 1.6 GB — 5 432 foods.
- Branded: CSV 428 MB zip / 2.9 GB — 433 403 foods.
- Full: CSV 460 MB zip / 3.1 GB.
- Licence ([API guide](https://fdc.nal.usda.gov/api-guide/)): "USDA FoodData Central data are in the public domain and they are not copyrighted. They are published under CC0 1.0 Universal (CC0 1.0)". No permission needed; USDA *requests* listing FoodData Central as source and notifying them. Suggested citation: "U.S. Department of Agriculture, Agricultural Research Service. FoodData Central, 2019. fdc.nal.usda.gov."
- Embed: yes (CC0). Ship a derived list (SR Legacy `description` + `food_category`) not the raw dump.
- CSV structure (UNVERIFIED — the field-description PDF could not be text-extracted; from memory): `food.csv` (fdc_id, data_type, description, food_category_id, publication_date), `food_category.csv` (id, code, description), `nutrient.csv` (id, name, unit_name, nutrient_nbr, rank), `food_nutrient.csv` (id, fdc_id, nutrient_id, amount, …), `food_portion.csv`, `measure_unit.csv`, `sr_legacy_food.csv` (fdc_id, NDB_number), `foundation_food.csv`.
### Dishes, cuisines, drinks
| Source | Licence | Embed? | Count | Notes |
|---|---|---|---|---|
| Wikidata dishes (`Q746549`) | CC0 | Yes | 7 480 with English label | SPARQL `?i wdt:P31/wdt:P279* wd:Q746549` |
| Wikidata cuisines (`Q1968435`) | CC0 | Yes | 219 | |
| Wikidata cocktails (`Q134768`) | CC0 | Yes | 305 | |
| Wikidata beer (`Q44` subclass tree) / wine (`Q282`) | CC0 | Yes | 420 / 3 025 | Wine tree is mostly appellations/brands; filter by P279 depth |
| IBA official cocktails | IBA site "© IBA 2026 – All rights reserved"; Wikipedia list CC BY-SA 4.0 | Names only are facts (102 cocktails: 34 Unforgettables, 34 Contemporary Classics, 34 New Era, 2024 revision); do not copy recipes/descriptions | 102 | Safest: take the names from Wikidata (`P31 Q134768`) rather than the IBA page |
| [Open Brewery DB](https://github.com/openbrewerydb/openbrewerydb) | MIT (LICENSE, © 2025 Open Brewery DB) | Yes | 11 931 breweries (US 8 308, DE 1 445, AU 514, BE 478, CA 283, NZ 243) | `breweries.csv`: id (UUID), name, brewery_type (micro 5 906, brewpub 3 947, closed 644, planning 635, regional 240, contract 208, large 137, proprietor, taproom, bar, nano, cidery, beergarden), address_1..3, city, state_province, postal_code, country, phone, website_url, longitude, latitude. Real businesses — fine for names, but generating "fake" data with real addresses/phones may be undesirable |
| Wikipedia lists (cuisines, dishes, IBA) | CC BY-SA 4.0 | No verbatim copies | | |
## Summary of safe picks
- Countries: `datasets/country-codes` (PDDL) or `annexare/Countries` (MIT); flag emoji derived from alpha-2.
- Languages: `langs` / `iso-639-1` (MIT, 639-1 only).
- Subdivisions: US-only list from Census (public domain); skip world list.
- Commerce words: faker `en` lists (MIT). Check digits: implement from the rules above; GS1 prefix table: hand-write the few ranges needed, use `952`/`2xx` for generated barcodes; ISBN range file: do not embed.
- Books: Gutenberg `pg_catalog.csv` (CC0), filtered.
- Music: MusicBrainz genre names (CC0, 2 202), ID3v1 list (192), Wikidata instruments (CC0).
- Food: USDA FDC SR Legacy/Foundation descriptions (CC0, cite USDA), Wikidata dishes/cuisines (CC0), Open Brewery DB (MIT).
- Avoid: mledoze (ODbL), lukes (CC BY-SA), stefangabos (CC BY-SA/LGPL), Debian iso-codes (LGPL), SIL 639-3 table (no-redistribution clause), GS1/ISBN-International pages (all rights reserved), Wikipedia tables verbatim (CC BY-SA), MusicBrainz tag associations (CC BY-NC-SA).
+114
View File
@@ -0,0 +1,114 @@
# Open data sources for real Swedish geographic and postal-address data
Researched 2026-09-16. Goal: Sweden → län → kommun → tätort/postort → street, with real postal codes, embeddable in an MIT-licensed library.
## TL;DR
| Need | Best source | Licence | Embed in MIT repo? |
|---|---|---|---|
| Län + kommun names/codes | SCB `kommunlankod-2026.xlsx` | CC0 1.0 | Yes, no attribution needed |
| Population per kommun (weighting) | SCB PxWeb API `BE0101A/BefolkningNy` | CC0 1.0 | Yes |
| Tätorter (2,017) with code, kommun, län, population | SCB open geodata WFS/GeoPackage `stat:Tatorter_2023` + PxWeb `MI0810A/LandarealTatortN` | CC0 1.0 | Yes |
| Streets + house numbers + postnummer + postort + kommun + coordinates (~3.9M points) | Lantmäteriet *Belägenhetsadress Nedladdning, vektor* (STAC, GeoPackage per kommun) | CC BY 4.0 **plus** personal-data terms, purpose review, Basic-auth download | **Not as-is** — raw redistribution is bound by the personal-data terms; a derived aggregate (street ↔ postnummer ↔ postort, no house numbers) is the realistic path but needs a stated purpose approved by Lantmäteriet. Legal review needed. |
| postnummer → postort (list only) | GeoNames `SE.zip` (18,887 rows, 1,780 postorter) | CC BY 4.0 | Yes with attribution — but source is an undated third-party file, likely stale |
| Street names per kommun/postort without Lantmäteriet | Trafikverket NVDB *Gatunamn* (Lastkajen) | CC0 1.0 | Yes, but no postnummer; needs account to download |
| Street names | OpenStreetMap | ODbL 1.0 | Effectively no (share-alike on the extracted database) |
| Municipal address files | OpenAddresses `sources/se/*` (16 kommuner) | Mixed: CC0 (Göteborg, Malmö), CC BY 4.0 (Helsingborg), unspecified (Stockholm 2016 dump, Uppsala) | Per source |
| Postal code register (authoritative) | PostNord → Postnummerservice Norden AB / Geposit AB | Commercial; "resale/sublicensing not permitted" | No |
## 1. Lantmäteriet
Since 2025-02-03 the EU High Value Datasets (HVD) products are fee-free but Lantmäteriet explicitly says they are *not* "öppna data" because conditions attach ([Avgiftsfria produkter](https://www.lantmateriet.se/sv/geodata/vara-produkter/avgiftsfria-produkter/)). Two tiers exist:
- **Öppna data** — licence "Creative Commons, CC0" ([Öppna data](https://www.lantmateriet.se/sv/geodata/vara-produkter/avgiftsfria-produkter/oppna-data/)). Verified on Geotorget product pages for *Topografi 50/100/250/1M Nedladdning, vektor* ("Villkor: Creative Commons, CC0", "Juridisk prövning: Nej", GeoPackage, download in Geotorget or via the *Geotorget Nedladdning* API).
- **Värdefulla datamängder (HVD)** — licence **CC BY 4.0** per the terms document *Användningsvillkor för värdefulla datamängder* (DNR LM2025/009266 v1.0, 2025-02-01, [PDF](https://www.lantmateriet.se/globalassets/geodata/geodataprodukter/anvandningsvillkor_for_vardefulla_datamangder.pdf)). Products with personal data use *Användningsvillkor för värdefulla datamängder som innehåller personuppgifter* (DNR LM2025/009269 v1.1, 2025-03-04, [PDF](https://www.lantmateriet.se/globalassets/geodata/geodataprodukter/anvandningsvillkor_for_vardefulla_datamangder_pu.pdf)).
Per-product status (Geotorget product pages, rendered 2026-09-16):
| Product | Fee | Terms | Legal review | Access | Format |
|---|---|---|---|---|---|
| [Belägenhetsadress Nedladdning, vektor](https://geotorget.lantmateriet.se/geodataprodukter/belagenhetsadress-nedladdning-vektor-api) | 0 | HVD terms **with personal data** | **Yes** | STAC API `https://api.lantmateriet.se/stac-vektor/v1`, collection `belagenhetsadresser` | GeoPackage, one zip per kommun |
| [Belägenhetsadress Direkt](https://geotorget.lantmateriet.se/geodataprodukter/belagenhetsadress-direkt-api) | 0 | HVD terms with personal data | Yes | REST `https://api.lantmateriet.se/distribution/produkter/belagenhetsadress/v4.2` (OAuth2 via API portal) | JSON / XML |
| [Belägenhetsadress Nedladdning, Inspire](https://www.lantmateriet.se/sv/geodata/vara-produkter/produktlista/belagenhetsadress-nedladdning-inspire/) | 0 | HVD terms with personal data | Yes (3–5 working days extra for security review) | Atom feed, login | GML; SWEREF 99 TM / ETRS89; semi-annual |
| [Ortnamn Nedladdning, vektor](https://geotorget.lantmateriet.se/geodataprodukter/ortnamn-nedladdning-vektor-api) | 0 | HVD terms (no personal data) | No | STAC collection `ortnamn`, one national zip (58.5 MB) | GeoPackage |
| [Ortnamn Direkt](https://geotorget.lantmateriet.se/geodataprodukter/ortnamn-direkt-api) | 0 | HVD terms | No | REST `…/distribution/produkter/ortnamn/v2.2` | JSON / XML |
| Kommun, län och rike (administrativ indelning) | 0 | STAC says `CC-BY-4.0`; a Lantmäteriet HVD presentation (GISS, 2025-03-04) tables it as **CC0** ("CC0 används som licens för tjänster som är gemensamma för HVD och NGP … Kommun, Län och Rike") — **conflicting, unverified which prevails** | No | STAC collection `kommun-lan-rike` (zip ~6 MB, yearly + "aktuell"); OGC API Features `https://api.lantmateriet.se/ogc-features/v1/administrativ-indelning` (collections `kommuner`, `lan`, `rike`, plus `-2025`/`-2026`; items endpoint returns 401) | GeoPackage / GeoJSON |
| [Topografi 10 Nedladdning, vektor](https://geotorget.lantmateriet.se/geodataprodukter/topografi-10-nedladdning-vektor) | 0 | HVD terms with personal data | **Yes** | Geotorget download / API | GeoPackage |
| Topografi 50 / 100 / 250 / 1M Nedladdning, vektor | 0 | **CC0** | No | Geotorget download / *Geotorget Nedladdning* API | GeoPackage |
Notes:
- *Administrativ indelning* is not a standalone Geotorget product; it is the STAC collection `kommun-lan-rike`, the OGC API above, an [Inspire download](https://www.lantmateriet.se/sv/geodata/vara-produkter/produktlista/administrativ-indelning-nedladdning-inspire/) (Atom/WCS), or a layer in Topografi 100/250/1M (CC0). Topografi 1M also carries "tätortsgränser och tätortssymboler".
- The STAC catalog root and collection/item listings are public (HTTP 200, no auth); it self-describes: "avgiftsfria och får användas enligt creative commons licens CC BY 4.0. För vissa informationsmängder kommer din användning att prövas juridiskt … och du behöver då godkänna särskilda användningsvillkor." Asset downloads on `dl1.lantmateriet.se` return **401 Basic** — credentials are the Geotorget account (private person) or a system account (organisation). Private accounts are "endast för privat bruk" ([konto-privatperson](https://geotorget.lantmateriet.se/konto-privatperson)); an MIT library is not private use, so apply as an organisation.
- **Record count.** 290 STAC items (one per kommun), total 324 MB zipped. Smallest Bjurholm 143 KB, largest Stockholm 9.1 MB, Göteborg 8.9 MB, Malmö 4.9 MB. Fastighetsregistret held 3,882,361 valid addresses in 2022 (2,886,796 street addresses + 964,737 rural) per [SCB/LM presentation](https://kartografiska.se/wp-content/uploads/5C_Foretagsadresser_SCB_LM.pdf); the OSM import thread cites 3.7 M. So ~3.9 M address points, not 4–5 M.
- **Fields** (JSON schema `belagenhetsadress-4.2.2.json`, matches *Nationell specifikation Adress* v1.0.1): `kommunkod`, `kommunnamn`, `kommundel` (geografisk kommundel), `adressomrade` + `adressomradestyp` (gatuadressområde / metertalsadressområde / byadressområde), `gardsadressomrade`, `adressplatsnummer`, `bokstavstillagg`, `lagestillagg`/`lagestillaggsnummer`, `adressplatstyp`, `adressplatsbeteckning`, `postnummer`, `postort` (both "beslutas av PostNord AB", only on current addresses), `popularnamn`, `distriktskod`/`distriktsnamn`, `adressattAnlaggning`, `anmarkningstyp`/`anmarkningstext`, `registerenhetsreferens` (property link), `objektidentitet`, `objektstatus`, `versionGiltigFran`, `geometri` (SWEREF 99 TM point). No län field — derive from kommunkod (first two digits).
- Lantmäteriet treats addresses as personal data via Fastighetsregisterlagen § 6: usage only for the purpose approved in the application, EU/EEA-only storage, spot checks, revocation (terms PU §§ 3.3–4.2). The OSM community reports Lantmäteriet refused ODbL compatibility and OpenAddresses ([#7657](https://github.com/openaddresses/openaddresses/issues/7657), [#7608](https://github.com/openaddresses/openaddresses/issues/7608)) has not been able to include it; the OSM thread ([127427](https://community.openstreetmap.org/t/lantmateriet-belagenhetsadress/127427), 91 posts, last 2026-07-25) says one mapper got approval for a full 3.7 M import but no formal import has happened.
- **Attribution text required (CC BY 4.0 + terms § 3.1):** product name as data source, "©Lantmäteriet", whether the data was processed, and that CC BY 4.0 applies — may be placed in accompanying documentation/metadata.
- Ortnamn: name types include `Tätort` ("minst 200 invånare"), Bebyggelse, Trakt, Kyrka, Anläggning, nature names; Ortnamnsregistret holds ~980,000 names (Lantmäteriet, figure not re-verified). Not delivered as CSV.
## 2. SCB (Statistics Sweden)
- **Licence.** Statistics and geodata published as open data: "Creative Commons 0 1.0 Universal, CC0" ([användningsvillkor](https://www.scb.se/om-scb/om-scb.se-och-anvandningsvillkor)); since 2021-07-01. Optional credit "Källa: SCB" / "Source: Statistics Sweden"; do not cite SCB as source for data you have processed. API limit 10 requests / 10 s per IP.
- **Län and kommuner with codes.** 21 län, 290 kommuner ([lan-och-kommuner](https://www.scb.se/hitta-statistik/regional-statistik-och-kartor/regionala-indelningar/lan-och-kommuner/)). Download: `https://www.scb.se/contentassets/7a89e48960f741e08918e489ea36354a/kommunlankod-2026.xlsx` (18 KB, HTTP 200). Codes: län 2 digits (01 Stockholm … 25 Norrbotten), kommun 4 digits = län + 2 (0180 Stockholm).
- **Population per kommun.** PxWeb API `https://api.scb.se/OV0104/v1/doris/sv/ssd/START/BE/BE0101/BE0101A/BefolkningNy` — region dimension has 312 values (riket, 21 län, 290 kommuner), years 1968–2024 (verified 2026-09-16). Also `BE0101C/BefArealTathetKon` (density, 1991–2025).
- **Tätorter.** 2,017 tätorter in 2023 (WFS `resultType=hits` on `stat:Tatorter_2023` → `numberOfFeatures="2017"`; SCB also states 2,017). Open geodata: `https://geodata.scb.se/geoserver/stat/wfs` (WFS/WMS) and GeoPackage download ([statistiska tätorter](https://www.scb.se/vara-tjanster/oppna-data/oppna-geodata/statistiska-tatorter/)). Attributes: `tatortskod`, `tatort`, `kommun`, `kommunnamn`, `lan`, `lannamn`, `area_ha`, `bef` (population), `ar`. Population per tätort also in PxWeb `MI0810/MI0810A/LandarealTatortN` (2,827 region codes, every 5 years to 2023). Also `Smaorter_2023`, `DeSO_2025`, `RegSO_2025`.
## 3. Postal codes
- **Owner.** PostNord Sverige owns and operates the postnummer system; PTS is supervisory authority; a Postnummerråd approves changes 3–4 times a year (sv.wikipedia). Register management is delegated to **Postnummerservice Norden AB**, machine access via **Geposit AB**. Postnummerservice's [terms](https://postnummerservice.se/en/information/kopvillkor): "free use within your own organization", "resale of data is not permitted", "sublicensing to third parties is not permitted" — unusable for an open library.
- **No open list from PostNord, Lantmäteriet or Digg.** dataportal.se search for "postnummer" (2026-09-16) yields only Postnummerservice's own commercial catalogue entry ("Svenska postnummer och postorter", published 2012), Helsingborg's municipal postnummer polygons, and Lantmäteriet's Inspire address entry. The old dataportal community called the postnummer system "slarvats bort till ett privat bolag". GitHub CSVs (`zegl/sweden-zipcode`, `lapplandi/sveriges-postnummer`, `beshrkayali/sverige_postnummer`) state no licence or provenance — do not use.
- **Lantmäteriet addresses carry `postnummer` + `postort`** on every current address point (set by PostNord), so the address product is the only lawful open-ish route to a *complete, current* postnummer ↔ postort ↔ street mapping.
- **GeoNames `SE.zip`** ([readme](https://download.geonames.org/export/zip/readme.txt)): CC BY 4.0 ("a link on your website to www.geonames.org is ok"). Downloaded and counted: **18,887 rows, 18,887 distinct postnummer, 1,780 postorter, 22 admin1 values (21 län + blank), 285 admin2 (kommun code) values; 3,055 rows lack kommun; accuracy: 16,115 rows = 4 (centroid of postal code area), 389 = 1, 2,383 blank**. Fields: country, postal code (`111 64` with space), place name, admin1 name/code (län), admin2 name/code (kommun 4-digit), admin3, lat, lon, accuracy. The GeoNames sources page lists Sweden's source as `www.pellesoft.se/upload/program/prg00745.zip` with **no date** (URL now dead) — a hobbyist upload, so freshness is unknown; the row count (18.9 k) exceeds today's ~17 k total codes, indicating retired codes are still present. The zip is rebuilt daily (Last-Modified 2026-09-16) but that says nothing about the underlying data. 96 KB zipped.
- **Counts.** Postnummerservice: "~17 000 postal codes", "10 500 deliverable postal codes", "1 743 postal cities"; Kartanalys: 10,839 geographic postnummer, 1,740 postorter; Wikipedia: >16,100 areas (2008). Changes quarterly.
- **Structure** (sv/en Wikipedia): 5 digits written `NNN NN`. Lower numbers further south, except 1xx xx = Stockholm. First digit/tens: 10–19 Stockholm, 20–29 Skåne (Malmö 20–21), 30–39 southern Sweden, 40–49 Göteborg area (Göteborg 40–41), 50–59, 60–69, 70–79, 80–89, 90–99 progressively north to Norrbotten (98x). Postorter come in 2-, 3- or 5-position sizes; in two-position places the 3rd digit (in three-position the 4th) marks delivery type: 0 = boxes/postal, 1 = boxes/business, 2–4 and 6–7 = street delivery, 5 = rural (lantbrevbäring), 8 = reply mail, 9 = competitions/temporary — with exceptions in big cities. Box codes are therefore not street-deliverable and should be excluded when generating street addresses.
## 4. Streets
- **Lantmäteriet** — see § 1; `adressomrade` with `adressomradestyp = gatuadressområde` is the street name.
- **Trafikverket NVDB** — CC0 1.0 Universal since 2017-04-03 ([trafikverket.se](https://www.trafikverket.se/e-tjanster/hamta-data-fran-trafikverket/) links `creativecommons.org/publicdomain/zero/1.0/deed.sv`; [OSM wiki](https://wiki.openstreetmap.org/wiki/Sweden/trafikverket)). Data product **Gatunamn** = "det officiellt adressbildande namnet på gatan" (municipal decision); Trafikverket is cleaning it in 2026 so it holds only names used in street addresses ([Vägnamn i NVDB](https://www.nvdb.se/sv/aktuellt/nyhetsarkiv/2026/vagnamn-i-nvdb/)). Download via Lastkajen (registration with e-mail + accept licence; Shapefile/GeoPackage), or the Datautbytesportal API. Gives street ↔ kommun (and geometry) but **no postnummer/postort**. Record count not published on the pages checked.
- **OpenStreetMap** — ODbL 1.0. A list of street names extracted from OSM is a Derivative Database: "Where you make our data or any Derivative Database available to others, it must continue to be licensed under the ODbL" ([OSMF FAQ](https://osmfoundation.org/wiki/Licence/Licence_and_Legal_FAQ)); only *Produced Works* (maps etc.) escape share-alike. Embedding an OSM-derived JSON in an MIT repo therefore forces that data file to ODbL with attribution "© OpenStreetMap contributors" + link to openstreetmap.org/copyright, and downstream consumers inherit share-alike on the data. Not recommended.
- **OpenAddresses** — 16 Swedish sources, all municipal (`sources/se/municipality_of_*.json`: Alingsås, Gislaved, Göteborg, Helsingborg, Höganäs, Kalmar, Kristinehamn, Malmö, Nacka, Sävsjö, Stockholm, Uppsala, Västerås, Vaxholm, Växjö, Österåker). OA does not relicense; licence is per source. Checked: Göteborg **CC0 1.0** (attribution "Göteborgs stad", shapefile, `gatunamn`), Malmö **CC0 1.0** (CSV with `ADRESSOMR`, `ADRESSPLAT`, `POSTNR`, `POSTORT`, cached 2024-11), Helsingborg **CC BY 4.0** ("Helsingborgs stad", GeoJSON with Gatunamn/Postnummer/Stad), Uppsala attribution required, licence name not stated (ArcGIS REST), Stockholm **no licence field**, data is a 2016 cache of a defunct WS. Usable for a few big cities only; no national coverage.
- **Unique street-name counts.** SCB (Lägenhetsregistret 2019, [artikel](https://www.scb.se/hitta-statistik/artiklar/2021/pa-ringvagen-bor-det-flest/)): "Antalet gator eller vägar med en eller flera adresser är 399 175 i hela landet" — this is streets counted per place, not unique names. Most frequent: Ringvägen 205, Skogsvägen 204, Björkvägen 200, Skolgatan 196, Storgatan 180 (in 180 kommuner, 55,360 residents). A national unique-name count was not found; expect well under 399 k (probably 100–200 k) — **unverified**.
## 5. Sizes
| Level | Count | Source |
|---|---|---|
| Län | 21 | SCB |
| Kommuner | 290 | SCB; 290 STAC items |
| Tätorter (2023) | 2,017 | SCB WFS hits |
| Postorter | ~1,740–1,780 | Postnummerservice 1,743; Kartanalys 1,740; GeoNames 1,780 |
| Postnummer | ~17,000 total, ~10,500–10,840 deliverable/geographic | Postnummerservice, Kartanalys |
| Streets with addresses (per place) | 399,175 (2019) | SCB |
| Unique street names | not published; likely 100–200 k | unverified |
| Address points | ~3.9 M (3,882,361 in 2022) | Lantmäteriet/SCB |
| Lantmäteriet address GeoPackages | 324 MB zipped, 290 files | STAC |
| GeoNames SE.zip | 96 KB (18,887 rows) | download |
| Ortnamn national GeoPackage | 58.5 MB zipped; ~980 k names | STAC; Lantmäteriet |
Embedding suggestion (size-driven, before licence): län + kommun + population (~20 KB) and tätorter with population (~100 KB) are trivially embeddable under CC0. A postnummer→postort table is ~300 KB raw. A national (street, postort, postnummer-range) table at ~400 k rows is roughly 15–25 MB JSON, ~3–5 MB gzipped — optional download pack territory; a default pack could carry the top N streets per kommun (e.g. 20 × 290 ≈ 6 k rows, <300 KB).
## 6. Licence compatibility with an MIT repo
| Source | Licence | Embed? | Required notice |
|---|---|---|---|
| SCB codes, population, tätorter | CC0 1.0 | Yes | None; optional "Källa: SCB" |
| Lantmäteriet Topografi 50/100/250/1M (admin boundaries, tätort symbols) | CC0 | Yes | None (courtesy credit suggested) |
| Lantmäteriet Ortnamn | CC BY 4.0 + HVD terms (no personal data, no review) | Yes, data file stays CC BY 4.0 inside the MIT repo | "Källa: Ortnamn Nedladdning, vektor, ©Lantmäteriet, bearbetad, CC BY 4.0" (product name, ©Lantmäteriet, processed-flag, CC BY 4.0) |
| Lantmäteriet kommun-län-rike | CC BY 4.0 per STAC (CC0 per LM slide — unresolved) | Yes | Same attribution form as above unless CC0 is confirmed |
| Lantmäteriet Belägenhetsadress | CC BY 4.0 + personal-data terms, purpose approval, EU-only storage, revocable | **Only with an approved application** whose stated purpose covers deriving and publishing an aggregated street/postnummer/postort list under CC BY 4.0. Raw points (house numbers, property links) should not be redistributed. Get written confirmation from geodatasupport@lm.se; treat as unresolved until then. | Product name, ©Lantmäteriet, "bearbetad", CC BY 4.0 |
| GeoNames SE | CC BY 4.0 | Yes | Link/credit to www.geonames.org; note staleness |
| Trafikverket NVDB Gatunamn | CC0 1.0 | Yes | None |
| OpenStreetMap | ODbL 1.0 | Not practical (share-alike on the data) | "© OpenStreetMap contributors" + ODbL if ever used |
| OpenAddresses Göteborg, Malmö | CC0 1.0 | Yes | None (Göteborg asks credit "Göteborgs stad") |
| OpenAddresses Helsingborg | CC BY 4.0 | Yes | "Helsingborgs stad" |
| OpenAddresses Stockholm/Uppsala | unspecified | No | — |
| PostNord/Postnummerservice files | commercial, no redistribution | No | — |
CC BY 4.0 and CC0 data files can sit in an MIT repo; the MIT licence covers the code, and a `DATA-LICENSES.md`/NOTICE lists each dataset with its licence and the attribution above. CC BY 4.0 § 2(a)(5)(B) (no technical protection measures) is irrelevant for a plain JSON file.
## Could not verify
- Whether the kommun-län-rike collection is CC0 or CC BY 4.0 (sources disagree).
- Whether Lantmäteriet will approve an application whose purpose is publishing a derived open dataset; OSM/OpenAddresses experience suggests refusal is likely.
- GeoNames Sweden source date; assume stale.
- National count of unique street names; Gatunamn (NVDB) record count.
- Exact number of active postnummer today (three sources: ~17 k total / 10.5–10.8 k deliverable).
+208
View File
@@ -0,0 +1,208 @@
# Open geographic / postal-address data outside Sweden — for fejkdata (MIT)
Researched 2026-09-16. "Embeddable in MIT repo" below means: the data licence permits redistribution and commercial use with at most an attribution notice, so the data can ship inside an MIT-licensed package with a separate `DATA-LICENSES` file. Anything share-alike (ODbL, CC BY-SA) is flagged as **not** embeddable. Items marked *(unverified)* were not confirmed against a primary source.
## 0. Summary table
| Country | Official open address register | Licence | Addresses | Postcode in register | Postcode data open? | Embed in MIT? |
|---|---|---|---|---|---|---|
| US | DOT National Address Database (NAD) | US public domain | ~80 M (Jun 2026) | ZIP (optional field; fill rate *unverified*) | ZCTA yes (public domain); USPS publishes no open ZIP file | Yes |
| Norway | Kartverket Matrikkelen – Adresse | CC BY 4.0 | 2.58 M (Jun 2024) | Yes (postnummer + poststed) | Yes: Bring postnummerregister, NLOD 2.0 | Yes |
| Denmark | Danmarks Adresseregister (DAR) via DAWA / Datafordeler | CC BY 4.0 | ~3.8 M | Yes | Yes (in DAR; 1,089 postnumre) | Yes |
| Finland | DVV building addresses (final release Feb 2025) + Posti PCF/BAF | CC BY 4.0 / Posti terms (free; redistribute with terms attached) | ~5 M buildings | Yes | Yes (Posti PCF/BAF) | Yes (Posti: attach terms) |
| UK (GB) | No open *address* register. OS Open UPRN (points only), OS Open Names (streets/places/postcodes), Code-Point Open (postcodes) | OGL v3 | 40 M UPRNs (no addresses); 870 k roads; 1.7 M postcodes | n/a | Yes for GB (Code-Point Open, OGL); Northern Ireland restricted | Yes (GB only; no house numbers) |
| Germany | No national register. 15 Länder publish Hauskoordinaten/Gebäudereferenzen (dl-de/by-2-0, dl-de/zero-2-0); Bavaria missing | mixed, mostly permissive | ~19 M points across Länder | Yes in Hauskoordinaten | **No** — Deutsche Post DATAFACTORY is proprietary; open PLZ lists are OSM-derived (ODbL) | Partly (Länder files yes; a national PLZ list is the problem) |
| Netherlands | BAG (Kadaster/PDOK) | CC0 1.0 | ~9.9 M | Yes | Yes (in BAG) | Yes |
| France | Base Adresse Nationale (BAN) | Licence Ouverte / Etalab 2.0 | >25 M | Yes | Yes (BAN; La Poste "base officielle des codes postaux", open licence) | Yes |
| Spain | Catastro INSPIRE Addresses (AD) + INE callejero | Catastro own licence (free, attribution) / INE legal notice | ~15.7 M points (mailwoman count) | Yes (AD:PostCode) | Correos' postcode layer closed since 2017; postcodes still appear per address in Catastro AD | Probably (read Catastro licence PDF first) |
| Australia | G-NAF (Geoscape via data.gov.au) | EULA based on CC BY 4.0 + "no mail-out" restriction | 15.95 M (Aug 2026) | Yes | In G-NAF; Australia Post file is non-commercial only | Yes (with EULA notice) |
| Canada | StatCan National Address Register (NAR) 2024; ODA v1 (2021) | StatCan Open Licence / OGL-Canada | 17.1 M / ~10 M | NAR: MAIL_POSTAL_CODE "for most addresses" | Canada Post's postal code file is **proprietary**; only what NAR/ODA carry | Yes (NAR); postcode completeness unverified |
## 1. Per-country detail
### United States
**Official address register — National Address Database (NAD), US DOT**
- Licence: US public domain (`http://www.usa.gov/publicdomain/label/1.0/`), stated on data.gov: https://catalog.data.gov/dataset/national-address-database-nad-text-file. No attribution required. Embeddable.
- Size: ~80 M records; current release compiled 2026-06-30, dataset updated 2026-09-03. Sources: https://data.transportation.gov/dataset/National-Address-Database-NAD-Text-File/fc2s-wawr/about_data, https://www.placekey.io/datasets/national-address-database (80 M; "24 fields" in the text export).
- Coverage: aggregated from state/local providers; a mix of fully, partially and non-participating states (placekey page; the DOT page https://www.transportation.gov/gis/national-address-database blocks scrapers, so the coverage map was not read).
- Formats: zipped flat text and File Geodatabase. The text file is ordered by spatial cluster, not by OID.
- Fields (NAD schema https://www.transportation.gov/sites/dot.gov/files/2023-07/NAD_Schema_202304.pdf, blocked for bots; list taken from the FGDC content-requirements doc https://www.fgdc.gov/organization/working-groups-subcommittees/address-sc/220810-nad-content-requirements-approved.pdf): address number parts (`AddNum_Pre`, `Add_Number`, `AddNum_Suf`), street name parts (`St_PreDir`, `St_PreTyp`, `St_Name`, `St_PosTyp`, `St_PosDir`, `StNam_Full`), sub-address (`Building`, `Floor`, `Unit`, `Room`), place names (`Inc_Muni`, `Post_City`, `Census_Plc`, `Uninc_Comm`), `County`, `State`, `Zip_Code`, `Plus_4`, `Longitude`, `Latitude`, `NatGrid`, `Addr_Type`, `Placement`, `NAD_Source`, `DataSet_ID`. FGDC: "A Zip Code is recommended but not mandatory"; a complete place name and state are mandatory. Expect some records without ZIP.
- Census Address Count Listing files give housing-unit counts per block, not addresses: https://census.gov/geographies/reference-files/2025/geo/addcountlisting.html. The Census MAF itself is not public.
**Administrative hierarchy — Census Bureau (public domain)**
- States: 50 + DC + PR + 4 island areas = 56 FIPS state codes ("States & Equivalents: 56" in the 2020 tallies). ANSI list: https://www.census.gov/library/reference/code-lists/ansi/ansi-codes-for-states.html (also carries FM, MH, PW, UM).
- Counties: 3,144 in the 50 states + DC; 3,234 including PR and island areas (2020 tallies https://www.census.gov/geographies/reference-files/time-series/geo/tallies.html; https://en.wikipedia.org/wiki/List_of_United_States_counties_and_county_equivalents). Connecticut switched to 9 planning regions as county equivalents in 2022 — use a current vintage.
- Places: 19,734 incorporated places + 12,454 CDPs (2020 tallies, incl. PR/island areas); "approximately 19,500" incorporated places in the 50 states (https://www.census.gov/library/stories/2020/05/america-a-nation-of-small-towns.html).
- TIGER/Line 2025 shapefiles and the pipe-delimited **Gazetteer files** (2026 vintage: states, counties, county subdivisions, places, ZCTAs; GEOID, name, area, centroid lat/lon; **no population**): https://www.census.gov/geographies/reference-files/time-series/geo/gazetteer-files.html. Gazetteer excludes island areas, includes PR.
- Population per place: **SUB-EST2025** CSVs (incorporated places + MCDs, April 2020 base through July 2025, released May 2026): https://www.census.gov/data/tables/time-series/demo/popest/2020s-total-cities-and-towns.html. Public domain. The right file for population-weighted draws.
**Postal codes**
- USPS publishes no open ZIP list file; counts quoted 41,541–41,695 active ZIPs (https://facts.usps.com/42000-zip-codes/ says 41,554). Commercial "USPS-licensed" databases exist.
- Census **ZCTA**: 33,791 ZCTAs (2020), public domain, in TIGER + Gazetteer + relationship files (ZCTA↔place, ZCTA↔county): https://www.census.gov/programs-surveys/geography/technical-documentation/records-layout/2020-zcta-record-layout.html. PO-box-only ZIPs have no ZCTA (https://en.wikipedia.org/wiki/ZIP_Code_Tabulation_Area).
- GeoNames `US.txt` (CC BY 4.0): ZIP → place name, state, county, lat/lon. Count not stated on the site *(≈41 k, unverified)*.
**Street names**
- TIGER/Line `featnames` relationship files (per county) hold every feature name with prefix/suffix parts; joined to `edges` and `addrfeat` (address ranges). Public domain. No published count of unique street names *(unverified; must be computed)*.
- NAD `St_Name`/`StNam_Full` distinct per `Post_City` is the cheaper source for "real streets in Denver".
**Sizes for the embed-vs-download decision (US)**
| Level | Count | Source |
|---|---|---|
| States + DC + PR + island areas | 56 | Census tallies 2020 |
| Counties / equivalents | 3,144 (50+DC) / 3,234 (with PR + islands) | Census tallies, Wikipedia |
| Incorporated places | 19,734 (all) / ~19,500 (50+DC) | Census tallies |
| CDPs | 12,454 | Census tallies |
| ZCTAs / USPS ZIPs | 33,791 / ~41,500 | Census tallies / USPS facts |
| Address points (NAD) | ~80 M | DOT |
| OpenAddresses US collections | NE 2.35 GB + South 7.11 GB + West 8.54 GB + Midwest 4.86 GB (zipped CSV); 2,381 US sources | https://batch.openaddresses.io/ (2026-09) |
| Unique street names | unknown; compute from NAD/TIGER | — |
### Norway
- **Matrikkelen – Adresse** (Kartverket). Licence CC BY 4.0 (https://data.norge.no/en/datasets/8f1151bf-ad63-3e47-adbb-143f876776ee/matrikkelen-adresse; Geonorge register https://register.geonorge.no/inspire-statusregister/matrikkelen-adresse/73bc0329-faac-419b-80ff-d113c0ffe6a0). Formats: CSV, SOSI, GML, SQL/PostGIS, FGDB, Atom feed, REST API `https://api.kartverket.no/adresser/v1/`. Files per kommune (daily) and fylke/country (weekly).
- Counts (Kartverket "Antall vegadresser og matrikkeladresser", extract 2024-06-01, https://www.kartverket.no/datakvalitet/rapporter/adresse/Antall%20vegadresser%20og%20matrikkeladresser.xlsx): **2,577,825 addresses** — 2,536,663 vegadresser (street addresses) + 41,162 matrikkeladresser (cadastral, no street name).
- Fields (CSV): adressenavn (street), nummer, bokstav, postnummer, poststed, kommunenummer, kommunenavn, grunnkrets, tettsted, coordinates (EPSG:25833 / 4258); bruksenhetsnummer only in non-CSV formats. Fylke = first two digits of kommunenummer.
- Hierarchy: 15 fylker + Svalbard, 357 kommuner (2024). Kartverket/SSB code lists are open *(SSB Klass API licence NLOD, unverified)*.
- Postal codes: Bring/Posten **Postnummerregister** — postnummer, poststed, kommunenummer, kommunenavn, kategori; TAB text + xlsx; **NLOD 2.0** per data.norge.no (https://data.norge.no/en/datasets/5e6847ba-156d-4e14-85d3-8d7f8b727523/postnummer-i-norge). Updated annually (Sept/Oct). Embeddable.
- OpenAddresses: 17 `no/` sources.
### Denmark
- **Danmarks Adresseregister (DAR)**, served by **DAWA** (`https://api.dataforsyningen.dk/`) and Datafordeler. Licence **CC BY 4.0**, credit "Klimadatastyrelsen" (https://datafordeler.dk/vejledning/brugervilkaar/danmarks-adresseregister-dar/). Embeddable.
- Counts: ~3.8 M adresser (Klimadatastyrelsen DAWA report); DAWA today: **1,089 postnumre, 53,976 vejnavne, 99 kommuner, 5 regioner** (queried 2026-09-16 via `?format=csv`).
- Fields (DAWA `adresser`): vejnavn, husnr, etage, dør, postnr, postnrnavn, supplerende bynavn, kommunekode/-navn, regionskode/-navn, sogn, koordinater (ETRS89/WGS84), DAR ids. Downloads: CSV, JSON, GeoJSON, per kommune or whole country; `adgangsadresser` = access address (building entrance, no floor/door). Docs: https://dawadocs.dataforsyningen.dk/dok/adresser.
- Hierarchy: 5 regioner → 98 kommuner (99 rows incl. Christiansø) → postnumre/byer. All in DAWA.
- Postcodes: in DAR (source PostNord, updated on change), same licence.
- OpenAddresses: `dk/countrywide` (1 source).
### Finland
- No single address-register download today. Sources:
- **DVV (Digital and Population Data Services Agency) building address open data** — CC BY 4.0; final official release 2025-02-14, distribution ended 2025-03-14; archived copies (CSV/GeoPackage, ~5 M buildings: street fi/sv, house number, postal code, municipality, WGS84) at https://markuskainu.fi/posts/2025-02-03-dvv-rakennusten-osoitetiedot/. Successor: **SYKE/Ryhti building addresses** via OGC API Features, CC BY 4.0 (https://avoindata.suomi.fi/data/fi/dataset/syke-rakennusten-osoitteet-ogc-api-features) — page blocked the fetch; verify fields and bulk download before relying on it.
- **NLS (Maanmittauslaitos)** topographic database address points — CC BY 4.0 (https://www.maanmittauslaitos.fi/en/opendata-licence-cc40); address points computed from road-name data.
- **Digiroad** (Fintraffic since 2026-01-01) — CC BY 4.0; road names + address number ranges (https://vayla.fi/en/transport-network/data/digiroad).
- **Posti postal code files** (https://www.posti.fi/en/for-businesses/customer-support/postal-code-services; directory https://www.posti.fi/webpcode/): **PCF** (postal code, names fi/sv + abbreviations, type code, region code/name, municipality code/name fi/sv, language code; daily), **BAF** (every street per postal code with odd/even building-number ranges + municipality; weekly; excludes Åland), **POM** (monthly changes). Fixed-width `.dat`, Unix LF. Terms (service description 2015-01-01): "The files are freely downloadable … The data can be disclosed to third parties but it must be ensured that the recipient of the data is aware of the terms of use for the service as well as the download date." No licence name, no fee, no commercial restriction stated. Embeddable if the terms PDF + download date ship with the data.
- Hierarchy: Statistics Finland classifications, CC BY 4.0: Municipalities 2026 https://stat.fi/en/luokitukset/kunta/kunta_1_20260101, Regions 2026 https://stat.fi/en/luokitukset/maakunta/maakunta_1_20260101. Posti PCF also carries municipality + region codes.
- OpenAddresses: `fi/countrywide-fi`, `fi/countrywide-sv` + 38 municipal sources.
### United Kingdom (Great Britain; NI separate)
- No open full-address register (AddressBase is paid). Open pieces, all **OGL v3**, attribution "Contains OS data © Crown copyright and database right [year]" (https://wiki.openstreetmap.org/wiki/Ordnance_Survey_OpenData_Licence; OS moved to OGL v3 in 2015):
- **OS Open Names** — "over 870 000 named and numbered roads, nearly 44 000 settlements and over 1.6 million postcodes"; 34 attributes incl. `name1`, `name2` (Welsh/Gaelic), `type`, `local_type` (populatedPlace: City/Town/Village/Hamlet/Suburban Area/Other Settlement; transportNetwork: Named Road, Numbered Road, Section Of …; other: Postcode), `postcode_district`, `populated_place`, `district_borough`, `county_unitary`, `region`, `country`, `GEOMETRY_X/Y`. CSV/GML/GeoPackage, quarterly. https://docs.os.uk/os-downloads/products/addresses-and-names-portfolio/os-open-names. Gives street → settlement → district → county → region → country, but **no house numbers** and only the postcode district per road.
- **Code-Point Open** — ~1.7 M GB postcode units; fields Postcode, Positional_quality_indicator, Eastings, Northings, Country_code, NHS codes, Admin_county_code, Admin_district_code, Admin_ward_code. No street or locality names. GB only; NI excluded. https://docs.os.uk/os-downloads/products/areas-and-zones-portfolio/code-point-open.
- **OS Open UPRN** — ~40 M UPRN + coordinates, no addresses. https://www.ordnancesurvey.co.uk/products/os-open-uprn.
- **ONS Postcode Directory (ONSPD)** — postcode → all admin geographies; OGL, but requires Royal Mail attribution ("Contains Royal Mail data © Royal Mail copyright and database right [year]") and NI (BT) postcodes are end-user-licence only (no commercial redistribution): https://www.ons.gov.uk/methodology/geography/licences. Drop NI rows.
- Hierarchy codes: ONS GSS codes (countries, regions, counties/UAs, districts, wards) via the ONS Open Geography Portal, OGL. Population: ONS mid-year estimates (OGL) at local-authority level; settlement population must come from GeoNames/Wikidata.
- OpenAddresses: 0 `gb/` sources.
- Practical: a deliverable GB address needs street + number + postcode unit; open data gives street + postcode *district* only. "Real street in a real town with a plausible district" is achievable; a real full postcode for that street is not.
### Germany
- **No national open address register.** BKG sells "Georeferenzierte Adressdaten" commercially. Open data is per Land: Hauskoordinaten / Gebäudereferenzen from 15 Länder under dl-de/by-2-0, dl-de/zero-2-0 or CC BY 4.0; **Bavaria missing** (~13 M people) — https://github.com/sister-software/mailwoman/issues/2300 (19.3 M permissive points vs 20.3 M in an ODbL build). Example: NRW "Gebäudereferenzen NW", dl-de/zero-2-0, semi-annual, https://open.nrw/dataset/172a0ba8-d470-47c1-ac89-b85f8190ac7e. Thuringia is dl-de/by-2-0. OpenAddresses lists 14 `de/*/statewide` sources (bb, bw, hb, he, hh, mv, ni, nw, rp, sh, sl, sn, st, th) — no `by`; Berlin appears as a Geoportal source *(check)*.
- Hierarchy: **BKG VG250** (Land → Regierungsbezirk → Kreis → Verwaltungsgemeinschaft → Gemeinde with AGS keys; VG250-EW adds population), dl-de/by-2-0: https://gdz.bkg.bund.de/index.php/default/open-data/verwaltungsgebiete-1-250-000-stand-01-01-vg250-01-01.html. **Destatis Gemeindeverzeichnis GV100/GV-ISys** (all Gemeinden with AGS, population, area, PLZ of the Verwaltungssitz), quarterly ASCII/Excel: https://www.destatis.de/DE/Themen/Laender-Regionen/Regionales/Gemeindeverzeichnis/_inhalt.html — licence dl-de/by-2-0 *(not stated on that page; Destatis default, unverified)*. 16 Länder, ~400 Kreise, ~10,800 Gemeinden *(approx.)*.
- **Postal codes are the blocker**: Deutsche Post **DATAFACTORY** is proprietary and licensed for internal use only (https://www.deutschepost.de/de/d/deutsche-post-direkt/datafactory.html; a FragDenStaat request for the data was refused). Open PLZ lists (OpenPLZ API, yetzt/postleitzahlen, Geofabrik postcode polygons) are **OSM-derived → ODbL** (OpenPLZ sources page https://www.openplzapi.org/en/sources/: streets and PLZ from OSM, municipalities from Destatis GV100). GeoNames `DE.txt` provenance not stated *(unverified)*. The Länder Hauskoordinaten carry PLZ per address, so a PLZ↔Ort↔Straße table can be derived legitimately for the 15 covered Länder; GV100 gives one PLZ per Gemeinde seat.
- OpenAddresses: 32 `de/` sources.
### Netherlands
- **BAG** (Basisregistratie Adressen en Gebouwen), Kadaster. Licence **CC0 1.0** (https://data.overheid.nl/en/dataset/basisregistratie-adressen-en-gebouwen--bag-; OSM wiki: "released under a Public Domain license"). Embeddable, no attribution.
- ~9.9 M addresses (mailwoman count). Object types: panden, verblijfsobjecten, **nummeraanduidingen** (huisnummer + huisletter + toevoeging + postcode), **openbare ruimten** (streets), **woonplaatsen**. Fields yield street, number, postcode (4 digits + 2 letters), woonplaats, gemeente (CBS code), RD/WGS84 point.
- Download: BAG Extract XML (monthly full + mutations) and BAG GeoPackage via PDOK Atom `https://service.pdok.nl/lv/bag/atom/bag.xml`; REST API `https://api.pdok.nl/lv/bag/`. data.overheid.nl mentions distribution fees for some Kadaster products; the PDOK extract/GeoPackage is free.
- Hierarchy: 12 provincies → 342 gemeenten (2025) → ~2,500 woonplaatsen. CBS gebiedsindelingen (CC BY 4.0); CBS "Kerncijfers wijken en buurten" for population.
- Postcodes: in BAG since 2012 (CC0). PostNL's product is proprietary but unnecessary.
- OpenAddresses: `nl/countrywide` (1 source).
### France
- **Base Adresse Nationale (BAN)** — **Licence Ouverte / Etalab 2.0** since 2020-01-01 (https://fr.wikipedia.org/wiki/Base_adresse_nationale; https://adresse.data.gouv.fr/outils/telechargements). Embeddable with attribution.
- >25 M addresses. Downloads: CSV national / per département / per commune, "CSV with BAN ids", addok format; daily updates with weekly/monthly archives; also MVT/WFS/WMS. CSV fields `id, id_fantoir, numero, rep, nom_voie, code_postal, code_insee, nom_commune, code_insee_ancienne_commune, nom_ancienne_commune, x, y, lon, lat, type_position, alias, nom_ld, libelle_acheminement, nom_afnor, source_position, source_nom_voie, certification_commune, cad_parcelles` *(from BAN docs, not re-verified this session)*.
- Hierarchy: INSEE **Code Officiel Géographique** (18 régions, 101 départements, ~34,900 communes), Licence Ouverte; INSEE populations légales per commune (Licence Ouverte).
- Postal codes: in BAN; also La Poste **"Base officielle des codes postaux"** (commune → code postal → libellé d'acheminement, INSEE code), open licence, https://datanova.laposte.fr/datasets/laposte-hexasmal and https://www.data.gouv.fr/datasets/base-officielle-des-codes-postaux.
- OpenAddresses: 107 `fr/` sources (one per département, fed weekly from BAN).
### Spain
- **Catastro INSPIRE Addresses (AD)** — per-municipality GML via Atom `http://www.catastro.minhap.es/INSPIRE/Addresses/ES.SDGC.AD.atom.xml`, refreshed ~6-monthly; WFS `http://ovc.catastro.meh.es/INSPIRE/wfsAD.aspx`. Fields: ThoroughfareName, address number, `AD:PostCode` (5 digits), AdminUnitName (municipio, provincia), point at building entrance or parcel centroid. Coverage: 95% of Catastro territory, **excludes País Vasco and Navarra** (own cadastres; Navarra has OpenAddresses source `es/nc/statewide`). Licence: "licencia de cesión de derechos que se obtendrá de manera automática", free, attribution "© Dirección General del Catastro" (https://www.catastro.hacienda.gob.es/webinspire/documentos/Conjuntos%20de%20datos.pdf; https://www.catastro.hacienda.gob.es/webinspire/index.html). mailwoman labels it CC BY 4.0 (15.66 M points); the Catastro PDF does not say "CC BY" — **read the "Descripción de la licencia" PDF before embedding**.
- **INE Callejero del Censo Electoral** — VIAS, TRAMOS, PSEUDOVIAS, UNIDADES POBLACIONALES; ASCII ZIP, national + per province, Jan/Jul each year, `https://www.ine.es/prodyser/callejero/caj_esp/caj_esp_MMYYYY.zip` (https://datos.gob.es/en/catalogo/ea0010587-callejero-de-censo-electoral). Licence = INE aviso legal (free reuse with source citation *(unverified; page blocked)*). Whether TRAMOS carry código postal: *(unverified)*.
- **CartoCiudad (IGN/CNIG)** — addresses, postcodes, toponyms, admin units; CC BY 4.0 (https://datos.gob.es/es/aplicaciones/direcciones-postales-de-cartociudad-espana). Its postcode polygon layer came from Correos and was **withdrawn in 2017** (Correos now sells it, ~€6k/yr) — https://www.nosolosig.com/articulos/asi-hice-el-mapa-de-los-codigos-postales-de-espana-con-sig-y-datos-abiertos. Per-address postcodes remain in Catastro AD and CartoCiudad addresses.
- Hierarchy: 17 CCAA + 2 ciudades autónomas → 50 provincias → 8,132 municipios; INE padrón gives population per municipio; INE codes are open.
- OpenAddresses: `es/countrywide` (Catastro) + `es/nc/statewide` + 4 others.
### Australia
- **G-NAF** (Geoscape via data.gov.au, https://data.gov.au/data/dataset/geocoded-national-address-file-g-naf). August 2026: **15,949,543 addresses** (15,108,510 principal). Quarterly. PSV tables (~5 GB unpacked, many tables to join) + **G-NAF Core** single simplified table; GDA94 or GDA2020.
- Licence: **G-NAF End User Licence Agreement, "based on CC BY 4.0"**, plus: must not be used to compile addresses for sending mail unless each is verified against a secondary source; must follow the Australian Privacy Principles. Preferred attribution: "Incorporates or developed using G-NAF © Geoscape Australia licensed by the Commonwealth of Australia under the Open Geo-coded National Address File (G-NAF) End User Licence Agreement." Embeddable with the EULA text and that notice; the mail restriction is irrelevant to fake data but must be passed on.
- Fields: number first/last with prefix/suffix, flat/level, street name + type + suffix, locality, state, postcode, lat/lon, mesh block; LGA via ABS codes.
- Hierarchy: 8 states/territories → LGAs (~560) → localities (~15,000 in the G-NAF locality table). **ABS ASGS Edition 3** boundaries/allocation files, CC BY 4.0: https://www.abs.gov.au/statistics/standards/australian-statistical-geography-standard-asgs/edition-3-july-2021-june-2026/access-and-downloads/digital-boundary-files. Population: ABS ERP by SA2/LGA (CC BY 4.0).
- Postcodes: in G-NAF. Australia Post's postcode file is free only as a **non-commercial PDF**; CSV products are paid (https://auspost.com.au/business/services/data-services/address-data/postcode-data). ABS "Postal Areas" (2,644 POAs, CC BY 4.0) are mesh-block approximations.
- OpenAddresses: `au/countrywide` + 7 statewide + councils (58 sources).
### Canada
- **National Address Register (NAR)**, Statistics Canada — December 2024 release, **17.1 M** records, all 13 provinces/territories; two CSVs per province (addresses, locations). Licence **Statistics Canada Open Licence** (use, reproduce, distribute, sell, sublicense; attribution "Adapted from Statistics Canada, National Address Register, 2024. This does not constitute an endorsement by Statistics Canada of this product."): https://www.statcan.gc.ca/en/reference/licence; user guide https://www150.statcan.gc.ca/n1/pub/46-26-0002/462600022024002-eng.htm. Fields: `CIVIC_NO, CIVIC_NO_SUFFIX, OFFICIAL_STREET_NAME/TYPE/DIR, APT_NO_LABEL, MAIL_STREET_NAME/TYPE/DIR, MAIL_MUN_NAME, MAIL_PROV_ABVN, MAIL_POSTAL_CODE, CSD_CODE, CSD_ENG_NAME, CSD_FRE_NAME, PROV_CODE, REPPOINT_LATITUDE/LONGITUDE, BG_X/Y, BU_USE`. Postal code present "for most addresses" — completeness per province unverified. Embeddable.
- **Open Database of Addresses (ODA) v1** (2021), ~10 M records from 99 datasets, OGL-Canada, CSV per province, postal code where the provider had it: https://www.statcan.gc.ca/en/lode/databases/oda. Older; prefer NAR.
- Hierarchy: StatCan Standard Geographical Classification 2021 — 13 provinces/territories → 293 census divisions → ~5,160 census subdivisions (municipalities); population per CSD from Census 2021. StatCan Open Licence.
- **Postal codes are proprietary**: Canada Post asserts copyright, sells the file, and sued Geocoder.ca in 2012 (https://opennorth.ca/resources/open-postal-code-data/). StatCan's PCCF is licensed from Canada Post and not redistributable. GeoNames `CA.txt` = **FSA (first 3 chars) only**. Only NAR/ODA `MAIL_POSTAL_CODE` gives open per-address postal codes.
- OpenAddresses: `ca/countrywide` + 221 sources; Canada collection 1.25 GB.
## 2. Cross-country hierarchy sources
| Source | Licence | Content | Notes |
|---|---|---|---|
| Debian **iso-codes** (`iso_3166-1.json`, `iso_3166-2.json`; XML deprecated) | LGPL-2.1 (https://salsa.debian.org/iso-codes-team/iso-codes) | All ISO 3166-1 countries and 3166-2 subdivisions with codes, types, parents; 90+ translations | LGPL on a data file is awkward inside MIT; ship it as a separately-licensed data file or consume it only at build time. ISO's own OBP is not open. |
| **GeoNames** `admin1CodesASCII.txt`, `admin2Codes.txt` | CC BY 4.0 | admin1/admin2 names + geonameids per country | admin1 codes follow ISO 3166-2 for some countries and FIPS/numeric codes for others. |
| `amckenna41/iso3166-2` (PyPI) | MIT (code); data compiled from Wikipedia/RestCountries | 250 countries, >5,000 subdivisions with names, lat/lon | Convenient; provenance mixed. |
| Per-country official lists (section 1) | see country | authoritative codes + population | Preferred for the 11 target countries. |
## 3. Localities with population
| Source | Licence | Size | Fit |
|---|---|---|---|
| **GeoNames** `cities500/1000/5000/15000` + `allCountries` (https://download.geonames.org/export/dump/) | CC BY 4.0 (attribution to geonames.org) | cities500 ≈185 k, cities1000 ≈130 k, cities5000 ≈50 k, cities15000 ≈25 k; fields geonameid, name, asciiname, alternatenames, lat/lon, feature class/code (PPL, PPLA…), country, admin1–4 codes, population, timezone | Best single cross-country locality+population source; population uneven (many zeros). Embeddable with attribution. |
| **Census SUB-EST2025** (US) | public domain | ~19.5 k places | Authoritative US population. |
| **Wikidata** (P1082 population, P131 located-in, P281 postal code) | CC0 | all | No attribution; needs a SPARQL/dump pipeline; quality varies. |
| **Who's On First** (https://www.whosonfirst.org/docs/licenses/) | CC0 for WOF's own work; 312 sources ranging CC0…CC BY 3.0/4.0, OGL | localities, regions, counties with hierarchy + population where sourced | Attribution list is per source; no share-alike found. Heavy (one GeoJSON per record). |
| **dr5hn/countries-states-cities-database** (https://github.com/dr5hn/countries-states-cities-database) | **ODbL 1.0** | 250 countries, 5,299 states, 153,765 cities, 844,248 postcodes (125 countries) | **Share-alike → not embeddable in MIT.** Community-maintained; README: data "may contain errors or lag behind geopolitical changes"; issue tracker shows large re-parenting fixes (8,727 French cities). Not a source of truth. |
## 4. Postal codes — GeoNames and per-country openness
- **GeoNames postal** (https://download.geonames.org/export/zip/, CC BY 4.0): TSV fields `country code, postal code, place name, admin name1, admin code1, admin name2, admin code2, admin name3, admin code3, latitude, longitude, accuracy (1–6)`. ~100 countries. Readme caveats: Canada and Netherlands partial codes only (full NL/CA files separate); UK from Royal Mail (© Royal Mail 2022 — attribution needed); Ireland/Malta first letters only; Argentina 5 chars; Brazil only `-000` codes; Chile/China partial. Lat/lon is algorithmic (matched to a GeoNames toponym or averaged from neighbours) — a locality centroid, not the postcode's. Per-country provenance is not documented, so DE's origin is unknown.
| Country | Open? | Best open source |
|---|---|---|
| US | Yes (ZCTA, public domain); USPS list not a file | Census ZCTA + GeoNames US.txt |
| Norway | Yes (NLOD 2.0) | Bring postnummerregister; also in Matrikkelen |
| Denmark | Yes (CC BY 4.0) | DAWA `postnumre` (1,089) |
| Finland | Yes (Posti terms) | Posti PCF (codes) + BAF (streets per code) |
| UK | GB yes (OGL, Code-Point Open); NI restricted | Code-Point Open + OS Open Names |
| Germany | **No** (Deutsche Post proprietary); OSM lists are ODbL | Derive PLZ↔Ort↔Straße from Länder Hauskoordinaten (15 Länder); GV100 for seat PLZ |
| Netherlands | Yes (CC0, in BAG) | BAG |
| France | Yes (Licence Ouverte) | La Poste hexasmal + BAN |
| Spain | Polygon layer closed since 2017; per-address codes open via Catastro AD (own licence) | Catastro AD `AD:PostCode`; INE callejero |
| Australia | In G-NAF (EULA/CC BY 4.0); Australia Post file paid/non-commercial | G-NAF locality table |
| Canada | **No** (Canada Post proprietary; GeoNames FSA only) | NAR `MAIL_POSTAL_CODE` where present |
## 5. OpenAddresses and OpenStreetMap
**OpenAddresses** (https://openaddresses.io/, https://batch.openaddresses.io/)
- Global collection 48.39 GB (2026-09-07); US in 4 regional collections (table in section 1); Canada 1.25 GB. 3,282 sources (1,586 flagged "need attention"). Sources per country in the batch list: US 2,381, CA 222, FR 107, AU 58, FI 40, DE 32, NO 17, ES 6, DK 1, NL 1, GB 0.
- Licensing: the repo LICENSE (BSD-3) covers source definitions only. **Each source keeps its own licence**, recorded in the source JSON and `state.txt`; "most sources only require attribution", but several are share-alike (ODbL, CC BY-SA) — e.g. OSM-derived and some municipal sources (https://geocode.earth/docs/reference/data_sources/). For MIT: filter by the per-source licence field, keep PD/CC0/CC BY/OGL/dl-de sources, generate the attribution file from `state.txt`. The old results.openaddresses.io coverage page is frozen at mid-2021.
- The official registers above are themselves the OA sources for NO/DK/FI/NL/FR/ES/AU (and NAD/state sources for the US); going to the register directly gives cleaner licensing and fresher data. OA's value is one uniform CSV schema (`LON,LAT,NUMBER,STREET,UNIT,CITY,DISTRICT,REGION,POSTCODE,ID,HASH`).
**OpenStreetMap** — ODbL 1.0
- Any *substantial* extract (systematic, >100 features, or an area with >1,000 inhabitants) is a Derivative Database; distributing it — including embedded in a library — requires offering it under ODbL (share-alike) plus attribution. OSMF guideline: https://osmfoundation.org/wiki/Licence/Community_Guidelines/Substantial_-_Guideline ("village map OK, town map not OK"; repeated small extractions count as one).
- Consequence: an embedded street/postcode list derived from OSM (OpenPLZ, Geofabrik postcode polygons, yetzt/postleitzahlen, OSM-derived dr5hn rows) cannot ship inside an MIT package without that data file being ODbL and tracked as a separate database; downstream users mixing it into their databases inherit share-alike. Keep OSM out of the default; at most an *optional* download pack clearly labelled ODbL.
## 6. Embed vs. optional download — sizes and a recommendation
**US (ships by default with en_US)**
- Embed (small, public domain): 56 states (+ USPS abbreviations, FIPS), 3,234 counties (GEOID, name, state), ~19.5 k incorporated places with population + county + lat/lon (SUB-EST2025 joined to the Gazetteer), 33,791 ZCTAs with the ZCTA→place/county relationship, and a compact **street-name pool per place** derived from NAD (distinct `St_Name`+`St_PosTyp` per `Post_City`/`State`, top-N by frequency). Places+ZCTAs ≈ 2–4 MB gzipped; a street pool for all ~19.5 k places would run to tens of MB — cap to the ~1,000 largest places or ~20 streets per place for the default pack.
- Download pack: full NAD (~80 M rows, multi-GB), TIGER `featnames`/`addrfeat` per county, OA US collections (~23 GB zipped). Real house numbers + ZIP+4 come only from here.
- Deliverability caveat: NAD ZIP is optional, and place naming mixes `Post_City` (USPS city) with `Inc_Muni`; use `Post_City` + `Zip_Code` for mailing realism.
**Other countries** — same shape: embed hierarchy + localities-with-population + postcode↔locality table (a few MB per country: NO's 2.6 M addresses collapse to ~4.5 k postcodes and ~90 k street names; DK 1,089 postcodes / 54 k street names; NL/FR/AU similar), with street pools for the largest localities; keep full address points as download packs fetched from the official Atom/CSV endpoints at runtime, never vendored (G-NAF ~5 GB, BAN CSV ~2 GB, BAG ~2 GB, NAR ~1 GB).
**Licence file needed in the repo** (`DATA-LICENSES.md`): CC BY 4.0 (GeoNames, Kartverket, Klimadatastyrelsen, Statistics Finland, DVV/SYKE, NLS, ABS, CartoCiudad), NLOD 2.0 (Bring), OGL v3 (OS, ONS), Licence Ouverte 2.0 (BAN, La Poste, INSEE), G-NAF EULA notice, StatCan Open Licence notice, dl-de/by-2-0 and dl-de/zero-2-0 (BKG, Länder), Posti terms + download date, Catastro "© Dirección General del Catastro". CC0/public domain (US Census, DOT NAD, BAG, Wikidata) need nothing but are worth listing.
## 7. Open questions / unverified
- NAD ZIP fill rate and the list of non-participating states (DOT site blocks fetches; read the NAD coverage map manually).
- Unique US street-name count — compute from NAD once downloaded.
- Catastro licence text ("Descripción de la licencia" PDF) — confirm redistribution inside a software package.
- INE aviso legal wording; whether callejero TRAMOS include código postal.
- Destatis GV100 licence statement; Berlin/Bayern Hauskoordinaten status.
- SYKE/Ryhti address API fields and bulk-download availability (avoindata.fi blocked the fetch).
- NAR `MAIL_POSTAL_CODE` completeness per province.
- GeoNames DE/US postal-code provenance and counts.
- BAN CSV field list re-verification.
+287
View File
@@ -0,0 +1,287 @@
# Fake-data libraries: data-category inventory
Researched 2026-09-16. Sources: official docs and current `master`/`main`/`next` source trees. "Big four" = @faker-js/faker (JS), Faker (Python), gofakeit (Go), Datafaker (Java). Abbreviations in tables: **JS**, **PY**, **GO**, **DF**, **BG** (Bogus), **MM** (Mimesis), **CH** (Chance), **RB** (Ruby faker).
## 1. Category modules and their generators
Legend: names are the library's own identifiers (camelCase = JS/DF/BG, snake_case = PY/MM/RB, PascalCase = GO). Columns list what each library offers under that module; a `—` means the module does not exist there.
### person
| Lib | Generators |
|---|---|
| JS | firstName, lastName, middleName, fullName, prefix, suffix, sex, sexType, gender, bio, jobArea, jobDescriptor, jobTitle, jobType, zodiacSign (all name methods take `sex: female|male|generic`) |
| PY | first_name/_female/_male/_nonbinary, last_name/_female/_male/_nonbinary, name/_female/_male/_nonbinary, prefix/_female/_male/_nonbinary, suffix/_female/_male/_nonbinary, language_name; separate providers: job (job, job_female, job_male), ssn (ssn; per-locale variants: vat_id, itin, ein, nif, cpf, personnummer …), passport (passport_number, passport_dob, passport_owner), profile (profile → job, company, ssn, residence, current_location, blood_group, website, username, name, sex, address, mail, birthdate; simple_profile) |
| GO | Person (struct), Name, NamePrefix, NameSuffix, FirstName, MiddleName, LastName, Gender, Age, Ethnicity, SSN, EIN, Hobby, SocialMedia, Bio, Contact, Email, Phone, PhoneFormatted, Teams |
| DF | Name: name, nameWithMiddle, fullName, firstName, femaleFirstName, maleFirstName, lastName, prefix, suffix, title, username; Demographic: race, educationalAttainment, demonym, sex, maritalStatus; Gender; Pronouns; BloodType; Mbti; Zodiac; Relationship; Hobby; Mood; Job; Passport; DrivingLicense; IdNumber (valid/invalid per country: US SSN, SE, ZA, SG FIN/UIN, CN, PT NIF, MX, PL PESEL, KR RRN, GE); CPF/CNPJ (BR); Nigeria; Australia |
| BG | Name: FirstName, LastName, FullName, Prefix, Suffix, FindName, JobTitle, JobDescriptor, JobArea, JobType; Person object (Gender, FirstName, LastName, FullName, UserName, Avatar, Email, DateOfBirth, Address, Phone, Website, Company); country extensions: Ssn/Ein (US), Personnummer/Samordningsnummer (SE), Cpr (DK), Henkilotunnus (FI), Fodselsnummer (NO), Pesel/Nip/Regon (PL), CodiceFiscale (IT), Cpf/Cnpj (BR), Sin (CA), Nino (GB), Cnp (RO), NationalNumber (BE), Nif/Nipc (PT) |
| MM | Person: first_name, surname/last_name, patronymic, full_name, name, title, username, password, email, birthdate, gender, gender_symbol, gender_code, sex, height, weight, blood_type, occupation, nationality, university, academic_degree, language, phone_number, telephone, identifier |
| CH | first, last, name, prefix, suffix, gender, age, birthday, ssn, cpf, cf (IT codice fiscale), profession |
| RB | Name (name, name_with_middle, first_name, male_first_name, female_first_name, neutral_first_name, last_name, prefix, suffix, initials), Gender, Demographic, IDNumber (per-country), Job, Relationship, Blood, DrivingLicence, Avatar, ChileRut, SouthAfrica, NationalHealthService |
### address / location
| Lib | Generators |
|---|---|
| JS (location) | buildingNumber, cardinalDirection, city, continent, country, countryCode, county, direction, language, latitude, longitude, nearbyGPSCoordinate, ordinalDirection, postalAddress, secondaryAddress, state({abbreviated}), street, streetAddress, timeZone, zipCode({state, format}) |
| PY (address) | address, building_number, city, city_suffix, country, country_code, current_country, current_country_code, postcode, street_address, street_name, street_suffix; locale extras e.g. en_US: city_prefix, secondary_address, administrative_unit, state_abbr, zipcode_plus4, postcode_in_state, military_ship/state/apo/dpo; ja_JP: prefecture, city, town, chome, ban, gou, building_name; fr_FR: department, department_name, department_number, region; es_ES: region/autonomous_community; de_DE: city_with_postcode, street_suffix_short/long; en_GB: county; separate **geo** provider: coordinate, latitude, longitude, latlng, local_latlng(country), location_on_land (real geonames.org places with tz) |
| GO (address) | Address (struct), City, Country, CountryAbr, State, StateAbr, Street, StreetName, StreetNumber, StreetPrefix, StreetSuffix, Unit, Zip, Latitude, LatitudeInRange, Longitude, LongitudeInRange |
| DF (Address) | streetName, streetAddressNumber, streetAddress, secondaryAddress, zipCode, postcode, eircode, zipCodePlus4, zipCodeByState, countyByZipCode, streetSuffix, streetPrefix, citySuffix, cityPrefix, city, cityName, state, stateAbbr, latitude, longitude, latLon, lonLat, timeZone, country, countryCode, buildingNumber, fullAddress, mailBox; plus Country (name, code2/3, capital, currency, flag), Nation (nationality, language, capital, flag), Compass, Mountain, Planet, Space, Locality (locale codes) |
| BG (Address) | ZipCode, City, StreetAddress, CityPrefix, CitySuffix, StreetName, BuildingNumber, StreetSuffix, SecondaryAddress, County, Country, FullAddress, CountryCode, State, StateAbbr, Latitude, Longitude, Direction, CardinalDirection, OrdinalDirection; premium Bogus.Locations: GPS, altitude, depth, geohash |
| MM (Address) | street_number, street_name, street_suffix, secondary_address, address, state/region/province/federal_subject/prefecture (abbr), postal_code/zip_code, country, country_code (A2/A3/numeric), country_emoji_flag, default_country, city, latitude, longitude, coordinates (DMS option), continent, calling_code/isd_code, iata_code, icao_code |
| CH | address, street, city, zip, postal (CA), postcode (GB), state, province, country, areacode, phone, latitude, longitude, altitude, depth, coordinates, geohash, locale |
| RB | city, street_name, street_address, secondary_address, building_number, mail_box, community, zip_code(state_abbreviation:), zip, postcode, time_zone, street_suffix, city_suffix, city_prefix, state, state_abbr, country, country_by_code, country_name_to_code, country_code, country_code_long, latitude, longitude, full_address, full_address_as_hash; Travel: airport, train_station; Locations: australia |
### company
| Lib | Generators |
|---|---|
| JS | name, buzzAdjective, buzzNoun, buzzPhrase, buzzVerb, catchPhrase, catchPhraseAdjective, catchPhraseDescriptor, catchPhraseNoun (locale data: adjective, descriptor, noun, legal_entity_type, name_pattern) |
| PY | company, company_suffix, catch_phrase, bs; locale extras (e.g. it_IT/pl_PL company_vat, nl_NL, ru_RU large_company, etc.) |
| GO | Company, CompanySuffix, BS, Blurb, BuzzWord, Slogan, Job, JobDescriptor, JobLevel, JobTitle |
| DF | Company: name, suffix, industry, profession, buzzword, catchPhrase, bs, logo, domainName, url; Business; Brand; IndustrySegments; Marketing; Team; Restaurant; Subscription; Twitter; University; Educator |
| BG | CompanySuffix, CompanyName, CatchPhrase, Bs |
| MM | Finance.company, company_type |
| CH | company, profession |
| RB | Company (name, suffix, industry, profession, type, catch_phrase, buzzword, bs, logo, ein, duns_number, swedish_organisation_number, czech_organisation_number, french_siren/siret, norwegian/australian/spanish/polish/russian/south_african/brazilian ids …), Business, Marketing, IndustrySegments, Team, University, Educator, Restaurant, Subscription, Construction |
### internet
| Lib | Generators |
|---|---|
| JS | displayName, domainName, domainSuffix, domainWord, email, exampleEmail, emoji, httpMethod, httpStatusCode, ip, ipv4, ipv6, jwt, jwtAlgorithm, mac, password, port, protocol, url, userAgent, username |
| PY | email, safe_email, free_email, company_email, ascii_* variants, free_email_domain, domain_name, domain_word, safe_domain_name, tld, hostname, dga, http_method, http_status_code, iana_id, image_url, ipv4, ipv4_network_class, ipv4_private, ipv4_public, ipv6, mac_address, nic_handle(s), port_number, ripe_id, slug, uri, uri_extension, uri_page, uri_path, url, user_name; user_agent provider (chrome, firefox, safari, opera, internet_explorer, platform tokens); emoji provider |
| GO | URL, UrlSlug, DomainName, DomainSuffix, IPv4Address, IPv6Address, MacAddress, HTTPStatusCode(Simple), HTTPMethod, HTTPVersion, LogLevel, UserAgent (+Chrome/Firefox/Opera/Safari/API), Username, Password, Emoji (+22 category fns), InputName, Svg |
| DF | Internet: username, emailAddress(name), safeEmailAddress, emailSubject, domainName/Word/Suffix, url, webdomain, image, httpMethod, password(opts), port, macAddress, ipV4Address, privateIpV4Address, publicIpV4Address, ipV4Cidr, ipV6Address, ipV6Cidr, slug, uuidv3/4/7, userAgent, botUserAgent; Http; Domain; Sip; Emoji; SlackEmoji; Credentials; Aws; Azure; Hashing; Fingerprint; Computer; Device; Camera; Drone; Robin |
| BG | Avatar, Email, ExampleEmail, UserName, UserNameUnicode, DomainName, DomainWord, DomainSuffix, Ip, Port, IpAddress, IpEndPoint, Ipv6, UserAgent, Mac, Password, Color, Protocol, Url, UrlWithPath, UrlRootedPath |
| MM | Internet: content_type, dsn, http_status_message/code, http_method, ip_v4/v6 (+object, cidr, with_port, special), cloud_region, asn, mac_address, hostname, url, uri, query_string/parameters, tld, user_agent, port, path, slug, public_dns, http_request/response_headers |
| CH | avatar, color, domain, email, fbid, google_analytics, hashtag, ip, ipv6, klout, tld, twitter, url |
| RB | Internet (email, username, password, domain_name, ip_v4/v6, mac_address, url, slug, user_agent, uuid, bot_user_agent …), Internet::HTTP, Omniauth, Stripe, X (Twitter), Avatar, Placeholdit, LoremFlickr |
### finance / payment / bank
| Lib | Generators |
|---|---|
| JS | accountName, accountNumber, amount, bic, bitcoinAddress, creditCardCVV, creditCardIssuer, creditCardNumber, currency, currencyCode, currencyName, currencyNumericCode, currencySymbol, ethereumAddress, iban, litecoinAddress, pin, routingNumber, transactionDescription, transactionType |
| PY | bank: aba, bank, bank_country, bban, iban, swift, swift8, swift11; credit_card: credit_card_number, credit_card_provider, credit_card_expire, credit_card_security_code, credit_card_full; currency: currency, currency_code, currency_name, currency_symbol, cryptocurrency(_code/_name), pricetag |
| GO | Price, CreditCard, CreditCardCvv, CreditCardExp, CreditCardNumber, CreditCardType, Currency, CurrencyLong, CurrencyShort, AchRouting, AchAccount, BitcoinAddress, BitcoinPrivateKey, BankName, BankType, Cusip, Isin |
| DF | Finance: creditCard(type), bic, iban(country), ibanSupportedCountries, usRoutingNumber; Money: currency, currencyCode, currencyNumericCode, currencySymbol; Currency; Stock; FinancialTerms; CryptoCoin; Coin; Business |
| BG | Account, AccountName, Amount, TransactionType, Currency, CreditCardNumber, CreditCardCvv, BitcoinAddress, EthereumAddress, RoutingNumber, Bic, Iban; GB SortCode, VatNumber |
| MM | Finance: bank, currency_iso_code, currency_symbol, cryptocurrency_iso_code/_symbol, price, price_in_btc, stock_ticker, stock_name, stock_exchange; Payment: cid, bitcoin_address, ethereum_address, credit_card_network/number/expiration_date, cvv, credit_card_owner |
| CH | cc, cc_type, currency, currency_pair, dollar, euro, exp, exp_month, exp_year |
| RB | Finance (credit_card, vat_number, ticker, stock_market), Bank (name, swift_bic, iban, account_number, routing_number, bsb_number), Currency, Crypto, Blockchain (bitcoin, ethereum, tezos, aeternity), Invoice, Stripe, Coin |
### commerce / product
| Lib | Generators |
|---|---|
| JS | department, isbn, price, product, productAdjective, productDescription, productMaterial, productName, upc |
| PY | barcode: ean, ean8, ean13, localized_ean*; isbn: isbn10, isbn13; sbn; (no product/commerce provider in core; community Ecommerce provider) |
| GO | Product (struct), ProductName, ProductDescription, ProductCategory, ProductFeature, ProductMaterial, ProductUPC, ProductAudience, ProductDimension, ProductUseCase, ProductBenefit, ProductSuffix, ProductISBN |
| DF | Commerce: department, productName, material, brand, vendor, price(min,max), promotionCode; Barcode; Code (isbn10/13, gtin8/13, ean, asin, imei); Appliance; GarmentSize; Size; Tire; House; ElectricalComponents |
| BG | Department, Price, Categories, ProductName, Color, Product, ProductAdjective, ProductMaterial, Ean8, Ean13 |
| MM | Code: locale_code, issn, isbn, ean, imei, pin |
| CH | — |
| RB | Commerce (color, department, material, product_name, price, promotion_code, brand, vendor), Barcode, Code (isbn, ean, asin, imei, npi, nric, rut, sin), Appliance, Device, Camera, House |
### vehicle / transport
| Lib | Generators |
|---|---|
| JS | bicycle, color, fuel, manufacturer, model, type, vehicle, vin, vrm; airline: aircraftType, airline, airplane, airport, flightNumber, recordLocator, seat |
| PY | automotive: license_plate, vin (per-locale plate formats) |
| GO | Car (struct), CarMaker, CarModel, CarType, CarFuelType, CarTransmissionType; Airline*: AircraftType, Airplane, Airport, AirportIATA, FlightNumber, RecordLocator, Seat |
| DF | Vehicle: vin, manufacturer, make, model(make), makeAndModel, style, color, upholstery(+Color/Fabric), transmission, driveType, fuelType, carType, engine, carOptions, standardSpecs, doors, licensePlate(state); Aviation; Transport; DrivingLicense |
| BG | Vin, Manufacturer, Model, Type, Fuel; GbRegistrationPlate |
| MM | Transport: manufacturer, car, airplane, vehicle_registration_code(locale) |
| CH | — |
| RB | Vehicle (vin, manufacture, make, model, make_and_model, style, color, transmission, drive_type, fuel_type, car_type, engine, car_options, standard_specs, doors, door, year, mileage, license_plate, singapore_license_plate, version), Drone, Travel::Airport, Travel::TrainStation |
### lorem / text / words
| Lib | Generators |
|---|---|
| JS | lorem: lines, paragraph(s), sentence(s), slug, text, word(s); word: adjective, adverb, conjunction, interjection, noun, preposition, sample, verb, words; hacker: abbreviation, adjective, ingverb, noun, phrase, verb; string: alpha, alphanumeric, binary, fromCharacters, hexadecimal, nanoid, numeric, octal, sample, symbol, ulid, uuid |
| PY | lorem: word(s), sentence(s), paragraph(s), text(s), get_words_list; misc: password, md5, sha1, sha256, binary, boolean, null_boolean, csv/dsv/psv/tsv, fixed_width, json, json_bytes, image, tar, zip, uuid4 |
| GO | Word families: Noun* (Common/Concrete/Abstract/Collective*/Countable/Uncountable/Proper/Determiner), Verb* (Action/Linking/Helping/Transitive/Intransitive), Adverb* (Manner/Degree/Place/Time*/Frequency*), Preposition* (Simple/Double/Compound), Adjective* (8 kinds), Pronoun* (8 kinds), Connective* (6 kinds), Word, Interjection, Sentence, Paragraph, LoremIpsum{Word,Sentence,Paragraph}, Question, Quote, Phrase{,Noun,Verb,Adverb,Preposition}, Comment; Hacker*, Hipster{Word,Sentence,Paragraph}; Letter(N), Digit(N), Vowel, Lexify, Numerify, RandomString, ShuffleStrings |
| DF | Lorem (word(s), sentence(s), paragraph(s), characters, fixedString, maxLengthSentence, supplemental), Text (character, upper/lowercase, text with symbol rules), Verb, Word, Hacker, Hipster, Shakespeare, Yoda, Joke, ChuckNorris, Matz, NatoPhoneticAlphabet, FunnyName, Marketing |
| BG | Lorem: Word(s), Letter, Sentence(s), Paragraph(s), Text, Lines, Slug; Hacker: Abbreviation, Adjective, Noun, Verb, IngVerb, Phrase; Rant: Review(s) |
| MM | Text: alphabet, level, text, sentence, title, words, word, quote, color, hex_color, rgb_color, answer, emoji |
| CH | paragraph, sentence, syllable, word, character, letter, string |
| RB | Lorem (word(s), character(s), sentence(s), paragraph(s), question(s), paragraph_by_chars, multibyte), Markdown, Hipster, Hacker, Adjective, Verbs, Quote, Quotes::Shakespeare/Chiquito/Rajnikanth, ChuckNorris, Emotion, Source (code snippets), Html, Json, Types, Alphanumeric, String, Boolean, NatoPhoneticAlphabet |
### date / time
| Lib | Generators |
|---|---|
| JS | anytime, between, betweens, birthdate({mode: age|year}), future, past, recent, soon, month, weekday, timeZone (locale data: month, weekday) |
| PY | date_time: date, date_time, date_object, time, time_object, date_between(_dates), date_time_between(_dates), date_this_century/decade/month/year, date_time_this_*, date_time_ad, date_of_birth(minimum_age, maximum_age), future_date(time), past_date(time), iso8601, unix_time, time_delta, time_series, timezone, pytimezone, am_pm, century, year, month, month_name, day_of_month, day_of_week |
| GO | Date, PastDate, FutureDate, DateRange, NanoSecond, Second, Minute, Hour, Month, MonthString, Day, WeekDay, Year, TimeZone, TimeZoneAbv, TimeZoneFull, TimeZoneOffset, TimeZoneRegion |
| DF | DateAndTime: future, past, between, birthday, birthdayLocalDate, duration, period; Time; TimeAndDate |
| BG | Past, PastOffset, Soon, SoonOffset, Future, FutureOffset, Between, BetweenOffset, Recent, RecentOffset, Timespan, Month, Weekday |
| MM | Datetime: date, datetime, time, timestamp, formatted_*, week_date, day_of_week, month, year, day_of_month, timezone, gmt_offset, periodicity, future/past_date(time), duration, bulk_create_datetimes |
| CH | ampm, date, hammertime, hour, millisecond, minute, month, second, timestamp, timezone, weekday, year |
| RB | Date (between, between_except, forward, backward, birthday, in_date_period, on_day_of_week_between), Time (between, between_dates, forward, backward) |
### phone
| Lib | Generators |
|---|---|
| JS | number({style: human|national|international}), imei |
| PY | phone_number, country_calling_code, msisdn (per-locale formats) |
| GO | Phone, PhoneFormatted |
| DF | PhoneNumber: cellPhone, cellPhoneInternational, phoneNumber, phoneNumberInternational, phoneNumberNational, extension, subscriberNumber |
| BG | PhoneNumber, PhoneNumberFormat |
| MM | Person.phone_number/telephone (mask), Address.calling_code |
| CH | phone, areacode |
| RB | PhoneNumber (phone_number, cell_phone, country_code, phone_number_with_country_code, cell_phone_with_country_code, cell_phone_in_e164, area_code, exchange_code, subscriber_number, extension) |
### science / medical
| Lib | Generators |
|---|---|
| JS | chemicalElement, unit; medical locale data (en only) |
| PY | — (community: Healthcare, Biology, Geoscience, Scientific) |
| GO | — |
| DF | Science: element, elementSymbol, unit, scientist, tool, quark, leptons, bosons; Medical, Medication, Disease, MedicalProcedure, CareProvider, Observation, BloodType, Measurement, Weather, Planet, Space, Mountain, Cannabis, LargeLanguageModel |
| BG | premium Bogus.Healthcare |
| MM | Science: rna_sequence, dna_sequence |
| CH | — |
| RB | Science (element, element_symbol, element_state, element_subcategory, scientist, modifier, tool), Space, Measurement, Cannabis, Medical (NationalHealthService) |
### music / animal / food / books / entertainment
| Lib | Generators |
|---|---|
| JS | music: album, artist, genre, songName; animal: bear, bird, cat, cetacean, cow, crocodilia, dog, fish, horse, insect, lion, petName, rabbit, rodent, snake, type; food: adjective, description, dish, ethnicCategory, fruit, ingredient, meat, spice, vegetable; book: author, format, genre, publisher, series, title |
| PY | — (core has none; community Music, Sci Fi) |
| GO | Song, SongName, SongArtist, SongGenre; PetName, Animal, AnimalType, FarmAnimal, Cat, Dog, Bird; Fruit, Vegetable, Breakfast, Lunch, Dinner, Snack, Dessert, Drink; Beer{Alcohol,Blg,Hop,Ibu,Malt,Name,Style,Yeast}; Book, BookTitle, BookAuthor, BookGenre; Movie, MovieName, MovieGenre; Celebrity{Actor,Business,Sport}; Minecraft (18 fns) |
| DF | Music: instrument, key, chord, genre; RockBand; Artist; Kpop; Hololive; Animal: name, scientificName, genus, species; Cat, Dog, Horse; Food: ingredient, allergen, spice, dish, fruit, vegetable, sushi, measurement; Apple, Beer, Cheese, Coffee, Dessert, IceCream, Tea; Book, Movie, Show, OscarMovie; ~75 entertainment franchises, ~31 videogames, 9 sports |
| BG | Music (premium Hollywood: movies, TV, actors) |
| MM | Food: vegetable, fruit, dish, spices, drink |
| CH | animal (with type), tv, radio, rpg, dice |
| RB | Music (+11 bands), Creature (animal, bird, cat, dog, horse), Food, Beer, Coffee, Tea, Dessert, Book, Books (4 series), Movie/Movies (17), TvShows (39), Games (26), JapaneseMedia (10), Sports (6), Fantasy::Tolkien, Religion::Bible, Superhero, DcComics, Kpop, Ancient, GreekPhilosophers, Cosmere |
### color / image / system / database / misc
| Lib | Generators |
|---|---|
| JS | color: cmyk, colorByCSSColorSpace, cssSupportedFunction, cssSupportedSpace, hsl, human, hwb, lab, lch, rgb, space; image: avatar, avatarGitHub, dataUri, personPortrait, url, urlLoremFlickr, urlPicsumPhotos; system: commonFileExt/Name/Type, cron, directoryPath, fileExt/Name/Path/Type, mimeType, networkInterface, semver; database: collation, column, engine, mongodbObjectId, type; git: branch, commitDate, commitEntry, commitMessage, commitSha; number: bigInt, binary, float, hex, int, octal, romanNumeral; datatype.boolean; science.unit |
| PY | color: color, color_name, safe_color_name, hex_color, safe_hex_color, rgb_color, rgb_css_color, color_hsl/hsv/rgb/rgb_float; file: file_extension, file_name, file_path, mime_type, unix_device, unix_partition; python: pybool, pydecimal, pydict, pyfloat, pyint, pyiterable, pylist, pyobject, pyset, pystr, pystr_format, pystruct, pytuple, enum; doi; emoji; user_agent |
| GO | Color, HexColor, RGBColor, HSLColor, SafeColor, NiceColors; Image, ImageJpeg, ImagePng; CSV, JSON, XML, SQL, FileExtension, FileMimeType, Template, Markdown, EmailText, FixedWidth; ID, UUID; Error* (9 kinds); AppName, AppVersion, AppAuthor; Language, LanguageAbbreviation, LanguageBCP, ProgrammingLanguage; School; Gamertag, Dice; Bool, Weighted, FlipACoin; Number/Int*/Uint*/Float* ranges |
| DF | Color, Image, File, App, ProgrammingLanguage, LanguageCode, Number, Bool, Unique, Options, Photography, Military, OlympicSport, Weather, Construction, Community, Chiquito … (263 providers total) |
| BG | Images: DataUri, PicsumUrl, PlaceholderUrl, LoremFlickrUrl; System: FileName, DirectoryPath, FilePath, CommonFileName, MimeType, CommonFileType/Ext, FileType, FileExt, Semver, Version, Exception, AndroidId, ApplePushToken, BlackBerryPin; Database: Column, Type, Collation, Engine; Randomizer (numbers, chars, Guid, Hash, Enum, WeightedRandom, Shuffle …) |
| MM | Hardware: resolution, screen_size, cpu, cpu_frequency, generation, cpu_codename, ram_type, ram_size, ssd_or_hdd, graphics, manufacturer, phone_model; Development: software_license, calver, version, stage, programming_language, os, boolean, system_quality_attribute; File, BinaryFile, Path, Cryptographic (uuid, hash, token, mnemonic), Numeric, Choice |
| CH | bool, falsy, floating, integer, natural, prime, hex, guid, hash, coin, dice, normal (Gaussian), n, unique, weighted, android_id, apple_token, bb_pin, wp7_anid, wp8_anid2 |
| RB | Color, File, Computer, Device, ProgrammingLanguage, App, Number, Boolean, Hash/Crypto, Json, Html, Markdown, Military, Compass, Coin, Mountain, Nation, Space, Time, Types, Slack Emoji, Vulnerability identifier |
Module presence summary:
| Module | JS | PY | GO | DF | BG | MM | CH | RB |
|---|---|---|---|---|---|---|---|---|
| person | x | x | x | x | x | x | x | x |
| address | x | x | x | x | x | x | x | x |
| company | x | x | x | x | x | (finance) | x | x |
| internet | x | x | x | x | x | x | x | x |
| finance | x | x | x | x | x | x | x | x |
| commerce/product | x | barcode only | x | x | x | code only | — | x |
| vehicle | x | plate/vin | x | x | x | x | — | x |
| airline | x | — | x | x | — | iata/icao | — | x |
| lorem/words | x | x | x | x | x | x | x | x |
| date | x | x | x | x | x | x | x | x |
| phone | x | x | x | x | x | x | x | x |
| science | x | — | — | x | — | dna | — | x |
| medical | data only | — | — | x | premium | — | — | x |
| music | x | — | x | x | — | — | — | x |
| animal | x | — | x | x | — | — | x | x |
| food | x | — | x | x | — | x | — | x |
| book | x | isbn | x | x | — | — | — | x |
| color | x | x | x | x | x | text | x | x |
| image | x | x | x | x | x | binaryfile | avatar | x |
| system/file | x | x | x | x | x | x | — | x |
| database | x | — | sql | — | x | dsn | — | — |
| git | x | — | — | — | — | — | — | — |
| hacker/hipster | x | — | x | x | x | — | — | x |
| national IDs | — | ssn per locale | SSN/EIN | IdNumber | ext. pkgs | identifier | ssn/cpf/cf | IDNumber |
| pop-culture | — | — | minecraft, movie | ~110 | premium | — | tv, rpg | ~85 |
| hardware | — | — | — | Computer/Device | — | x | — | Computer/Device |
| weather/space | — | — | — | x | — | — | — | Space |
## 2. Locale coverage and address data structure
### Locale counts
| Lib | Locales | Structure |
|---|---|---|
| JS | 77 dirs under `src/locales` (incl. `base`, `en_BORK`, `en_AU_ocker`) | One dir per locale, one file per module per key (`en/location/city_name.ts`). Fallback chain per Faker instance: `[de_CH, de, en, base]`; first locale that defines a key wins. `base` holds locale-independent data (ISO codes, time zones). A key may be explicitly `null` to mean "not applicable" (e.g. no zip codes for HK) so it does not fall through to `en`. https://fakerjs.dev/guide/localization.html |
| PY | 126 locale codes listed; per provider: address 67, person 86, ssn 56 | Python subclasses per `providers/<provider>/<locale>/__init__.py` overriding tuples/formats of the base. Missing provider for a locale falls back to `en_US`. Coverage is uneven (`am_ET` = phone only). Multi-locale `Faker(['it_IT','en_US'])` with optional weights. https://faker.readthedocs.io/en/master/locales.html |
| GO | 1 (English/US). No locale API; `Language`/`LanguageBCP` only emit language codes | Static Go maps in `data/*.go`. Open issue #352 "generate data for specific country/language" unanswered. https://github.com/brianvoe/gofakeit/issues/352 |
| DF | 127 yml files in `src/main/resources` (incl. 49 country-only files like `_SE.yml` and 78 language(-region) files); README says "60+" | One YAML per locale (`sv-SE.yml`) under `sv-SE: faker: address: …`; `en/` split into 264 per-provider files. Locale chain `de_CH → de → en`, key-by-key. Values are templates `#{Name.first_name}`. Custom yml via `faker.addPath(locale, path)`. Country file (`_SE.yml`) layered over language file. |
| BG | 50 (`data/*.locale.json`; README says 46) | Verbatim copy of faker.js (v5-era) locale JSON, merged with `data_extend/*.locale.json` by a gulp task, shipped as BSON. Falls back to `en`. https://github.com/bchavez/Bogus/wiki/Creating-Locales |
| MM | 47 (53 dirs incl. `global`, `int`, `bin`, template) | Per locale six JSON files: `address, datetime, finance, food, person, text`. Everything else (internet, payment, code, hardware…) is locale-independent. |
| CH | 1 (en; `postal` CA and `postcode` GB formats exist, `locale()` only returns codes) | Data inline in `chance.js`. |
| RB | 58 yml files ("over 40" in README) | One YAML per locale `lib/locales/<code>.yml` (`en` split into `en/*.yml`), I18n gem, key-level fallback to `en`. |
### Address data: real or synthetic, and correlation
| Lib | Cities | Street names | Postcodes | Hierarchy / correlation |
|---|---|---|---|---|
| JS | `en`: real list `city_name` (~1000 US cities) is one of 5 `city_pattern`s; the other 4 are `{prefix} {firstName}{suffix}` style synthetic. `de`: 327 real cities; `sv`: no `city_name`, pattern-only synthetic. Varies per locale. | `en`: `street_pattern` = `{firstName} {street_suffix}` / `{lastName} …` (synthetic); `de`: 1800 real street names (from one NRW region). | Format masks (`#####`, `#####-####`). `en_US` has `postcode_by_state` so `zipCode({state:'CA'})` gives a state-prefixed range. | None between city↔state↔zip. `county` is a flat list. `timeZone` is per-locale. `nearbyGPSCoordinate` is the only "correlated" geo function. Issue #983: en_CA returned US cities; PR #2141 added real cities for ZA locales — real-city coverage is locale-by-locale volunteer work. |
| PY | Per locale. `en_US`/`en_GB`: synthetic (`{city_prefix} {first_name}{city_suffix}`). `de_DE` 400+ real, `sv_SE` 45 real, `pl_PL` real, `ja_JP` real prefectures/cities/towns, `es_ES` real provinces + regions. | `en_US`: `{first_name} {street_suffix}`; `sv_SE`: prefix+suffix ("Björkgatan"); `de_DE`: synthetic name+Straße; `pl_PL`: real street-name pools. | Masks. `en_US`: `postcode_in_state` with real per-state numeric ranges (still random within range, "not correlated with cities"). `fr_FR`: postcode built from a real department number (correlated to department, not to city). `en_GB`: real outward-area letters but random assembly. `sv_SE`: `%####`. | No country→state→city hierarchy anywhere; `ja_JP` has prefecture–city pairs in data but methods pick independently. `geo.local_latlng(country)` returns real geonames places (name, lat, lon, country, tz) — the only real, correlated location record in the library. `address_formats` are weighted (`25.0` standard vs `1.0` military). |
| GO | Real list of major US cities | Real-looking suffix lists; StreetName is a pool | `#####` mask | None: `Address()` picks city, state, zip independently; lat/lon are uniform random on the globe, not near the city. Open issue #196 "Better address" unanswered. |
| DF | `en`: `city` formats are all synthetic (`#{city_prefix} #{Name.first_name}#{city_suffix}`); `cityName` draws from a separate real list. `sv-SE`: real `city_name` list; `de`: synthetic. | `en`: `#{Name.first_name} #{street_suffix}`; `de`: real `street_root` list combined randomly. | Masks; `en-US` `postcode_by_state` (`350##`), `zipCodeByState`, `countyByZipCode` (en-US only), `eircode` (IE). | None by default. Issue #1551 (Locale.CHINA yields "福建省南京市") open as "proposal". Issue #1477 unresolved directive for city names in some locales. |
| BG | Whatever faker.js v5 had per locale (`en`: synthetic patterns only — real `city_name` list postdates the copy) | synthetic | Masks; `en_US` `postcode_by_state` | None. Issue #481 "Invalid Zip Codes (US)" open. Issue #342 "City returns people's names" (side effect of `{firstName}` patterns). |
| MM | Real lists per locale (`en`: real US cities; `sv`: real municipalities incl. small towns) | `en`: real San Francisco street names; `sv`: real Swedish street names | `postal_code_fmt` mask per locale; `sv` has none | `state` (name+ISO 3166-2 abbr) real, but no link to city or postcode; `default_country` tied to locale. Open question #1584 on choosing default city per locale. |
| CH | Synthetic (3-syllable word) | Synthetic word + real suffix | Real-format masks | None |
| RB | `en`: synthetic patterns (`city_with_state` exists but is `city, state` random) ; some locales have real lists | `{first_name} {street_suffix}` | `#####`; `zip_code(state_abbreviation:)` uses per-state pattern (`900##`); doc says "may not be an actual state zip" | `full_address_as_hash` returns fields but does not correlate them. Issue #2881 (state vs state_abbr mismatch, country vs country_code mismatch) open; #1958 asked for geocodable city+zip; #2899 fake zips closed "not planned". |
**Which libraries correlate any address fields at all:** Python Faker (`fr_FR` postcode↔department, `en_US` zip↔state range, `geo.local_latlng` real records), faker-js/Datafaker/Bogus/Ruby (`zip↔state` range only, en_US), Mimesis (locale↔country only). None ships a country→region→city→postcode hierarchy or city↔lat/lon.
### Other localized fields (what varies per locale)
- JS: `sv` overrides color, commerce, company, date, internet, location, person, phone_number, team; `en` defines 24 modules; `base` 9. So finance/vehicle/animal/food/science are effectively English-only.
- PY: per-locale providers exist for address, automotive (plate formats), bank (IBAN/BBAN/SWIFT by country), company (VAT IDs), currency, date_time (month/day names), internet (TLDs, free email domains, user-name formats), job, lorem (word lists), person, phone_number, ssn (national IDs with check digits), passport, color, barcode (EAN prefixes), misc.
- DF `de.yml` overrides: address, company, compass, creature, lorem, hipster, name, color, commerce, book, university, chuck_norris, space, music, games, food, simpsons, dr_who, vehicle. Country files (`_SE.yml`) mainly phone_number and address formats.
- MM: only address, datetime, finance, food, person, text are localized.
### Name weighting
| Lib | Weighted by frequency? |
|---|---|
| PY | Yes, default `use_weighting=True`. `en_US`: ~500 female + ~500 male first names weighted from SSA decade tables, ~1000 surnames weighted by US Census; `sv_SE`: ~1000/1000/1000 weighted from Skatteverket/ISOF. Plus weighted `address_formats`, `random_element(OrderedDict)`. |
| JS | No. Plain lists (`en`: ~650 female, ~550 male, ~400 generic first names). `helpers.weightedArrayElement` exists for user data only. |
| GO | No. ~1500 first, ~500 last, plain slices. `Weighted(options, weights)` for user data. |
| DF | No. `en`: ~1200 male, ~4400 female first names, ~1000 last, unweighted. |
| BG | No (faker.js data); `Randomizer.WeightedRandom` for user data. |
| MM | No. `en`: ~4000 female, ~3000 male, ~2500 surnames, unweighted. |
| CH | No; `weighted()` helper only. |
| RB | No. |
## 3. Common complaints (issues and posts)
- **Uncorrelated address parts** — the single most repeated request across every library:
- faker-ruby #1958 "Need an address object that has a street, city, zip that correlate" (geocoding breaks) https://github.com/faker-ruby/faker/issues/1958; #2881 "Consistent addresses" (state ≠ state_abbr, country ≠ country_code) https://github.com/faker-ruby/faker/issues/2881; #2899 "Fake Zip Codes" closed not-planned https://github.com/faker-ruby/faker/issues/2899; #275 zips like 90099 https://github.com/stympy/faker/issues/275
- fzaninotto/Faker #625 "Faking Legitimate Addresses" https://github.com/fzaninotto/Faker/issues/625
- java-faker #378 "Italy, Lima" / "USA, Moscow" https://github.com/DiUS/java-faker/issues/378; Datafaker #1551 wrong province–city for China https://github.com/datafaker-net/datafaker/issues/1551
- gofakeit #196 "city may not be in the same state as the zip" https://github.com/brianvoe/gofakeit/issues/196
- Bogus #481 invalid US zips https://github.com/bchavez/Bogus/issues/481
- Python Faker #697 "city in country" https://github.com/joke2k/faker/issues/697; #1812 `states_postcode` fails for territories
- **Synthetic city names** ("East Jarretmouth", "Nord Müller-stadt"): faker-js #2021 merged `city`/`cityName` and documented that only some patterns are real https://github.com/faker-js/faker/issues/2021; Bogus #342 `Address.City` returns people's names https://github.com/bchavez/Bogus/issues/342; faker-js PR #2141 / #2127 / #3792 adding "real" or "more realistic" cities per locale.
- **Localized data falling through to English**: faker-js #983 en_CA cities were US cities https://github.com/faker-js/faker/issues/983; Datafaker #1477 unresolved directive for city names https://github.com/datafaker-net/datafaker/issues/1477; #1358 ru-MD city names; Python #2432 vi_VN outdated administrative units.
- **Fields within one record don't match**: Python #1185 profile name vs username/email https://github.com/joke2k/faker/issues/1185; #1420 email consistent with name https://github.com/joke2k/faker/issues/1420 (both stale, no maintainer answer); gofakeit `Person()` picks gender and first name independently and email from a fresh name.
- **No frequency weighting / flat distributions**: blood types come out uniform when O+ and A+ are >70% of population (https://www.statology.org/how-to-validate-enhance-faker-profile-data-generation/); Python's `use_weighting` is the only built-in answer and only some locales supply weights; #1815 weighting silently dropped after adding a provider https://github.com/joke2k/faker/issues/1815.
- **Formats that produce impossible values**: Python #1849 US phone numbers that can't exist; #1868 fr_FR postcodes with <5 digits; faker-js #1159 unrealistic BIC; #3429 routing numbers fixed to use a real Federal Reserve lookup table; Ruby #1123 wrong pt-BR zip ranges.
- **No locales at all** in gofakeit (#352, open) and Chance.
- **"Column-by-column" generation, no relationships between columns/tables** as a general critique of Faker-style tools (https://securityboulevard.com/2026/04/best-synthetic-data-generation-tools-and-platforms-compared-for-2026/).
- **Uniqueness**: faker-js removed `faker.unique` and tells users to roll their own (https://fakerjs.dev/guide/unique.html); Python `fake.unique` raises `UniquenessException` after retries.
## 4. Notable and unusual features
- **Datafaker** — expression language `#{Name.first_name}`, `#{regexify '[a-z]{4,10}'}`, `#{numerify '##'}`, `#{bothify}`, `#{letterify}`, `#{templatify}`, `#{examplify}`, `#{options.option 'A','B'}`, `#{csv …}`, `#{json …}`, method calls with args (`#{date.birthday 'yy DDD'}`); nested resolution (https://www.datafaker.net/documentation/expressions/). `Schema.of(field(...), compositeField(...))` with transformers to CSV, JSON, YAML, XML, TOML, SQL (batch inserts, Postgres/Oracle/Spark dialects, ARRAY/MULTISET/ROW), `@FakeForSchema` POJO population (https://www.datafaker.net/documentation/schemas/). `faker.collection(...).len(3,5)`, infinite `stream()`, `faker.unique()`, sequences, custom yml via `addPath`. 263 providers, the largest pop-culture catalogue (≈110 franchises/games).
- **gofakeit** — Go `text/template` based `Template()`, `Markdown()`, `EmailText()`, `FixedWidth()`; `Generate("{firstname} {regex:[a-z]{5}} {randomstring:[a,b]}")` mini-language; `Struct(&v)` fill via `fake:"{city}"`, `fakesize`, `format` tags; `Slice`, `Map`, `Regex`; `CSV/JSON/XML/SQL` output; `AddFuncLookup` registers custom functions with parameter metadata (used by its CLI and HTTP server); `Weighted()`; word families split by grammatical role (46 Noun/Verb/Adjective/Pronoun/Connective sub-kinds); `Error*` generators; Minecraft.
- **Mimesis** — `Field`/`Fieldset`/`Schema` builder (`Field(locale)("person.full_name")`, keyed lookups with `key=` post-processors like `maybe`, `romanize`), `Schema(schema=lambda: {...}, iterations=n)` exporting to JSON/CSV/pickle, relational data with `SchemaRef` foreign keys, custom field handlers, factory-boy integration, fastest Python generator; locale = six JSON files; ISO country codes in A2/A3/numeric; DMS coordinates.
- **Bogus** — `Faker<T>().RuleFor(x => x.Prop, f => …)` fluent rules with `StrictMode`, `CustomInstantiator`, `FinishWith`, `Rules()`, `Ignore`, rule sets; local vs global seeding, `UseDateTimeReference`; `Bogus.Distributions.Gaussian`; country ID extension packages; premium Locations/Healthcare/Hollywood/Text packages; Roslyn analyzer; `AutoBogus` auto-fills any class.
- **faker-js** — locale fallback stack with explicit `null` = "not applicable"; `mergeLocales`; `helpers.fake('{{person.firstName}}')` template; `helpers.fromRegExp`, `weightedArrayElement`, `multiple`, `uniqueArray`, `mustache`; pluggable `Randomizer` (32/53-bit Mersenne), `Distributors` (uniform, exponential); `date.birthdate({mode:'age'|'year'})`; `location.nearbyGPSCoordinate`; `git` module; `person.bio`; typed per-locale definition files generated by script.
- **Python Faker** — `use_weighting` (frequency-weighted names/formats from census data); `OrderedDict` weighted `random_element`; `fake.unique`, `fake.optional`; multi-locale instance with weights; `geo.local_latlng` real geonames records; `misc` file/archive/CSV/JSON/image generators; `pystr_format`; pytest fixture; CLI; 24 community providers (healthcare, market data, observability, airtravel, education, security, pyspark …).
- **Chance** — `weighted`, `normal` (Gaussian), `unique`, `n`, `mixin`, `set` (override data), `pick/pickone/pickset`; mobile IDs (android_id, apple_token, bb_pin, wp7/8 anid); geohash/altitude/depth.
- **Ruby faker** — `full_address_as_hash`, `Faker::Config.locale`, `Faker::UniqueGenerator`, `Faker::Config.random`; `Types` generator for random Ruby types; company registration numbers for ~12 countries; deep pop-culture catalogue (~85 themed generators).
+145
View File
@@ -0,0 +1,145 @@
# Swedish (sv_SE) fake-data sources — research 2026-09-17
Scope: names, identifiers, phone, company, plates/cars, words, dates/misc, other. Geography excluded.
"Verified" = fetched and read this session. Local copies of inspected files: `/home/lilleman/.claude/jobs/e377c8d4/tmp/`.
## Licence summary (embed in MIT repo?)
| Source | Licence (exact) | Embed in MIT? | Attribution |
|---|---|---|---|
| SCB Statistikdatabasen + geodata | CC0 1.0 Universal | Yes | None required; SCB asks "Källa: SCB" for statistics taken from scb.se |
| SCB files on scb.se (xlsx: names, SNI, SSYK, SUN) | Not stated per file; scb.se terms: cite "Källa: SCB" | Yes (flag: CC0 statement names only statistikdatabasen/geodata) | "Källa: SCB" |
| Skatteverket testpersonnummer / testsamordningsnummer | CC0 1.0 (in DCAT) | Yes | None |
| Skatteverket namn på nyfödda, efternamn lists | No licence in DCAT/page; "fri att använda, inga avtal eller avgifter" (PSI law 2022:818) | Yes, flag | Cite Skatteverket |
| Bolagsverket statistics CSVs | CC BY 2.5 SE | Yes | Attribution required |
| Bolagsverket/SCB bulk company files | Not stated ("värdefulla datamängder", free, no agreement) | Derive stats only; do not embed raw | Cite |
| Bankinfrastruktur clearing CSV | Not stated; accuracy disclaimer | Facts only; flag | Cite BSAB |
| Bankgirot PDFs | "© Bankgirocentralen BGC AB. All rights reserved." (Informationsklass: Öppen) | No verbatim copy; facts only | — |
| PTS numbering plan PDF | Not stated (no licence on pts.se or in PDF) | Facts only; flag | Cite PTS |
| Transportstyrelsen blocked combinations | Not stated on page | Facts (a list of 99 codes); flag | Cite |
| SALDO morphology (Språkbanken) | CC BY 4.0 | Yes | Cite: Språkbanken (2017) SALDO's morphology, DOI 10.23695/agcm-ny22 |
| Språkbanken word statistics | CC BY 4.0 | Yes | Cite Språkbanken |
| Folkets lexikon | CC BY-SA 2.5 Generic | No (share-alike) | — |
| SAOL (svenska.se) | Copyright, written permission required | No | — |
| Hunspell sv_SE (DSSO) | MPL 1.1 / GPL 2 / LGPL 2.1 (tri); yeager/hunspell-sv LGPL-3.0 | Avoid | — |
| Mobility Sweden registrations | "ange källa" (no open licence) | Top-N facts only | Cite Mobility Sweden |
| Wikipedia (sv/en) | CC BY-SA 4.0 | No copying; facts only | — |
| CLDR (via ICU/Intl) | Unicode licence | Yes | Unicode notice |
## 1. Person names
### SCB (historic, frozen)
- SCB stopped producing name statistics from 2024 and refers to Skatteverket (Namnsök page, verified). Statistikdatabasen folder BE0001 now holds only "Äldre tabeller som inte längre uppdateras" (BE0001D newborns, BE0001G whole population).
- Whole-population xlsx (verified, downloaded): https://www.scb.se/contentassets/9fe7dbb460994c72b835163dbc491ef9/namn-med-minst-tva-barare-31-december-2022.xlsx — 13.1 MB, last-modified 2024-11-21, reference date 2022-12-31, names with >= 2 bearers, names in UPPERCASE.
| Sheet | Rows (incl. 4 header rows) | Columns |
|---|---|---|
| Efternamn | 411,802 | Efternamn, Antal bärare |
| Förnamn kvinnor | 91,247 | Förnamn, Antal bärare |
| Förnamn män | 79,128 | Förnamn, Antal bärare |
| Tilltalsnamn kvinnor | 57,785 | Tilltalsnamn, Antal bärare, Medelålder (only where >= 10 bearers, else "-") |
| Tilltalsnamn män | 49,404 | same |
- Gender split: a name appearing in both kvinnor/män sheets gives P(female) = count_k / (count_k + count_m).
- Middle names: "Förnamn" = every given name a person carries; "Tilltalsnamn" = the one used. Förnamn-minus-tilltalsnamn frequency approximates middle-name usage. Legal "mellannamn" (a second surname) is a different concept, no longer grantable under namnlagen 2016:1013 (not verified this session).
- Junk rows to filter: single letters ("A"), initials ("A-C", "K C", "JR"), "A:SON".
- Legacy PxWeb tables (folder BE0001G, verified via web UI): BE0001T06AR, BE0001TNamn10 (tilltalsnamn >= 10 bearers 1999–2020), BE0001T100, BE0001T08AR, BE0001T07AR, BE0001FNamn10 (förnamn >= 10 bearers), BE0001F100, BE0001T03Ar (efternamn top), BE0001ENamn10 (efternamn >= 10 bearers 1999–2020). PxWebApi v1 (`api.scb.se/OV0104/v1/doris/sv/ssd/START/BE/BE0001/BE0001G/...`) returns HTTP 400 and PxWebApi 2 (`/v2beta/api/v2/tables/BE0001T06AR`) returns "Non-existent table" — these tables are web-UI only now (verified). Use the xlsx.
- Licence: SCB terms page (verified): "Statistik och geodata som SCB tillgängliggör som öppna data i statistikdatabasen och i vår geodataplattform har licensen Creative commons 0 1.0 Universal, CC0" and "Om du använder eller sprider statistik från scb.se ska dock som källa alltid SCB anges". The xlsx sits on scb.se, not in statistikdatabasen — flag; treat as CC0 + "Källa: SCB".
### Skatteverket (current)
- Namn på nyfödda (2021+), annual: DCAT https://skatteverket.entryscape.net/store/26/metadata/416; CSV `fbf_namn_nyfodda.csv` https://skatteverket.entryscape.net/store/26/resource/417; JSON API https://skatteverket.entryscape.net/rowstore/dataset/da2556d0-c717-45e8-a1d8-3320161d3a7d/json?_limit=100&_offset=0 (47,529 rows, verified). Columns: rangordning, kön (Kvinna/Man), fodelsear, uppdateringsdatum, namn, gruppering (Kommun/…), grupperingsvärde (e.g. ALE), antal. Names in Title case. No licence field in DCAT — flag.
- Most common surnames (> 2,000 bearers, ~500 names), 2026 text file (verified, 517 lines, format `Andersson 211808`): https://www.skatteverket.se/download/18.70685bee19c85dd5dd03b59/1775049117448/Fria_efternamn_namn_antal_textfil_2026.txt (also two PDFs). Page: https://www.skatteverket.se/privat/folkbokforing/namn/bytaefternamn/sokblanddevanligasteefternamnen.4.515a6be615c637b9aa48e09.html. Licence not stated.
- Skatteverket general open-data terms (verified): "Vår öppna data är fri att använda och kräver varken några avtal eller innebär några avgifter" under lag 2022:818; no CC licence named except on the test-number datasets (CC0).
- Skatteverket publishes no whole-population first-name list as a file; its statistikportalen "Så många har ett visst namn" is an interactive search (page 404 on fetch; unverified).
Recommendation: embed SCB 2022 xlsx-derived lists (top-N per sheet with counts) + Skatteverket 2026 surname counts for freshness.
## 2. Identifiers
### Personnummer (Folkbokföringslagen 1991:481 18 §, verified via lagen.nu)
- 10 digits: YYMMDD + födelsenummer (3 digits) + kontrollsiffra. Födelsenummer odd = male, even = female. Separator "-" ; "+" once the person turns 100. 12-digit form YYYYMMDDNNNC (no separator).
- Check digit: Luhn/mod 10 over the 9 digits YYMMDDNNN (weights 2,1,2,1,… from the left; sum digits of two-digit products; check = (10 − sum mod 10) mod 10).
- Before 1990 the first two födelsenummer digits encoded county; now random (Wikipedia, unverified officially).
### Samordningsnummer
- Same layout; day + 60 (3rd → 63); individual number random 001–999, odd male / even female; same Luhn (Skatteverket rättslig vägledning, via search excerpt; page itself blocked). Now regulated by Lag (2022:1697) om samordningsnummer (18 a § FBL repealed) — flag.
- Example from Skatteverket: man born 1970-10-03, no. 239 → 701063-2391.
### Official test numbers (Skatteverket, CC0 1.0)
- Testpersonnummer dataset: https://www.dataportal.se/sv/datasets/6_67959/testpersonnummer (DCAT: https://admin.dataportal.se/store/6/metadata/67959?recursive=dcat). "Testpersonnummer med sekelsiffror för åren 1890-2025. Nya testnummer för kommande år läggs ut i december varje år." Formats: Excel (one file), CSV (split into many period files), JSON API https://skatteverket.entryscape.net/rowstore/dataset/b4de7df7-63c0-4e7e-bb59-1f156a591763/json?_limit=100 — resultCount 43,895 (verified), one column `testpersonnummer`, 12 digits.
- Observed pattern (verified on 1890s CSV file fully and API samples at offsets 0/10000/30000/43880): every day gets two numbers, födelsenummer 238 (female) and 239 (male), through 2026-12-31; 1890s use 980/981 (plus a few 982/983/936). All Luhn-valid. Skatteverket blocks these from ever being assigned.
- Testsamordningsnummer: DCAT https://skatteverket.entryscape.net/store/9/metadata/153; years 1914–2025; CSV per year; sample file 2022 (1,498 rows): days 61–91, individual numbers 238/239, Luhn-valid.
- Should a generator restrict to these? Skatteverket does not mandate it, but only these are guaranteed never to collide with a real person; any other Luhn-valid number may be real. Recommended: default to the official series (date + 238/239, or 980/981 for 1890s) and embed nothing but the rule (~530 kB if the full list were embedded). Flag: the 238/239 rule was sampled, not exhaustively verified for every year.
- Common practice elsewhere (Faker etc.): random date + random 3 digits + Luhn — valid but collision-prone.
### Organisationsnummer
- 10 digits NNNNNN-NNNN, last = Luhn over first 9 (same algorithm as personnummer). Digits 3–4 ("month") are always >= 20, distinguishing from personnummer (Wikipedia; not found on an official page — flag).
- First digit (gruppnummer), Bolagsverket page (verified): 5 = aktiebolag, filialer, banker, försäkringsbolag, europabolag; 9 = handelsbolag, kommanditbolag; 7 or 8 = bostadsrättsföreningar, ekonomiska föreningar, näringsdrivande ideella föreningar etc.; 2 or 8 = trossamfund; 20… = state agencies (Bolagsverket itself 202100-5489); 3 = foreign companies. Wikipedia adds: 1 = dödsbon, 2 = stat/region/kommun, 6 = samfällighetsföreningar, 8 = ideella föreningar/stiftelser. Enskild firma = owner's personnummer.
- VAT number (Skatteverket, verified): "SE" + 12 digits = organisationsnummer (10) + "01" ("De två sista siffrorna är alltid 01"). Written SE556047352101.
### Bankgiro / Plusgiro (Bankgirot "10-modul" PDF 2016-12-01, verified)
- Bankgironummer: 7 or 8 digits, last digit mod-10 (Luhn) check; printed 991-2346 (7) / 5555-5551 (8), i.e. hyphen before the last four digits. 90-konton (charity) 900-000x…904-999x, always 7 digits (Wikipedia).
- Plusgironummer: 2–8 digits, last digit Luhn (Wikipedia/samlogic; not from an official Nordea page — flag). Printed with hyphen before check digit, e.g. "12 34 56-7" (convention, unverified).
### Clearing numbers / bank accounts
- Bankinfrastruktur i Sverige AB CSV (verified, 57 rows, 40 actors): https://www.bankinfrastruktur.se/media/1melztro/tabell-over-clearingnummer-250305.csv (also `/media/sqfjsvkp/clearingnummertabell-for-nedladdning.csv`, 63 rows, newer, includes Zimpler 2130-2139). Columns: `Clearingnummer;Aktör;BIC;IBAN ID;Konto-typ;Metod IBAN konvertering`. Semicolon, UTF-8 BOM. Page: https://www.bankinfrastruktur.se/framtidens-betalningsinfrastruktur/iban-och-svenskt-nationellt-kontonummer. No licence; disclaimer "garanterar inte att publicerade uppgifter är korrekta".
- Key ranges: Nordea 1100–1199, 1400–2099, 3000–3399, 3410–3999, 4000–4999 (3300 and 3782 = personkonto); Danske 1200–1399, 2400–2499; SEB 5000–5999, 9120–9124, 9130–9149; Handelsbanken 6000–6999; Swedbank 7000–7999 (4-digit) and 8000–8999 (5-digit clearing with own check digit); Länsförsäkringar 3400–3409, 9020–9029, 9060–9069; Skandiabanken 9150–9169; ICA 9270–9279; Avanza 9550–9569; Nordnet 9100–9109; Klarna 9780–9789.
- Bankgirot "Bankernas kontonummer" 2024-02-22 (verified text): Typ 1 = 4-digit clearing + 7-digit account (11 digits) with mod-11 check, weights 1,10,9,…,1; comment 1 = weigh clearing minus first digit + 7 digits; comment 2 = whole clearing + 7 digits. Typ 2 = clearing not part of the account: Handelsbanken 9-digit account mod-11 (comment 2); Swedbank 8000–8999 up to 10 digits mod-10 (comment 3); Danske 9180–9189, Nordea personkonto 3300/3782, Sparbanken Syd 9570–9579 10 digits mod-10 (comment 1).
- IBAN SE (Bankinfrastruktur page, verified): 24 chars = "SE" + 2 check digits + 3-digit IBAN ID (from CSV column, e.g. Nordea 300, SEB 500, Handelsbanken 600, Swedbank 800, Danske 120, LF 902) + 17-digit account (zero-padded; clearing included/excluded per "Metod IBAN konvertering" 1–3).
## 3. Phone
- Source: PTS "The Swedish numbering plan for telephony according to ITU-T E.164", 2024-01-08 (verified, text extracted): https://pts.se/globalassets/globala-block/nummertillstand/the-swedish-numbering-plan-for-telephony-according-to-itu---2024-01-08.pdf. Swedish "nrplansammanstallning-2026-02-04.pdf" and https://nummer.pts.se/NbrPlanSearch are behind a Radware bot check (both curl and Playwright blocked). No licence stated by PTS anywhere found; pts.se says contact pts@pts.se for data.
- 264 geographic area codes (PDF text yields 263 "Area code for" rows; one likely split across a page break — flag). Area code = trunk "0" + NDC of 1–3 digits (08 Stockholm, 031 Göteborg, 040 Malmö, 0480 Kalmar). Geographic N(S)N (NDC + subscriber) max 9 digits, min 7 (2-digit NDC) or 8 (3-digit NDC); Stockholm min 7. Each row in the PDF gives NDC, max/min length, "Area code for <name>", so a `{ndc, name}` list can be built from it.
- Mobile: NDC 70, 72, 73, 76, 79 — N(S)N exactly 9 digits, i.e. 07X-XXX XX XX (10 digits with trunk 0). Others: 71 mobile broadband/M2M (13 digits), 74 paging, 75 personal numbering, 77 shared cost, 20 freephone, 10 location-independent (10 AXX XX XX, A=1–8), 378 M2M fixed (10 digits).
- Formatting (Språkrådet/Isof frågelådan, via search excerpts): hyphen after area code, subscriber digits grouped 2–3 from the left: `08-668 01 50`, `031-123 45 67`, `0480-123 45`, `070-123 45 67` (also `0701-23 45 67` accepted). International: `+46 70 123 45 67` (drop trunk 0).
- Wikipedia "Lista över riktnummer i Sverige" cites the same PTS document (CC BY-SA — do not copy).
## 4. Company
- Legal forms + suffixes (Bolagsverket, lag 2018:1653 om företagsnamn): aktiebolag must contain "aktiebolag" or "AB" (verified); public companies add "(publ)"; handelsbolag "HB" / "handelsbolag", kommanditbolag "KB" / "kommanditbolag", ekonomisk förening "ekonomisk förening" / "ek. för.", enskild firma no suffix — statutory wording for HB/KB/ek.för. not verified on the fetched page (flag).
- Legal-form codes (SCB "juridisk form", from Skatteverket organisationsnummer page and SCB variabelbeskrivning): 10 fysiska personer, 21 enkla bolag, 22 partrederier, 23 värdepappersfonder, 31 handelsbolag/kommanditbolag, 32 gruvbolag, 41 bankaktiebolag, 42 försäkringsaktiebolag, 43 europabolag, 49 övriga aktiebolag, 51 ekonomiska föreningar, 53 bostadsrättsföreningar, 54 kooperativ hyresrättsförening, 61 ideella föreningar, 62 samfälligheter, 63 registrerade trossamfund, 71 familjestiftelser, 72 övriga stiftelser, 81 statliga enheter, 82 kommuner, 84 regioner, 87 offentliga korporationer, 88 hypoteksföreningar, 91 oskiftade dödsbon, 92 ömsesidiga försäkringsbolag, 93 sparbanker, 94 understödsföreningar, 95 arbetslöshetskassor, 96 övriga svenska juridiska former, 98 utländska juridiska personer — codes beyond 49 quoted from memory of the SCB list, verify against https://www.scb.se/contentassets/8a8eb5c3d45f461ea93482f8e8d4de4f/variabelbeskrivning-api.pdf (flag).
- Distribution (Bolagsverket ftgstat_oppna.csv, CC BY 2.5 SE, 404,779 rows, latin-1, verified): https://www.bolagsverket.se/statistik/ftgstat_oppna.csv. Columns: ar, manad, handelse (1 = nyregistrerade, 2 = alla registrerade, 3 = avslutade), regfam, län/kommun, then counts per form: AB, BAB (bankaktiebolag), BF (bostadsförening), BRF, EK (ekonomisk förening), E (enskild näringsidkare registered at Bolagsverket only), SE (europabolag), FL (filial), FAB (försäkrings-AB), HB, I (näringsdrivande ideell förening), KB, KHF, MB, SF, SB (sparbank), TSF (trossamfund), BFL, OFB, SCE, S, EGTS, FOF, TPAB, OTPB, TPF. New registrations 2025 (computed): AB 50,859; E 8,468; HB 1,557; BRF 398; EK 319; FL 269; KB 176; I 59. Column description xlsx: https://www.bolagsverket.se/download/18.46f4138717c599ee403aac05/1638951733112/beskrivning-av-csv-fil.xlsx.
- Company-name register: Bolagsverket "värdefulla datamängder" bulk file https://vardefulla-datamangder.bolagsverket.se/bolagsverket/bolagsverket_bulkfil.zip (250 MB, updated 2026-09-14; org.nr, names, legal form, business description, address) and SCB bulk https://vardefulla-datamangder.bolagsverket.se/scb/scb_bulkfil.zip (71 MB; adds SNI up to 5 codes, status). Free, no agreement; licence not stated on the page (statistics CSVs are CC BY 2.5 SE) — flag. Use offline to derive name-token frequencies; do not embed rows. API (org.nr lookup only; name search "planned, no date"): https://bolagsverket.se/apierochoppnadata/hamtaforetagsinformation/apiforatthamtaforetagsinformation.3988.html.
- Naming rules (Bolagsverket "Välja företagsnamn", verified): must be distinctive — not only a generic activity ("Bilverkstad AB") or a bare surname; no confusion with existing names/trademarks; no professional titles without credentials; no apparent internet address; "aktiebolag/AB/handelsbolag" ignored when comparing names. Typical patterns (unverified heuristic, derive from bulk file): `<Surname> <Trade> AB`, `<Surname>s <Trade>`, `<Place> <Trade> AB`, `<Fantasy word> AB`, `<Initials> <Trade> HB`, suffix words Bygg, Konsult, Holding, Invest, Fastigheter, Förvaltning, Entreprenad, Teknik, Design.
- SNI 2007 (SCB, verified xlsx, 94 kB): https://www.scb.se/contentassets/d43b798da37140999abf883e206d0545/sni2007.xlsx — sheets Detaljgrupp (5-digit, 821 codes), Undergrupp (615), Grupp (272), Huvudgrupp (88), Avdelning (21 letters). Code format `01.110` + Swedish label. SNI 2025 replaces it from 2025-12-08 (835/651/287/87/22): https://www.scb.se/globalassets/sni-2025.xlsx (93 kB). Skatteverket/Bolagsverket transition status unverified. Licence: no statement on the classification pages; SCB terms apply ("Källa: SCB") — flag.
## 5. Licence plates and cars
- Format: `ABC 123` (since 1973) and `ABC 12A` (since 2019-01-16; last character a letter, never O). Letters never used: I, Q, V, Å, Ä, Ö. Digits 001–999 (000 not issued). Personal plates 2–7 characters, may use any letters, must not mimic the standard pattern. Source: sv.wikipedia Registreringsskyltar i Sverige (facts; not confirmed on a Transportstyrelsen page — flag).
- Blocked letter combinations, Transportstyrelsen (verified, page updated 2026-03-31, 99 codes): https://www.transportstyrelsen.se/sv/vagtrafik/fordon/aga-kopa-eller-salja-fordon/registreringsskyltar/byte-av-registreringsnummer/Sparrade-bokstavskombinationer/
APA ARG ASS BAJ BSS CUC CUK CUM DUM ETA ETT FAG FAN FEG FEL FEM FES FET FNL FUC FUK FUL GAM GAY GEJ GEY GHB GUD GYN HAT HBT HKH HOR HOT KGB KKK KUC KUF KUG KUK KYK LAM LAT LEM LOJ LSD LUS MAD MAO MEN MES MLB MUS NAZ NRP NSF NYP OND OOO ORM PAJ PKK PLO PMS PUB RAP RAS ROM RPS RUS SEG SEX SJU SOS SPY SUG SUP SUR TBC TOA TOK TRE TYP UFO USA WAM WAR WWW XTC XTZ XUK XXL XXX ZEX ZOG ZPY ZUG ZUP ZOO
- Car makes, Mobility Sweden 2025 full-year new passenger cars (272,987 total; press release 2026-01-02 "Svagt fordonsår 2025", via search excerpt; "ange källa" required): Volvo 48,961; Volkswagen 38,677; Toyota 22,189; Kia 19,922; Mercedes-Benz 17,372; Skoda 16,664; BMW 15,001; Audi 14,887; Peugeot 9,059; Polestar 7,601. Top models 2025: Volvo XC60 17,933, Volvo EX40, VW ID.7, Tesla Model Y 5,822 (press excerpts). Monthly xlsx: https://mobilitysweden.se/statistik/Nyregistreringar_per_manad_1. Not an open licence — embed as a curated list with attribution.
- Transportstyrelsen fordonsstatistik (open data, by vehicle type only, not make): https://www.transportstyrelsen.se/sv/om-oss/statistik-och-analys/statistik-inom-vagtrafik/fordonsstatistik/. Trafa (https://www.trafa.se/vagtrafik/fordon/) publishes by owner/technical/region, no make/model list found; licence not stated.
## 6. Words
- SALDO morphology (recommended): https://sprakbanken.se/en/resources/saldom — CC BY 4.0, 128,036 entries, LMF XML `saldom.xml` 254,107,385 bytes (https://svn.spraakbanken.gu.se/sb-arkiv/pub/lmf/saldom/saldom.xml, 2017-09-19). Fields: lemma, POS, inflection paradigm, full word forms with MSD. Derive a small `{lemma, pos}` list (nouns/adjectives/verbs) offline; embed with the citation "Språkbanken (2017). SALDO's morphology. https://doi.org/10.23695/agcm-ny22".
- Frequencies: Språkbanken word statistics https://sprakbanken.se/resurser/ordstatistik — `stats_all.txt.zip` 763.87 MB, CC BY 4.0, 2025-04-22, aggregated over modern corpora (Korp). Column layout not confirmed (FAQ entry not found; per-corpus files are tab-separated, 6 columns incl. word form, POS, lemgram, SALDO sense, compound flag, frequency) — flag; inspect a per-corpus stats file first (smaller).
- Folkets lexikon (KTH): CC BY-SA 2.5 — share-alike, not suitable for an MIT-embedded derivative. SAOL: svenska.se "får inte användas … utan skriftligt tillstånd" — unusable. Hunspell sv_SE/DSSO: MPL/GPL/LGPL — avoid.
## 7. Dates, money, email, domains
- CLDR sv-SE (verified with Node 24.18.1 ICU): short date `2026-09-17`; long `17 september 2026`; full `torsdag 17 september 2026 kl. 16:30`; months januari…december (lowercase), short `jan. feb. mars apr. maj juni juli aug. sep. okt. nov. dec.`; weekdays måndag…söndag, short `mån tis ons tors fre lör sön`; currency `1 234,56 kr`; number `1 234 567,891`. ICU uses U+00A0 (NBSP) as group separator; typographic guidance says hard space, decimal comma.
- Språkrådet (search excerpts): ISO 8601 `ÅÅÅÅ-MM-DD` in formal/technical text, `6 juli 2022` in running text; clock `14.30` (period; colon accepted); amounts `1 234,50 kr`.
- Public holidays, Lag (1989:253) om allmänna helgdagar (verified): nyårsdagen 1/1, trettondedag jul 6/1, långfredagen, påskdagen, annandag påsk, Kristi himmelsfärdsdag (6th Thursday after Easter), pingstdagen, nationaldagen 6/6, midsommardagen (Saturday 20–26 June), alla helgons dag (Saturday 31 Oct–6 Nov), juldagen 25/12, annandag jul 26/12, plus Sundays. Semesterlagen 3 a § (verified): "Med söndag jämställs allmän helgdag samt midsommarafton, julafton och nyårsafton". Easter needs computing (Gregorian).
- Email domains: no official statistic found (Internetstiftelsen "Svenskarna och internet" does not report providers). Common providers by general knowledge only (unverified): gmail.com, hotmail.com, outlook.com, icloud.com, yahoo.com/yahoo.se, live.se, telia.com (Telia is phasing out its mail service), bredband.net, comhem.se, spray.se, me.com. Flag as curated.
- TLDs (Internetstiftelsen registrar stats JSON, verified): .se active domains 1,475,495 (2026-07); .nu 203,940 (2026-07). https://internetstiftelsen.se/registrar-stats/se/activedomainsmonthend.json, `/nu/…`. Licence not stated.
## 8. Other
- Occupations, SSYK 2012 (SCB, verified xlsx): https://www.scb.se/contentassets/0c0089cc085a45d49c1dc83923ad933a/ssyk-2012-koder.xlsx (50 kB; sheets 1/2/3/4-siffer; 429 four-digit codes) and job-title index https://www.scb.se/contentassets/0c0089cc085a45d49c1dc83923ad933a/ssyk_index_webb_2026.xlsx (213 kB, 2026-03-18, thousands of yrkesbenämningar → code). Licence: SCB terms ("Källa: SCB") — flag as with SNI.
- Education, SUN 2020 (SCB): https://www.scb.se/contentassets/aeeedec0e28c465aa524429407dcd5ba/sun-2020_niva_inriktning2.xlsx (304 kB; levels + orientations). Population 25–64 (SCB "Befolkningens utbildning 2024", via search excerpt): ~10 % förgymnasial, ~40 % gymnasial, ~31 % eftergymnasial >= 3 years (remainder shorter post-secondary). Approximate — flag.
- Blood groups (geblod.nu 2007 via sv.wikipedia Blodgruppsfördelning; geblod.nu today only says RhD+ ≈ 85 %): A+ 37, O+ 32, B+ 10, AB+ 5, A− 7, O− 6, B− 2, AB− 1 (%).
- Regional/kommun data: out of scope (geography).
## Unverified / flagged items
1. CC0 applies by SCB's statement to statistikdatabasen/geodata; xlsx files on scb.se (names, SNI, SSYK, SUN) carry only "Källa: SCB" terms.
2. Skatteverket name datasets and surname file: no licence field; only "fri att använda" under PSI law.
3. Testpersonnummer 238/239 rule: sampled (1890s file fully, offsets 0/10k/30k/43.9k), not exhaustively checked; per-period CSV list has many files.
4. Organisationsnummer "digits 3–4 >= 20" and group digits 1/6 — Wikipedia only.
5. Plusgiro length/format — Wikipedia/blog only.
6. PTS: no licence; 264 vs 263 area codes counted in the PDF text; Swedish 2026 PDF and nummer.pts.se blocked (Radware).
7. Plate format details (000 excluded, O excluded as last letter, 2019-01-16) — Wikipedia only.
8. Legal-form codes above 49 — from memory; verify against the SCB variabelbeskrivning PDF.
9. Word-statistics column layout; email-provider list; company-name patterns; education shares — heuristic/approximate.
+616
View File
@@ -0,0 +1,616 @@
# Open data sources and format rules — US/locale-neutral fake-data categories
Researched 2026-09-17 for embedding as data files in an MIT-licensed library. Geography/addresses and Swedish data are out of scope (covered elsewhere). Each item that was not read first-hand from the cited page/file is marked **UNVERIFIED**.
## Licence key (what each licence means inside an MIT repo)
| Licence | Embed? | Obligation |
|---|---|---|
| Public domain (US federal work, 17 USC §105), CC0 1.0, PDDL, Unlicense | Yes | None; cite the source in the data file header as courtesy |
| MIT, BSD-2, ISC | Yes | Keep the copyright + permission notice next to the copied data |
| Unicode License v3 (OSI-approved 2023-11-17, verified) | Yes | Notice in LICENSE/NOTICE or docs |
| WordNet License (verified at wordnet.princeton.edu) | Yes | Notice + disclaimer on all copies; no Princeton name in advertising |
| W3C Document License 2023 | Yes (data tables) | Notice "includes material copied from or derived from <title> <URI>" |
| CC BY 3.0 / 4.0 | Yes | Creator, licence link, indicate changes (verified deed) |
| Apache-2.0 (code), LGPL-2.1/3.0 (data) | Awkward | LGPL data files stay LGPL; keep them out of an MIT repo, use as reference for facts only |
| CC BY-SA, ODbL, CC BY-NC-SA, GFDL | No | Share-alike / non-commercial — incompatible |
| No licence stated, "all rights reserved", proprietary directories (SWIFT BIC, Fed RTN, CUSIP, ICAO 8643, GS1 page, ISO OBP, SAE WMI, whatismybrowser, useragents.me, MySQL manual, Pantone) | No | Generate synthetically from the format rules instead |
## Embed verdict at a glance
| Category | Use (clean) | Use with notice | Avoid |
|---|---|---|---|
| 1 Names/IDs | SSA baby names (CC0), Census 2010 surnames (PD), IRS EIN prefixes / ITIN ranges, SSA SSN rules | — | — (prefix/suffix lists: hand-curate) |
| 2 Finance | ISO 4217 via datasets/currency-codes (PDDL) or SIX list-one.xml; algorithms (ABA, Luhn, ISIN, CUSIP, IBAN mod 97, EIP-55, Bech32) | npm currency-codes (MIT), GLEIF BIC-LEI (custom notice), CLDR currency symbols (Unicode-3.0) | Fed routing directory, SwiftRef BIC directory, real CUSIP lists, php-iban (LGPL) verbatim |
| 3 Company | NAICS 2022 (PD), SEC company_tickers.json (free reuse), GLEIF ELF list (CC0) | — | — |
| 4 Internet | All IANA registries (CC0): media types, HTTP status, ports, TLDs, IPv4/IPv6 special-purpose; RFC 5737/3849/9637 ranges; locally-administered MAC | mime-db (MIT), Linguist languages.yml (MIT), top-user-agents (MIT) / user-agents (BSD-2), IEEE OUI (IEEE "does not assert copyright" — second-hand) | useragents.me, whatismybrowser, MySQL manual tables |
| 5 Vehicle | NHTSA vPIC makes/models/WMIs/value lists (PD); VIN check-digit rule | — | SAE WMI DB, openalpr plate patterns (AGPL), Wikipedia lists verbatim |
| 6 Airline | OurAirports (PD/Unlicense), Wikidata airlines + aircraft types (CC0) | OPTD airlines (CC BY 4.0) | OpenFlights + npm mirrors (ODbL), ICAO Doc 8643 + scraped mirrors, OpenSky aircraft DB |
| 7 Science | PubChem periodic table (PD), NASA planetary fact sheet (PD), Wikidata animals/plants/moons (CC0), Stanford/AABB blood types (facts) | BIPM SI brochure (CC BY 4.0), GBIF backbone (CC BY 4.0 + citation), faker-js lists (MIT) | Bowserinator periodic table (CC BY-SA 3.0) |
| 8 Text | Moby POS (PD), 12dicts PD lists, XKCD colours (CC0), Lorem ipsum/Cicero, Bartlett via Gutenberg (PD) | WordNet 3.x (WordNet License), Open English WordNet (CC BY 4.0), SCOWL (MIT-like), CSS named colours (W3C notice), Google Books ngrams (CC BY 3.0) | wordfreq data (CC BY-SA), hermitdave (CC BY-SA), Norvig count_1w (LDC-derived), Wikiquote |
| 9 Dates | IANA tzdb (PD) | CLDR cldr-dates-full (Unicode-3.0) | — |
| 10 ISO | datasets/country-codes (PDDL), LoC ISO 639-2 (PD, UNVERIFIED layout), Census state.txt (PD) | annexare/Countries, i18n-iso-countries, langs, iso-639-1 (all MIT) | mledoze (ODbL), lukes + stefangabos (CC BY-SA), Debian iso-codes (LGPL), SIL 639-3 table (no-redistribution clause), olahol iso-3166-2 (unclear) |
| 11 Commerce | Check-digit algorithms (EAN/UPC/GTIN, ISBN-10/13, ISSN, IMEI Luhn); GS1 demo prefix 952 | faker-js en commerce word lists (MIT) | GS1 prefix page + ISBN range file verbatim, Wikipedia tables verbatim |
| 12 Books/music/food | Gutenberg pg_catalog.csv (CC0), MusicBrainz genre names (CC0), ID3v1 genres, Wikidata instruments/dishes/cuisines/cocktails (CC0), USDA FoodData Central (CC0) | Open Brewery DB (MIT) | MusicBrainz tag associations (CC BY-NC-SA), IBA site text |
## Corrections to the brief's assumptions
- Bowserinator/Periodic-Table-JSON is CC BY-SA 3.0 (not MIT) and has 119 rows; use PubChem CSV (118, PD).
- wordfreq code is Apache-2.0 and its data CC BY-SA 4.0 — avoid the data.
- CSS Color 4 has 148 named colours (rebeccapurple included).
- Gutenberg #27889 Bartlett is the 9th edition (1903); the 1919 10th is not on Gutenberg (UNVERIFIED elsewhere).
- CLDR file path is `cldr-dates-full/main/<locale>/ca-gregorian.json`; current package 48.2.0.
- Diners Club US/Canada is a Mastercard co-brand on 55; replace the loose "Maestro 50/56–58/6xxx" rule with the explicit IINs in part 2.
- Maestro/Discover/Diners/JCB/UnionPay lengths vary 12–19 — generate per the table, not a fixed 16.
- The 987-65-4320..4329 SSN advertising range is repeated by third parties but not found on any ssa.gov/POMS page (UNVERIFIED).
- DUNS dropped its mod-10 check digit in 2006 — no validation rule today.
- IEEE's OUI "does not assert copyright" statement is only recorded second-hand (Debian/FSF); no current IEEE page restates it.
- Sites blocking automation (403/Cloudflare): ssa.gov, www2.census.gov, swift.com, redcrossblood.org, loc.gov, gs1.org (WebFetch), standards-oui.ieee.org (needs browser UA). Download those manually with a browser UA.
## Recommended default picks for an en_US library
- Given names: SSA `names.zip` aggregated to top-N per sex (CC0). Surnames: Census `Names_2010Census.csv` top-N (PD).
- Countries: `datasets/country-codes` CSV (PDDL); flag emoji derived from alpha-2. Currencies: same file's ISO 4217 columns or `datasets/currency-codes`; symbols from CLDR `en` (Unicode-3.0).
- Internet: IANA registries + mime-db; UAs from templates (no data) or `top-user-agents` (MIT).
- Vehicles: vPIC `GetMakesForVehicleType/car` (195) + models per make; VIN generated with check digit.
- Airports: OurAirports rows with `iata_code` and type large/medium (4,569). Airlines: Wikidata (CC0).
- Words: WordNet 3.0 index files (with LICENSE) or Moby POS (PD) for POS-tagged words; XKCD + CSS for colours.
- Timezones: `zone1970.tab` names only; offsets via `Intl` at runtime.
- Check digits: implement from the rules in parts 1, 2 and 4 — no data needed.
---
# Part 1 — Names, finance, company: sources and rules
Researched 2026-09-17. "Embed" = ship as a data file in an MIT repo. Everything not marked UNVERIFIED was read from the cited page or file this session.
## 1. US person names and identifiers
### SSA baby names (given names by year and sex)
- Page: https://www.ssa.gov/oact/babynames/limits.html — national zip https://www.ssa.gov/oact/babynames/names.zip (7 MB), state zip https://www.ssa.gov/oact/babynames/state/namesbystate.zip (23 MB), territories 228 kB.
- Format (per SSA page + readme as mirrored by https://github.com/hackerb9/ssa-baby-names): one `yobYYYY.txt` per year 1880–2025, CSV `name,sex,count`, name 2–15 chars, sex `M`/`F`, sorted by sex then count desc. Names with < 5 occurrences in a year are excluded.
- Counts: ~100,364 distinct names, ~2.02 M rows across all years (hackerb9 snapshot; year of snapshot UNVERIFIED). Top 1000 names cover 71.5 % of 2025 births (SSA page).
- Licence: data.gov catalog entry https://catalog.data.gov/dataset/baby-names-from-social-security-card-applications-national-data lists licence **CC0 1.0** (`https://creativecommons.org/publicdomain/zero/1.0/`), publisher SSA. US federal work, no attribution required; embeddable.
- Note: ssa.gov returns 403 to non-browser user agents (curl/WebFetch); download with a browser UA or via the data.gov link.
### Census 2010 surnames
- Page: https://www.census.gov/topics/population/genealogy/data/2010_surnames.html
- Files: full set https://www2.census.gov/topics/genealogy/2010surnames/names.zip (CSV + XLSX, `Names_2010Census.csv`), top 1000 `Names_2010Census_Top1000.xlsx`, docs https://www2.census.gov/topics/genealogy/2010surnames/surnames.pdf.
- Fields (from surnames.pdf): `name, rank, count, prop100k, cum_prop100k, pctwhite, pctblack, pctapi, pctaian, pct2prace, pcthispanic`. `(S)` = suppressed percentage. Names are UPPERCASE.
- Count: 162,253 surnames occurring ≥ 100 times (covers 90.1 % of people with a recorded surname); the CSV also has a trailing `ALL OTHER NAMES` row (UNVERIFIED — www2.census.gov returned 403 to curl this session; download with a browser).
- Licence: US Census Bureau work → US federal government work, public domain in the US (17 U.S.C. §105). Census only asks for citation (https://www.census.gov/about/policies/citation.html: "U.S. Census Bureau, [Table], [Product], [Vintage], [URL], accessed on [date]"). Embeddable; cite source in the data file header.
### Name prefixes / suffixes
- No authoritative enumerated list exists. USPS Publication 28 §33 (https://pe.usps.com/text/pub28/28c3_011.htm) defines the fields: Name Prefix, First Name, Middle Name or Initial, Surname, **Suffix Title** = "maturity (e.g., JR, SR) and professional (e.g., PHD, DDS) suffixes"; example "MR. WALTER W. WITHERSPOON JR.".
- Recommendation: hand-curate. Prefixes `Mr., Mrs., Ms., Mx., Dr., Rev.`; suffixes `Jr., Sr., II, III, IV, PhD, MD, DDS, Esq.`. Rule of thumb: generational suffix (Jr/Sr/II–IV) only with a male name in traditional US usage; never combine Jr. with Sr.; at most one generational suffix.
### SSN
- Structure: `AAA-GG-SSSS` (area 3, group 2, serial 4). No check digit (Wikipedia).
- Randomization since **2011-06-25** (https://www.ssa.gov/employer/randomization.html): "Previously unassigned area numbers were introduced for assignment excluding area numbers 000, 666 and 900-999." Area no longer geographic.
- Invalid per SSA POMS RM 10201.035 (https://secure.ssa.gov/apps10/poms.nsf/lnx/0110201035): area `000`, `666`, `900–999`; group `00`; serial `0000`.
- Generator rule: area ∈ [001,899] \ {666}; group ∈ [01,99]; serial ∈ [0001,9999].
- Advertising SSNs (https://www.ssa.gov/history/ssn/misused.html): `078-05-1120` (Woolworth wallet card, 1938; > 40,000 people claimed it) and `219-09-9999` (1940 Social Security Board pamphlet). Both are structurally valid — exclude from output.
- `987-65-4320`–`987-65-4329` "reserved for advertising": repeated by searchbug.com, ssn-check.org, a Missouri state standard (https://oa.mo.gov/sites/default/files/CC-SocialSecurityNumberNamingStandardV050305.pdf, which even misprints it as 987-54-4329). **UNVERIFIED** — not found on any ssa.gov page or POMS this session; Wikipedia no longer carries it. Safe to use these ten as the library's "obviously fake" examples, but do not cite SSA for it.
### ITIN (IRS)
- IRS Publication 4757 (Rev. 6-2026), p. 5 (https://www.irs.gov/pub/irs-pdf/p4757.pdf): "All valid ITINs are nine-digit numbers in the same format as the SSN (9XX-8X-XXXX), beginning with a "9" and the 4th and 5th digits ranging from 50 to 65, 70 to 88, 90 to 92, and 94 to 99."
- Group `93` = ATIN (adoption TIN); `89` reserved (Intuit/TaxSlayer support pages, not IRS — UNVERIFIED).
- Generator rule: `9` + 2 digits + group ∈ {50–65, 70–88, 90–92, 94–99} + 4 digits (serial `0000` not documented as invalid for ITIN; avoid it anyway).
### EIN (IRS)
- Format `NN-NNNNNNN`. Page: https://www.irs.gov/businesses/small-businesses-self-employed/how-eins-are-assigned-and-valid-ein-prefixes ("The first two digits of your EIN show the IRS campus that assigned it or if you applied online").
- Valid prefixes (union of the table, 2026-09-17):
`01 02 03 04 05 06 10 11 12 13 14 15 16 20 21 22 23 24 25 26 27 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 71 72 73 74 75 76 77 80 81 82 83 84 85 86 87 88 90 91 92 93 94 95 98 99`
By campus: Andover 10,12 · Atlanta 60,67 · Austin 50,53 · Brookhaven 01–06,11,13,14,16,21–23,25,34,51,52,54–59,65 · Cincinnati 30,32,35–38,61 · Fresno 15,24 · Kansas City 40,44 · Memphis 94,95 · Ogden 80,90 · Philadelphia 33,39,41–43,46,48,62–64,66,68,71–77,85–88,91–93,98,99 · Internet 20,26,27,33,39,41,42,45–47,81–88,92,93,99 · SBA 31.
- Not valid: 00, 07, 08, 09, 17, 18, 19, 28, 29, 49, 69, 70, 78, 79, 89, 96, 97. No check digit.
## 2. Finance
### ABA routing number (RTN)
- Wikipedia https://en.wikipedia.org/wiki/ABA_routing_transit_number: 9 digits; check: `(3(d1+d4+d7) + 7(d2+d5+d8) + (d3+d6+d9)) mod 10 = 0`.
- First two digits: `00` US Government; `01–12` Federal Reserve districts (normal banks); `21–32` thrifts (historic, still valid); `61–72` non-bank payment processors/clearinghouses; `80` traveler's checks. Other prefixes unused (Wikipedia does not say so explicitly — inferred).
- Generator rule: prefix ∈ {01–12, 21–32, 61–72}, 6 random digits, compute 9th as `(10 − (3d1+7d2+d3+3d4+7d5+d6+3d7+7d8) mod 10) mod 10`.
- Open list of real RTNs: the Federal Reserve E-Payments Routing Directory (https://www.frbservices.org/resources/routing-number-directory/index.html) says "may not be sold, re-licensed, or otherwise used for commercial gain" → **not embeddable**. Generate synthetically.
### Payment card numbers (Wikipedia "Payment card number", https://en.wikipedia.org/wiki/Payment_card_number; all Luhn-checked)
| Network | IIN ranges | Length | CVV |
|---|---|---|---|
| Visa | 4 | 13, 16, 19 | 3 |
| Mastercard | 51–55, 2221–2720 | 16 | 3 |
| American Express | 34, 37 | 15 | 4 (front) |
| Discover | 6011, 644–649, 65, 622126–622925 (UnionPay co-brand) | 16–19 | 3 |
| JCB | 3528–3589 | 16–19 | 3 |
| Diners Club International | 30 (covers 300–305, 3095), 36, 38, 39 | 14–19 | 3 |
| Diners Club US & Canada | 55 (Mastercard co-brand) | 16 | 3 |
| UnionPay | 62 | 16–19 | 3 |
| Maestro | 5018, 5020, 5038, 5893, 6304, 6759, 6761, 6762, 6763 | 12–19 | 3 |
| Maestro UK | 6759, 676770, 676774 | 12–19 | 3 |
| Dankort 5019 · Mir 2200–2204 · RuPay 60,65,81,82,508 · Troy 65,9792 · UATP 1 (15) · Verve 506099–506198, 507865–507964, 650002–650027 (16/18/19) | | | |
- CVV lengths: https://en.wikipedia.org/wiki/Card_security_code ("three-digit ... Visa, Mastercard, and Discover"; "American Express is a four-digit code on the front").
- The requested "Maestro 50/56–58/6xxx" is a looser legacy rule; use Wikipedia's explicit IINs.
- Wikipedia text is CC BY-SA 4.0 — copy the *facts* (not prose) into your table; facts are not copyrightable.
### Published test cards
- Stripe: https://docs.stripe.com/testing — Visa 4242424242424242, Mastercard 5555555555554444, Amex 378282246310005, Discover 6011111111111117, Diners 3056930009020004 (14), JCB 3566002020360505, UnionPay 6200000000000005; any future expiry, any CVC (4 for Amex).
- PayPal: https://developer.paypal.com/tools/sandbox/card-testing/ — Visa 4005519200000004, 4012000033330026, 4012000077777777, 4012888888881881, 4217651111111119, 4500600000000061, 4772129056533503, 4915805038587737; Mastercard 2223000048400011; Amex 371449635398431, 376680816376961; Diners 36461510000039, 36461510000013; Maestro 6304000000000000, 5063516945005047; JCB 3636500000000260, 3636500000000989; CUP 6200680000000004, 6200680000000038.
- Reuse: neither page carries an open licence (Stripe docs are under Stripe's site terms). The numbers themselves are Luhn-valid facts, widely reproduced; embedding the numbers with a "from Stripe/PayPal test docs" note is low-risk but **UNVERIFIED** as an explicit grant. Alternative: generate Luhn-valid numbers from the IIN table above — no dependency on either vendor.
### BIC / SWIFT (ISO 9362)
- https://en.wikipedia.org/wiki/ISO_9362: 4 letters bank code + 2 letters ISO 3166-1 country + 2 alphanumeric location + optional 3 alphanumeric branch (8 or 11 chars; `XXX` = primary office). Location second char `0` = test BIC, `1` = passive participant, `2` = reverse billing.
- Generator rule: `[A-Z]{4}[A-Z]{2}[A-Z2-9][A-NP-Z0-9]([A-Z0-9]{3})?` — avoid `0`/`1` as the 8th char for a "live" BIC.
- SWIFT's BIC Directory is proprietary (SwiftRef licence, redistribution needs a separate "SwiftRef Redistribution License": https://www.swift.com/myswift/ordering/order-products-services/swiftref-redistribution-license) → **not embeddable**.
- Open alternative: GLEIF BIC-to-LEI mapping, https://www.gleif.org/en/lei-data/lei-mapping/download-bic-to-lei-relationship-files, files at https://mapping.gleif.org/api/v2/bic-lei/ (monthly zip; `LEI-BIC-20260828.zip` → `lei-bic-20260828T000000.csv`, columns `LEI,BIC`, 39,347 rows). Licence: "BIC/LEI Mapping Table License Agreement" (https://www.gleif.org/lei-data/lei-mapping/download-bic-to-lei-relationship-files/2017-12-21_annex-2_bic-to-lei-mapping-table-license-agreement_final.pdf) — a CC0-style grant "for any purpose whatsoever, including ... commercial" but **conditional on reproducing the notice** "SWIFT © and database rights [month year]. All rights reserved. This Mapping Table has been developed by SWIFT. Any use of the Mapping Table ... is subject to the BIC/LEI Mapping Table License Agreement ... For the latest BIC information and updates, always refer to www.swift.com/bic." Embeddable in an MIT repo with that notice in the data file; it is a list of real live BICs (bank names not included — join to GLEIF LEI data if names are wanted).
### ISIN (https://en.wikipedia.org/wiki/International_Securities_Identification_Number)
- 12 chars: 2-letter country + 9 alphanumeric NSIN + 1 check digit.
- Check: convert letters A=10…Z=35 (ASCII − 55) to produce a digit string, then Luhn over that string (double every second digit from the right, sum digits, check makes total ≡ 0 mod 10). Pitfall: work on the *expanded* digit string; a transposed letter pair can pass.
### CUSIP (https://en.wikipedia.org/wiki/CUSIP)
- 9 chars: 6 issuer + 2 issue + 1 check. Values: digits as-is, A=10…Z=35, `*`=36, `@`=37, `#`=38.
- Check: for i in 1..8, v = value; if i even, v ×= 2; sum += v div 10 + v mod 10; check = (10 − sum mod 10) mod 10.
- Real CUSIPs are proprietary (CGS/FactSet); generate synthetic ones only.
### IBAN (ISO 13616)
- Validation (https://en.wikipedia.org/wiki/International_Bank_Account_Number): move first 4 chars to end, letters → 10–35, integer mod 97 must equal 1; max 34 chars; check digits 02–98.
- SWIFT registry: https://www.swift.com/standards/data-standards/iban-international-bank-account-number, PDF release 101 https://www.swift.com/sites/default/files/files/iban-registry-v101.pdf (release 100 was Oct 2025); TXT also published (URL **UNVERIFIED** — swift.com returns 403/errors to automation this session). Terms: free of charge; no explicit licence text found (**UNVERIFIED**). Country count: 89 countries per Wikipedia (Dec 2024).
- Embeddable per-country length tables already extracted from the registry:
- php-iban `registry.txt` (https://github.com/globalcitizen/php-iban, **LGPL-3.0**): pipe-separated, 121 rows (116 official + unofficial), columns `country_code|country_name|domestic_example|bban_example|bban_format_swift|bban_format_regex|bban_length|iban_example|iban_format_swift|iban_format_regex|iban_length|bban_bankid_start_offset|...|country_sepa|swift_official|...|currency_iso4217|central_bank_url|central_bank_name|membership`. LGPL data in an MIT repo is awkward; use it as a *reference* to build your own table (formats are facts) rather than copying the file.
- ibankit-js (https://github.com/koblas/ibankit-js, **Apache-2.0**, registry v95) — same caveat.
- Best route: own table `{country, length, bban_format}` derived from the SWIFT PDF (facts), cite SWIFT.
### Currencies (ISO 4217)
- ISO says use is free: https://www.iso.org/iso-4217-currency-codes.html — "ISO allows free-of-charge use of its country, currency and language codes from ISO 3166, ISO 4217 and ISO 639". Official list from SIX: https://www.six-group.com/dam/download/financial-information/data-center/iso-currrency/lists/list-one.xml (+ `.xls`, `list-three` historic). Fields `CtryNm, CcyNm, Ccy, CcyNbr, CcyMnrUnts`; published 2026-01-01; 280 entity rows, 178 distinct codes. No symbols. No licence text inside the XML.
- Alternatives:
- datasets/currency-codes (https://github.com/datasets/currency-codes): **PDDL** (public domain); `data/codes-all.csv` fields `Entity, Currency, AlphabeticCode, NumericCode, MinorUnit, WithdrawalDate`; built from SIX list one + three; no symbols. Embeddable.
- npm `currency-codes` v2.2.0 (https://github.com/freeall/currency-codes): **MIT**; fields `code, number, digits, currency, countries[]`; generated from the SIX XML; no symbols. Embeddable.
- umpirsky/currency-list (https://github.com/umpirsky/currency-list): **MIT**; code → localised name only (311 entries in `data/en_US/currency.json`, includes historic); no symbols, no numeric codes, no decimals.
- Debian iso-codes (https://salsa.debian.org/iso-codes-team/iso-codes): **LGPL-2.1+**; `data/iso_4217.json` has `alpha_3, name, numeric` only (179 entries); no symbols/minor units. LGPL → avoid embedding.
- Symbols: none of the above carry them. Unicode CLDR (`common/main/en.xml` currency symbols, Unicode License, MIT-compatible with notice) is the usual source — **UNVERIFIED this session**.
### Crypto addresses
- Bitcoin (https://en.bitcoin.it/wiki/List_of_address_prefixes): P2PKH version 0x00 → leading `1`; P2SH 0x05 → leading `3`; Base58Check 25–34 chars (alphabet excludes `0OIl`; last 4 bytes = first 4 of double-SHA256 of version+payload). Testnet `m`/`n`, `2`, `tb1`.
- Bech32 (BIP-173, https://github.com/bitcoin/bips/blob/master/bip-0173.mediawiki, BSD-2-Clause): HRP `bc`, separator `1`, charset `qpzry9x8gf2tvdw0s3jn54khce6mua7l`, all-lowercase (or all-uppercase), max 90 chars; P2WPKH `bc1q…` = 42 chars, P2WSH `bc1q…` = 62 chars.
- Bech32m (BIP-350, https://github.com/bitcoin/bips/blob/master/bip-0350.mediawiki): witness v1+ (P2TR `bc1p…`, 62 chars), checksum constant `0x2bc830a3` instead of 1.
- Ethereum EIP-55 (https://eips.ethereum.org/EIPS/eip-55, CC0): `0x` + 40 hex; keccak256 of the lowercase hex (no 0x, as ASCII); for each hex letter, uppercase iff the corresponding hash nibble ≥ 8 (bit 4·i set). Digits unchanged.
## 3. Company
### NAICS 2022 (US Census)
- Page https://www.census.gov/naics/ ; files: 6-digit list https://www.census.gov/naics/2022NAICS/6-digit_2022_Codes.xlsx (82 kB; **1,012** six-digit codes, 111110…928120, verified by parsing), 2–6 digit https://www.census.gov/naics/2022NAICS/2-6%20digit_2022_Codes.xlsx, structure https://www.census.gov/naics/2022NAICS/2022_NAICS_Structure.xlsx, descriptions https://www.census.gov/naics/2022NAICS/2022_NAICS_Descriptions.xlsx, manual PDF https://www.census.gov/naics/reference_files_tools/2022_NAICS_Manual.pdf.
- Format: xlsx, column A code (numeric), column B title (trailing spaces in some titles — trim).
- Licence: no statement on the page or in the manual front matter; US federal government work → public domain in the US (17 U.S.C. §105). Embeddable; cite Census/OMB.
### SEC EDGAR company names
- `https://www.sec.gov/files/company_tickers.json`: object keyed `"0".."N"` of `{cik_str, ticker, title}`; **10,422** entries (2026-09-17). `company_tickers_exchange.json`: `{"fields":["cik","name","ticker","exchange"],"data":[...]}`.
- `https://www.sec.gov/Archives/edgar/cik-lookup-data.txt`: 40 MB, **1,059,372** lines, `NAME:CIK:` (`COMPANY NAME:0001234567:`), all filers incl. individuals — filter to the ticker list for company names.
- Access: declare `User-Agent: Company Name email@domain`, ≤ 10 req/s (https://www.sec.gov/os/webmaster-faq). Licence: "All Government-created content on sec.gov and EDGAR public filing content are free to access and reuse" (same FAQ). Embeddable; the ticker file (~800 kB) is the practical one.
### DUNS
- https://en.wikipedia.org/wiki/Data_Universal_Numbering_System: 9 digits, no significance; shown `NN-NNN-NNNN` or plain. Had a mod-10 check digit until ~Dec 2006; **dropped** since (expanded the pool by 800 M). So: any 9 digits are structurally valid today; no validation rule to enforce. Optional DUNS+4 suffix (4 alphanumerics) is user-assigned.
### Legal entity suffixes per country
- GLEIF ISO 20275 Entity Legal Forms code list v1.6 (2026-02-19): page https://www.gleif.org/en/lei-data/code-lists/iso-20275-entity-legal-forms-code-list ; CSV https://www.gleif.org/lei-data/code-lists/iso-20275-entity-legal-forms-code-list/2026-02-19-elf-code-list-v1.6.csv (758 kB; xlsx alongside).
- Columns: `ELF Code, Country of formation, Country Code (ISO 3166-1), Jurisdiction of formation, Country sub-division code (ISO 3166-2), Entity Legal Form name Local name, Language, Language Code (ISO 639-1), Entity Legal Form name Transliterated name, Abbreviations Local language, Abbreviations transliterated, Date created, ELF Status ACTV/INAC, Modification, Modification date, Reason`.
- Counts: 4,002 rows; 130 countries; 3,790 ACTV / 212 INAC; 1,518 rows have a local abbreviation (e.g. US `N.A.`, `FSA`; SE `AB`… multiple abbreviations separated by `;`). US has 737 rows (per-state forms). Abbreviations are the field you want for `Inc`, `LLC`, `GmbH`, `AB`, `SAS`, `Pty Ltd`; many rows have none, so filter on non-empty abbreviation and status ACTV.
- Licence: GLEIF Open Data page (https://www.gleif.org/en/about/open-data): "The data on GLEIF's website is provided under a Creative Commons (CC0) license"; LEI Data Terms of Use: "provided under the CC0 licence, see CC0 1.0 Universal". The ELF page itself names no licence — that CC0 explicitly covers this code list is **UNVERIFIED**, but GLEIF treats all its published data this way. Embeddable.
## Quick embed verdict
| Item | Embed? | Licence / attribution |
|---|---|---|
| SSA baby names | yes | CC0 (data.gov) |
| Census 2010 surnames | yes | US gov PD; cite Census |
| EIN prefixes, ITIN/SSN rules | yes (rules) | facts from IRS/SSA |
| NAICS 2022 | yes | US gov PD |
| SEC company_tickers.json | yes | free to reuse (SEC) |
| GLEIF ELF code list | yes | CC0 (GLEIF) |
| GLEIF BIC-LEI (real BICs) | yes, with mandatory SWIFT notice | custom permissive licence |
| ISO 4217 (SIX XML / datasets PDDL / npm MIT) | yes | free use (ISO), PDDL, MIT |
| Fed routing directory, SwiftRef BIC directory, CUSIP lists | no | restricted/proprietary — generate synthetically |
| php-iban / ibankit registries | copy facts only | LGPL-3.0 / Apache-2.0 |
| Stripe / PayPal test cards | numbers only, note source | no explicit grant (UNVERIFIED) |
# Part 2 — Internet, vehicle, airline
Verdict key: **EMBED** = safe to vendor into an MIT repo; **EMBED+ATTR** = ok with a notice; **AVOID** = licence incompatible or unclear.
## 4. Internet
### IANA registries (all of them)
- Licence: IANA/IETF statement at https://www.iana.org/help/licensing-terms — "IANA and IETF intend that the Protocol Registries may be freely used by any party for any purpose", data dedicated under **CC0 1.0**. No attribution required. **EMBED**.
- All registries below share this licence; ship one `SOURCES`/NOTICE line pointing at that URL.
| Registry | CSV URL | Rows (2026-09-16) | Notes |
|---|---|---|---|
| Media types | `https://www.iana.org/assignments/media-types/<type>.csv`, type ∈ application, audio, font, haptics, image, message, model, multipart, text, video (also `example`) | application 1800, audio 165, font 6, haptics 3, image 88, message 27, model 42, multipart 17, text 105, video 97 (header excluded) ≈ **2350** | Fields: `Name,Template,Reference`. Template is the full `type/subtype`; some rows are DEPRECATED/OBSOLETED in Name (filter). Format: RFC 6838 §4.2 — `type "/" subtype`, restricted chars, max 127 chars each, case-insensitive. |
| HTTP status codes | https://www.iana.org/assignments/http-status-codes/http-status-codes-1.csv | 75 rows, **64 assigned** (306 and 418 are "(Unused)"; 104 is TEMPORARY) → 62 usable | Fields: `Value,Description,Reference`. Value is 3 digits 1xx–5xx. |
| IPv4 special-purpose | https://www.iana.org/assignments/iana-ipv4-special-registry/iana-ipv4-special-registry-1.csv | 26 blocks | Fields: `Address Block,Name,RFC,Allocation Date,Termination Date,Source,Destination,Forwardable,Globally Reachable,Reserved-by-Protocol`. |
| IPv6 special-purpose | https://www.iana.org/assignments/iana-ipv6-special-registry/iana-ipv6-special-registry-1.csv | 27 blocks | Same fields. Contains `2001:db8::/32` and `3fff::/20`. |
| Service names / ports | https://www.iana.org/assignments/service-names-port-numbers/service-names-port-numbers.csv | 14,534 rows; 11,731 with both name and a single port; **6,102 distinct named ports**; 6,008 tcp-named; 707 tcp names < 1024 | Fields: `Service Name,Port Number,Transport Protocol,Description,Assignee,Contact,Registration Date,Modification Date,Reference,Service Code,Unauthorized Use Reported,Assignment Notes`. Port Number may be empty, single, or a range `a-b`. Well-known = 0–1023, registered 1024–49151, dynamic 49152–65535 (RFC 6335). ~1.1 MB; embed only the named-port subset. |
| TLDs | https://data.iana.org/TLD/tlds-alpha-by-domain.txt | **1,401** TLDs (first line is a `# Version …` comment) | Upper-case ASCII; IDNs as `XN--…` punycode. Same IANA CC0 terms (UNVERIFIED that data.iana.org is explicitly covered by the licensing page; it is the same registry data). Lower-case on use. |
### MIME extension mapping: jshttp/mime-db
- https://github.com/jshttp/mime-db — **MIT** (GitHub API spdx MIT; npm `license: MIT`), version 1.54.0 on npm.
- `db.json` (CDN: https://cdn.jsdelivr.net/npm/mime-db/db.json): object keyed by lower-case type → `{source, extensions[], compressible, charset}`.
- Counts (1.54.0): **2,522 types**, **1,015 with extensions**; source = iana 2,136, apache 275, nginx 13, custom 98.
- Sources: IANA registry (CC0), Apache httpd `mime.types` (Apache-2.0), nginx `mime.types` (BSD-2). **EMBED+ATTR** (keep the MIT copyright notice). Data updates are not semver-breaking per README.
### IPv4 / IPv6 generation rules (RFCs, no data file needed)
- RFC 5737 documentation: `192.0.2.0/24` (TEST-NET-1), `198.51.100.0/24` (TEST-NET-2), `203.0.113.0/24` (TEST-NET-3). https://www.rfc-editor.org/rfc/rfc5737.html
- RFC 1918 private: `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16`. Also useful: `100.64.0.0/10` CGNAT (RFC 6598), `198.18.0.0/15` benchmarking (RFC 2544), `169.254.0.0/16` link-local.
- IPv6 documentation: `2001:db8::/32` (RFC 3849) and `3fff::/20` (RFC 9637, "expands on the existing 2001:db8::/32 … with the reservation of an additional, larger prefix"). ULA `fc00::/7` (use `fd00::/8`), link-local `fe80::/10`.
- Validation: generate inside those blocks; avoid network/broadcast (`.0`/`.255`) if the consumer expects host addresses. Format IPv6 per RFC 5952 (lower-case hex, `::` once, longest zero run).
### MAC OUI (IEEE MA-L)
- CSV: https://standards-oui.ieee.org/oui/oui.csv (3.8 MB, **40,161 rows**, fields `Registry,Assignment,Organization Name,Organization Address`; Assignment is 6 hex chars). Server returns HTTP 418 to non-browser user agents; fetch with a browser UA. Also MA-M `oui28.csv`, MA-S `oui36.csv`, CID `cid.csv`.
- Licence: no licence on the IEEE pages themselves (site footer "© IEEE – All rights reserved" is generic). Debian's `ieee-data` copyright file (https://metadata.ftp-master.debian.org/changelogs/main/i/ieee-data/unstable_copyright) records IEEE's 2014 statement: *"IEEE does not assert any copyright in the OUI Public Listing or attempt to restrict distribution of the listing in any way. The IEEE Registration Authority does, however, strongly encourage those who use the list to obtain it directly from IEEE…"* — same text in the FSF directory. Debian ships it in `main` on that basis. **EMBED+ATTR** (cite the statement); the org-address column is not needed — keep `Assignment,Organization Name` only (~1 MB → few hundred KB). UNVERIFIED: no current IEEE page restates it (regauth FAQ and IPR pages do not mention it).
- Alternative needing no data: locally administered unicast MAC — first octet's bit 1 (U/L) = 1, bit 0 (I/G) = 0 → second hex digit ∈ {2, 6, A, E}, e.g. `x2:xx:xx:xx:xx:xx`. Guaranteed never to collide with a vendor OUI. Notations: `01:23:45:67:89:ab`, `01-23-45-67-89-AB`, Cisco `0123.4567.89ab`.
### User-agent strings
| Source | Licence | Verdict |
|---|---|---|
| https://www.useragents.me/ | No licence stated anywhere on the page ("© 2022–2025 Useragents.me"); `/api` returns 404 now; JSON blobs embedded in page, weekly updates | **AVOID** (unlicensed = all rights reserved) |
| https://explore.whatismybrowser.com/useragents/explore/ + API legal https://developers.whatismybrowser.com/api/about/legal/ | Proprietary: "You may NOT share the unique database download URL, or the downloaded file itself … EXCEPT for making it publicly available on the internet" | **AVOID** |
| npm `user-agents` (intoli) https://github.com/intoli/user-agents | **BSD-2-Clause**, v2.1.185, daily-rebuilt dataset from Intoli traffic, weighted by frequency; fields `userAgent, platform, vendor, appName, deviceCategory, screenWidth/Height, viewportWidth/Height, connection{…}, oscpu, cpuClass, pluginsLength` | **EMBED+ATTR** (keep BSD notice). Count not stated on the page (UNVERIFIED; dataset is a few thousand rows). |
| npm `top-user-agents` (microlinkhq) https://github.com/microlinkhq/top-user-agents | **MIT** (GitHub API). Top 100 UA strings from microlink.io traffic, weekly; JSON: https://cdn.jsdelivr.net/gh/microlinkhq/top-user-agents@master/src/index.json (+ desktop.json, mobile.json) | **EMBED+ATTR** — smallest clean option |
- None of the above is CC0. Alternative: generate UAs from templates (Chrome/Firefox/Safari/Edge strings are fixed-form; only version numbers vary), no data needed.
### Programming languages: GitHub Linguist
- https://raw.githubusercontent.com/github-linguist/linguist/main/lib/linguist/languages.yml — **MIT** (repo LICENSE, "Copyright (c) 2017 GitHub, Inc."). **EMBED+ATTR**.
- **835** top-level entries; **562** with `type: programming` (others: data, markup, prose). Per-entry fields: `type, color, extensions, aliases, tm_scope, ace_mode, language_id, group, interpreters, …`.
### Operating systems
- No open dataset found beyond awesome-lists. OS names/versions are facts — hand-author a small table (Windows 10/11, Windows Server, macOS versions+names, Ubuntu/Debian/Fedora/Arch/Alpine, iOS, Android, FreeBSD…). `user-agents` (BSD-2) carries `platform`/`oscpu` values if a data-driven list is wanted.
### Database engines / collations
- Engine names: hand-author (facts).
- MySQL manual (https://dev.mysql.com/doc/refman/8.4/en/preface.html): Oracle terms — "you may not … copy, reproduce … distribute … any part" except distributing the docs with the software. **AVOID copying tables from MySQL docs.** Instead dump `SHOW COLLATION` from a `mysql:8.4.x` container (server output is factual, not the manual) — ~286 collations / 41 charsets in 8.x (UNVERIFIED count). PostgreSQL docs are under the PostgreSQL licence (permissive, keep notice) — `pg_collation` dump or docs both fine. MariaDB KB is CC BY-SA — avoid.
## 5. Vehicle
### NHTSA vPIC
- API base `https://vpic.nhtsa.dot.gov/api/vehicles/`, `?format=json|xml|csv`. Endpoints verified: `GetAllMakes` (**12,363 makes**, fields `Make_ID,Make_Name`; 612 KB JSON, mostly small trailer/custom builders — filter with `GetMakesForVehicleType/car` → 195 makes, `/truck` → 207), `GetModelsForMake/honda` (362 models, `Make_ID,Make_Name,Model_ID,Model_Name`), `DecodeWMI/1HG`, `GetWMIsForManufacturer/honda` (45 WMIs, fields `WMI,Name,Country,VehicleType,Id,…`), `GetVehicleVariableList`, `GetVehicleVariableValuesList/<display name>` (must be the display name, URL-encoded, e.g. `Fuel%20Type%20-%20Primary`; the camel-case name returns 0 rows).
- Value lists (counts): **Fuel Type - Primary 14**, **Body Class 71**, **Transmission Style 12**, **Drive Type 23**, **Vehicle Type 9**. Fields `Id,Name`.
- Bulk: https://vpic.nhtsa.dot.gov/downloads/ — `vPICList_lite_2026_08.bak.zip` (MS SQL Server 2019+ backup, 190 MB, updated 2026-08-14), also `.plain.zip` / `.custom.zip`. Page says standalone DB "limited to VIN decoding"; makes/models/variables still via API. Contains `Wmi`, `Make`, `Model`, `WMIYearValidChars` tables (UNVERIFIED table names).
- Licence: FAQ https://vpic.nhtsa.dot.gov/api/home/index/faq — "No, NHTSA is a government agency and the services provided on the API are free for use by the public as an offering as a part of our Open Data initiatives"; no rate-limit number stated, batch jobs asked to run at night. Work of the US federal government → public domain under 17 USC §105; data.gov entry lists licence as "unknown". **EMBED** (note source; no attribution legally required).
### VIN (ISO 3779 / 49 CFR 565, North America)
- 17 chars from `0-9 A-H J-N P R-Z` (**I, O, Q never**). Positions: 1–3 WMI, 4–8 VDS, 9 check digit, 10 model year, 11 plant, 12–17 serial (last 4 must be digits for US, positions 12–14 alphanumeric — 49 CFR 565.15).
- Check digit: transliterate `A=1 B=2 C=3 D=4 E=5 F=6 G=7 H=8 J=1 K=2 L=3 M=4 N=5 P=7 R=9 S=2 T=3 U=4 V=5 W=6 X=7 Y=8 Z=9`, digits as-is; weights `8 7 6 5 4 3 2 10 0 9 8 7 6 5 4 3 2`; sum mod 11 → digit, remainder 10 → `X`.
- Model year (pos 10): excludes `I O Q U Z 0`; `A=2010 … H=2017 J=2018 K=2019 L=2020 M=2021 N=2022 P=2023 R=2024 S=2025 T=2026 V=2027 W=2028 X=2029 Y=2030`, `1=2031…9=2039` (30-year cycle; `1–9` also = 2001–2009). Source: https://en.wikipedia.org/wiki/Vehicle_identification_number (rules are facts; do not copy prose).
### WMI list
- SAE J1044 / SAE WMI database is paid (https://www.sae.org/standards/j1044_202501-world-manufacturer-identifier) — **AVOID**.
- Wikipedia "List of WMIs" — CC BY-SA — **AVOID**.
- Use vPIC: `GetWMIsForManufacturer/<name or id>` per manufacturer, or the `Wmi` table in the standalone `.bak`. Public domain. Country prefix rule (first char): `1,4,5` USA, `2` Canada, `3` Mexico, `J` Japan, `K` Korea, `S` UK, `W` Germany, `Y` Sweden/Finland, `Z` Italy, `L` China (ISO 3780 facts).
### US licence plate formats
| Source | Licence | Verdict |
|---|---|---|
| openalpr `runtime_data/postprocess/us.patterns` (https://github.com/openalpr/openalpr) | **AGPL-3.0** | **AVOID** (291 lines, `@`=letter `#`=digit — useful only as a reference to check hand-authored patterns) |
| Wikipedia "United States license plate designs and serial formats" | CC BY-SA | **AVOID copying**; formats themselves are facts |
| jonnii/platekit | MIT, but SVG rendering, no serial patterns | not useful |
- No MIT/CC0 per-state pattern list found. Recommendation: hand-author a per-state table (facts); current standard passenger formats as listed by Wikipedia (verify against DMV pages before use): AL `0AXXXXX`/`00AXXXX`, AK `ABC 123`, AZ `XXX 1XX`, AR `ABC 12D`, CA `1ABC123` (digit-3 letters-3 digits), CO `ABC-D12`, CT `AB·12345`, DE `123456`, DC `AB-1234`, FL `ABC D12`, GA `ABC1234`, HI `ABC 123`, ID `A 1234U`, IL `AB 12345`, IN `123ABC`, IA `ABC 123`, KS `1234ABC`, KY `ABC123`, LA `123 ABC`, ME `123·ABC`, MD `1AB2345`, MA `1ABC 23`, MI `ABC 1234`, MN `ABC-123`, MS `ABC 123`, MO `AB1 C2D`, MT `0-AB1234`, NE `ABC 123`, NV `123·A45`, NH `123 4567`, NJ `D12-ABC`, NM `123-ABC`, NY `ABC-1234`, NC `ABC-1234`, ND `123 ABC`, OH `ABC 1234`, OK `ABC-123`, OR `123 ABC`, PA `ABC1234`, RI `1AB 234`, SC `123ABC`, SD `0A1 234`, TN `ABC 1234`, TX `ABC-1234`, UT `A12 3BC`, VT `ABC 123`, VA `ABC-1234`, WA `ABC1234`, WV `X1A 2345`, WI `ABC-1234`, WY `1A-123A`. Most states omit I/O/Q from serials (state-specific, UNVERIFIED per state).
## 6. Airline
### OurAirports
- https://ourairports.com/data/ — "All data is released to the Public Domain"; mirror repo https://github.com/davidmegginson/ourairports-data is **Unlicense** (GitHub API). Nightly updates. **EMBED**.
- Raw URLs: `https://davidmegginson.github.io/ourairports-data/{airports,airport-frequencies,runways,navaids,countries,regions}.csv`.
| File | Rows | Fields |
|---|---|---|
| airports.csv (12.7 MB) | **86,089**; by type: small_airport 42,734, heliport 23,216, closed 13,524, medium 4,106, seaplane_base 1,273, large 1,174, balloonport 62 | `id,ident,type,name,latitude_deg,longitude_deg,elevation_ft,continent,iso_country,iso_region,municipality,scheduled_service,icao_code,iata_code,gps_code,local_code,home_link,wikipedia_link,keywords` |
| — with `iata_code` | **9,055** (large 1,171, medium 3,398, small 4,233, seaplane 153, heliport 100); `scheduled_service=yes` 4,335; US with IATA 2,036 (873 large/medium) | Suggested subset: `iata_code != '' AND type IN (large_airport, medium_airport)` → 4,569 rows |
| countries.csv | **249** | `id,code,name,continent,wikipedia_link,keywords` |
| regions.csv | **3,987** | `id,code,local_code,name,continent,iso_country,wikipedia_link,keywords` (code = ISO 3166-2, e.g. `US-KS`) |
- No airline file. Derived: datasets/airport-codes on datahub (PDDL) — no need, use upstream.
### OpenFlights
- https://openflights.org/data.php — airports/airlines/routes/planes under **ODbL 1.0** (+ DbCL); airline/plane data partly from Wikipedia (GFDL/CC BY-SA). 5,888 airlines (2012). **AVOID** (share-alike). npm `airline-codes` (npow, ISC) is a straight OpenFlights sync → inherits ODbL → **AVOID**.
### Airline IATA/ICAO codes — open options
| Source | Licence | Count | Verdict |
|---|---|---|---|
| Wikidata SPARQL (https://query.wikidata.org/sparql) | **CC0** (https://www.wikidata.org/wiki/Wikidata:Licensing) | **2,928** items with IATA code (P229), **3,644** with ICAO code (P230) | **EMBED**. Query: `SELECT ?item ?itemLabel ?iata ?icao ?callsign WHERE { ?item wdt:P229 ?iata . OPTIONAL { ?item wdt:P230 ?icao } OPTIONAL { ?item wdt:P432 ?callsign } SERVICE wikibase:label { bd:serviceParam wikibase:language "en" } }`. Includes defunct airlines — filter `MINUS { ?item wdt:P576 ?dissolved }` and/or P31 = Q46970 (airline). |
| OpenTravelData `optd_airlines.csv` (https://github.com/opentraveldata/opentraveldata) | **CC BY 4.0** | 1,620 rows; `^`-separated; fields `pk^env_id^validity_from^validity_to^3char_code^2char_code^num_code^name^name2^alliance_code^alliance_status^type^wiki_link^flt_freq^alt_names^bases^key^version^parent_pk_list^successor_pk_list` | **EMBED+ATTR** (attribution notice required) |
- Format rules: IATA designator = 2 alphanumeric chars (`[A-Z0-9]{2}`, at least one letter in practice; a third optional char exists in the standard but is unused); "controlled duplicates" share a code between non-overlapping regional carriers. ICAO designator = 3 letters `[A-Z]{3}`, unique. Accounting/prefix code = 3 digits (ticket number prefix, e.g. `016` United).
### Aircraft types
| Source | Licence | Verdict |
|---|---|---|
| ICAO Doc 8643 (https://www.icao.int/operational-safety/doc-8643-aircraft-type-designators) | ICAO copyright notice: "None of the materials … may be used, reproduced or transmitted … without permission in writing from ICAO"; API Data Service is paid (25 free trial calls) | **AVOID** — including GitHub mirrors such as ColtJD45/icao-aircraft-designator-list (MIT-labelled, 7,388 rows, fields `manufacturer,model,type_designator,description,engine_type,engine_count,wtc`) because the underlying data is scraped ICAO content (UNVERIFIED provenance) |
| OpenSky aircraft database (https://opensky-network.org/data/aircraft; CSV https://s3.opensky-network.org/data-samples/metadata/aircraftDatabase.csv) | "The aircraft database is unlicensed and does not fall under our terms of use … offered as is"; built from registries, openflights.org and ICAO Doc 8643; citation requested; page says it is no longer up to date | **AVOID** ("unlicensed" ≠ open; upstream includes ODbL/ICAO). Fields for reference: `icao24,registration,manufacturericao,manufacturername,model,typecode,serialnumber,linenumber,icaoaircrafttype,operator,operatorcallsign,operatoricao,operatoriata,owner,…,engines,…,categoryDescription` |
| Wikidata | **CC0** | **550** items with ICAO aircraft type designator (P8305); also IATA aircraft code P9040 (UNVERIFIED property id), manufacturer P176 | **EMBED**. Query: `SELECT ?item ?itemLabel ?icao ?mfrLabel WHERE { ?item wdt:P8305 ?icao . OPTIONAL { ?item wdt:P176 ?mfr } SERVICE wikibase:label { bd:serviceParam wikibase:language "en" } }` |
- Format: ICAO type designator 2–4 alphanumerics (`[A-Z0-9]{2,4}`, e.g. `B738`, `A20N`, `E190`); IATA aircraft code 3 alphanumerics (`738`, `32N`).
### Airport / flight / seat format rules
- IATA airport code `[A-Z]{3}`; ICAO location indicator `[A-Z0-9]{4}` (US `K…`, Canada `C…`, UK `EG…`, Sweden `ES…`). Use OurAirports rows so codes are real.
- Flight number: 2-char IATA designator + 1–4 digits, no leading zeros in display (`QF9`, `AA1234`); systems cap at 0001–9999. Conventions (not rules): 1–999 mainline, 3000–5999 regional affiliates, ≥6000 codeshares, 8xxx charters, ≥9000 ferry/positioning. ICAO form: 3-letter designator + 1–4 alphanumerics (`AFR1`). Source: https://en.wikipedia.org/wiki/Flight_number.
- Seat: `row + letter`. Rows typically 1–~60 (A380 up to ~90); many airlines skip 13 (and 14/17 on some carriers). Letters `A–K` **skipping I** (looks like 1/l); narrow-body 3-3 = `ABC DEF` (A/F windows, C/D aisles); wide-body 3-4-3 = `ABC DEFG HJK`, 2-4-2 = `AB DEFG JK`. Regex: `^[1-9][0-9]?[A-HJK]$`.
# Part 3 — Science/nature, Text, Dates: open data sources for an MIT fake-data library
Researched 2026-09-17. "OK to embed" = redistributable inside an MIT repo with the stated notice. Anything not fetched first-hand is marked **UNVERIFIED**.
## 7. Science / nature
### Periodic table
| Source | Licence | Embed in MIT repo | Fields | Rows | Format |
|---|---|---|---|---|---|
| PubChem periodic table — CSV https://pubchem.ncbi.nlm.nih.gov/rest/pug/periodictable/CSV, JSON https://pubchem.ncbi.nlm.nih.gov/rest/pug/periodictable/JSON (page: https://pubchem.ncbi.nlm.nih.gov/periodic-table/) | US-government work, public domain. NLM policy (https://www.ncbi.nlm.nih.gov/home/about/policies/): "Information that is created by or for the US government on this site is within the public domain … may be freely distributed and copied. However, it is requested that in any subsequent use of this work, NLM be given appropriate acknowledgment." | Yes; add an acknowledgment line ("Element data: PubChem/NLM"). | `AtomicNumber, Symbol, Name, AtomicMass, CPKHexColor, ElectronConfiguration, Electronegativity, AtomicRadius, IonizationEnergy, ElectronAffinity, OxidationStates, StandardState, MeltingPoint, BoilingPoint, Density, GroupBlock, YearDiscovered` (17) | 118 (Z=1..118, verified) | CSV; JSON is `{Table:{Columns:{Column:[...]},Row:[{Cell:[...]}]}}` |
| NIST periodic table https://www.nist.gov/pml/periodic-table-elements | Public domain (17 USC 105); NIST asks for citation (https://www.nist.gov/open/copyright-fair-use-and-licensing-statements-srd-data-software-and-technical-series-publications) | Yes, but | PDF only (no CSV/JSON) — not usable as a data source | 118 | PDF |
| Bowserinator/Periodic-Table-JSON https://github.com/Bowserinator/Periodic-Table-JSON | **CC BY-SA 3.0** (LICENSE.md, verified; data derived from Wikipedia) | **No** — ShareAlike is incompatible with MIT redistribution | 33 keys incl. `name, symbol, number, atomic_mass, category, phase, period, group, block, shells, electron_configuration, summary, cpk-hex, …` | **119** (includes hypothetical Ununennium) | JSON, CSV |
| faker-js `src/locales/en/science/chemical_element.ts` https://github.com/faker-js/faker (branch `next`) | MIT (LICENSE verified) | Yes, keep MIT notice | `{symbol, name, atomicNumber}` | 118 | TS |
Validation rules: `AtomicNumber` 1..118 unique; `Symbol` 1–2 letters, first uppercase, unique; `Name` unique. PubChem leaves numeric fields empty for Z≥104 where unknown; `ElectronConfiguration` for Og is `"[Rn]7s2 7p6 5f14 6d10 (predicted)"` and `StandardState` is `"Expected to be a Gas"` — strip "(predicted)"/"Expected to be a" if you want an enum. `AtomicMass` for unstable elements is the mass number of the most stable isotope as a bare decimal (e.g. `295.216`), no brackets. `CPKHexColor` is 6 hex digits without `#`, empty for some.
### SI units
- Source: BIPM SI Brochure, 9th ed. (2019; current revision v4.01, 2026) https://www.bipm.org/en/publications/si-brochure ; PDF https://www.bipm.org/documents/20126/41483022/SI-Brochure-9-EN.pdf
- Licence: **CC BY 4.0** — stated on the brochure page and on BIPM's copyright page (https://www.bipm.org/en/copyright): content may be adapted, copied, used in commercial products with "appropriate acknowledgement of the BIPM and its source"; BIPM name/logo must not imply endorsement. Verified via the HTML pages; the PDF's own front-matter licence line is **UNVERIFIED** (could not text-extract the PDF here).
- Embed: Yes — the unit *names/symbols* are facts (not copyrightable); attribute "SI Brochure, BIPM, CC BY 4.0" anyway.
- Data (cross-checked against https://en.wikipedia.org/wiki/International_System_of_Units):
- 7 base units: second s (time), metre m (length), kilogram kg (mass), ampere A (electric current), kelvin K (thermodynamic temperature), mole mol (amount of substance), candela cd (luminous intensity).
- 22 coherent derived units with special names: radian rad, steradian sr, hertz Hz, newton N, pascal Pa, joule J, watt W, coulomb C, volt V, farad F, ohm Ω, siemens S, weber Wb, tesla T, henry H, degree Celsius °C, lumen lm, lux lx, becquerel Bq, gray Gy, sievert Sv, katal kat.
- 24 prefixes: quetta Q 10^30, ronna R 10^27, yotta Y 10^24, zetta Z 10^21, exa E 10^18, peta P 10^15, tera T 10^12, giga G 10^9, mega M 10^6, kilo k 10^3, hecto h 10^2, deca da 10^1, deci d 10^-1, centi c 10^-2, milli m 10^-3, micro µ 10^-6, nano n 10^-9, pico p 10^-12, femto f 10^-15, atto a 10^-18, zepto z 10^-21, yocto y 10^-24, ronto r 10^-27, quecto q 10^-30.
- Format rules for valid output: unit symbols are case-sensitive (`s` vs `S`, `m` vs `M`); prefix and unit symbol are joined without a space (`kJ`, `µs`); a space separates number and unit (`12 kg`), except the degree/minute/second of plane angle; `kg` already carries a prefix — never `µkg`; micro is Greek mu U+03BC (U+00B5 micro sign is the legacy compatibility character); ohm is U+03A9 (U+2126 OHM SIGN is deprecated); `°C` not `° C`.
- MIT alternative: faker-js `en/science/unit.ts` — 29 `{name, symbol}` objects, MIT.
### Animals by class
- **Wikidata** (https://www.wikidata.org/wiki/Wikidata:Licensing): all structured data **CC0** — no attribution required. Embed: Yes.
- Feasibility, live SPARQL (https://query.wikidata.org/sparql), pattern `?t wdt:P31 wd:Q16521; wdt:P105 wd:Q7432; wdt:P171* wd:<class>; wdt:P1843 ?cn FILTER(LANG(?cn)='en')` (taxon, rank species, parent-taxon chain, English common name):
- Mammalia Q7377: **5,759** species with ≥1 English common name.
- Aves Q5113: **12,432**.
- Reptilia Q10811: **16,369** — suspiciously large; Wikidata's parent chain likely pulls birds/Sauropsida in. Use Squamata Q122422 + Testudines Q223044 + Crocodilia Q1387 (or check the chain) instead.
- Insecta Q1390 and Plantae Q756: **timed out (502/504) every attempt** — trees too big for the 60 s public endpoint. Options: paginate by order (`P171` one level at a time), use the weekly JSON dump, or take GBIF instead.
- Sample rows (mammals): `Notoryctes typhlops` → "Central Desert Marsupial Mole", "Itjaritjari", "Marsupial Mole", "Southern marsupial mole", "Southern Marsupial Mole" — i.e. **several P1843 values per taxon, inconsistent casing, and indigenous-language names tagged `en`**. Validation: pick one name per taxon (prefer the one matching the item's English `rdfs:label`), normalise case, drop names that equal the scientific name.
- **GBIF Backbone Taxonomy** https://www.gbif.org/dataset/d7dddbf4-2cf0-4f39-9b2a-bb099caae36c — licence **CC BY 4.0** (verified via https://api.gbif.org/v1/dataset/d7dddbf4-2cf0-4f39-9b2a-bb099caae36c, `license: http://creativecommons.org/licenses/by/4.0/legalcode`); required citation text: "GBIF Secretariat (2023). GBIF Backbone Taxonomy. Checklist dataset https://doi.org/10.15468/39omei". Embed: Yes with that citation. Vernacular names come via `/v1/species/{key}/vernacularNames`, each row from a *different* source checklist with its own licence (some CC BY-NC) — record and check per row. GBIF's site terms page (https://www.gbif.org/terms) was unreachable (timeout) — **UNVERIFIED** beyond the API field.
- **faker-js** `src/locales/en/animal/*` (MIT; counts = quoted string lines on branch `next`): bird 821, snake 533, dog 493, cow 467, horse 342, rodent 161, insect 129, fish 95, cat 55, cetacean 52, rabbit 49, type 44 (generic: "bat, bear, bee, …, zebra"), pet_name 42, crocodilia 24, bear 8, lion 7. These are breeds/species mixed with common names, no class field; fine as MIT drop-ins. (A WebFetch summariser claimed 1,036 dogs; the regex count is 493 — treat the exact count as UNVERIFIED.)
### Plants
- Same Wikidata approach (CC0) — but the Plantae query timed out; paginate by order or family (`?t wdt:P171 ?family . ?family wdt:P105 wd:Q35409`) or use the dump.
- GBIF Plantae (kingdomKey 6) vernacular names, CC BY 4.0, same per-row-licence caveat.
- faker-js has **no** plant list (en locale dirs: airline, animal, app, book, color, commerce, company, database, date, finance, food, hacker, internet, location, lorem, medical, music, person, phone_number, science, team, vehicle, word).
- USDA PLANTS database (public domain, has common names) — **UNVERIFIED**, not fetched.
### Planets / moons
- NASA NSSDCA Planetary Fact Sheet https://nssdc.gsfc.nasa.gov/planetary/factsheet/ — 10 columns (Mercury, Venus, Earth, Moon, Mars, Jupiter, Saturn, Uranus, Neptune, Pluto) × 20 rows (mass 10^24 kg, diameter km, density, gravity, escape velocity, rotation period, day length, distance from Sun, perihelion, aphelion, orbital period, orbital velocity, inclination, eccentricity, obliquity, mean temp °C, surface pressure, number of moons, rings Y/N, magnetic field Y/N). HTML table only. NASA content "generally not subject to copyright in the United States" (https://www.nasa.gov/nasa-brand-center/images-and-media/); NASA asks to be acknowledged; the NASA insignia must not be used. Embed: Yes with "Source: NASA NSSDCA".
- Wikidata (CC0): planets = `?p wdt:P31/wdt:P279* wd:Q634; wdt:P397 wd:Q525` → Mercury Q308, Venus Q313, Earth Q2, Mars Q111, Jupiter Q319, Saturn Q193, Uranus Q324, Neptune Q332 (plus hypothetical Theia Q1053432 — filter out). Moons = `?m wdt:P397 <planet>; wdt:P31/wdt:P279* wd:Q2537` (natural satellite; note **Q2199 is "dwarf planet"**, not moon): Saturn 165, Jupiter 85, Uranus 29, Neptune 16, Earth 5, Mars 2. Caveats: Wikidata lags IAU counts (Saturn has 274 recognised moons as of 2025 — UNVERIFIED figure); Earth's 5 includes quasi-satellites; provisional designations (`S/2004 S 3`) are labels too — filter `^S/\d{4}` if you want proper names only.
### Blood type distribution
- **US** (Stanford Blood Center https://stanfordbloodcenter.org/donate-blood/blood-donation-facts/blood-types/, verified; cites AABB Technical Manual 18th ed.): O+ 37.4 %, A+ 35.7 %, B+ 8.5 %, AB+ 3.4 %, O− 6.6 %, A− 6.3 %, B− 1.5 %, AB− 0.6 % (sums to 100.0). Facts — freely embeddable; cite Stanford/AABB.
- American Red Cross https://www.redcrossblood.org/donate-blood/blood-types.html — **UNVERIFIED**: site returns HTTP 403 (Akamai) to both WebFetch and the headless browser. Search snippets say the Red Cross gives type O by ethnicity (Latino 57 %, African American 51 %, Caucasian 45 %) rather than an 8-way national table.
- **Global**: no authoritative figure found. Wikipedia "Blood type distribution by country" gives a population-weighted row (O+ 38.4, A+ 27.3, B+ 8.1, AB+ 2.0, O− 13.1, A− 8.1, B− 2.0, AB− 0.01) that its own editors flag "unreliable source"; WorldAtlas gives O+ 42, A+ 31, B+ 15, AB+ 5, O− 3, A− 2.5, B− 1, AB− 0.5 with no source. Recommend shipping only the US table, or per-country tables from Wikipedia (facts; Wikipedia prose is CC BY-SA but the numbers are cited to national blood services — cite those).
- Validation: ABO ∈ {A, B, AB, O}; Rh ∈ {+, −}; weights must sum to 100 ± rounding.
## 8. Text
### English word lists with part of speech
| Source | Licence | Embed | Counts | Format |
|---|---|---|---|---|
| Princeton WordNet 3.0 / 3.1 https://wordnet.princeton.edu/license-and-commercial-use ; download https://wordnetcode.princeton.edu/wn3.1.dict.tar.gz | "WordNet License" (SPDX `WordNet`), BSD-style: permission to use/copy/modify/distribute "for any purpose and without fee or royalty … provided that … the following copyright notice and statements, including the disclaimer … appear on ALL copies … including modifications". Notice text: "WordNet 3.0 Copyright 2006 by Princeton University. All rights reserved." + AS-IS disclaimer; no use of Princeton's name in advertising. | Yes — ship the LICENSE text alongside the data | WordNet 3.0 (wnstats(7WN), verified): unique strings noun 117,798 / verb 11,529 / adj 21,479 / adv 4,481 (total 155,287); synsets 82,115 / 13,767 / 18,156 / 3,621 | WNDB: `index.noun|verb|adj|adv` (one lemma per line: `lemma pos synset_cnt p_cnt [ptr…] sense_cnt tagsense_cnt synset_offset…`), `data.<pos>` (synsets with glosses). Multi-word lemmas use `_` for spaces; lower-case; entries can contain digits/apostrophes/hyphens. |
| Open English WordNet 2025 https://en-word.net/ ; https://github.com/globalwordnet/english-wordnet | **CC BY 4.0**; cite "Open English WordNet … derived from Princeton WordNet" | Yes, with attribution | 135,969 words, 107,519 synsets (README, verified) | `english-wordnet-2025.zip` (WNDB, 9.2 MB), `english-wordnet-2025.xml.gz` (WN-LMF, 10.8 MB), `english-wordnet-2025-json.zip` (9.5 MB), `.ttl.gz` (16.9 MB) |
| SCOWL v2 / ESDB https://github.com/en-wl/wordlist (branch v2) ; classic http://wordlist.aspell.net/ | Custom MIT-like (Copyright 2000-2026 Kevin Atkinson: use/copy/modify/distribute/sell "provided that the above copyright notice appears in all copies and that both the above copyright notice and this notice appear in supporting documentation"). Sources 12dicts + ENABLE2K are public domain; lists above size 80 add the UKACD notice; **POS assignments partly from WordNet, so the WordNet notice "MIGHT apply"** (their words). Australian-English parts carry an extra notice. | Yes, ship the Copyright file | Sizes 35 (small), 50 (medium), 60 (spell-check default), 70 (large), 80 (incl. game words), 85 (archaic). No per-size counts published in the v2 README. Classic SCOWL sizes 10–95 — counts **UNVERIFIED** (README not reachable). | v2: `scowl.db` (SQLite), `scowl.txt`, Python `libscowl`. ESDB rows carry POS, spelling (A/B/Z/C/D) and region codes. |
| 12dicts http://wordlist.aspell.net/12dicts-readme/ | Public domain ("I explicitly release them to the public domain, but request acknowledgment"), **except** `2of12inf` and the `2+2+3*` lists, which depend on AGID and inherit its terms | Yes (PD lists); acknowledge Alan Beale | 3esl ≈22k, 6of12 ≈32k, 2of12 ≈41k, 2of12inf ≈82k, 3of6game ≈65k, 5d+2a ≈68k, 3of6all ≈83k, 2+2+3lem ≈84k, 2of5core ≈4.7k, 6phrase ≈22k, neol2016 ≈600 | Plain text, one word per line; marker suffixes `+ ! ^ &` (not POS). **No POS tags** in any current 12dicts list. |
| Moby Part-of-Speech II https://www.gutenberg.org/ebooks/3203 ; file https://www.gutenberg.org/files/3203/files/mobypos.txt | Public domain ("Public Domain material by grant from the author, January, 2001", in the package README) | Yes, no notice needed | **233,356** lines (verified) | One entry per line: `word\CODES`, CRLF line ends (e.g. `A-line\NA`, `a tempo\Avh`). Codes in priority order: `N` noun, `p` plural, `h` noun phrase, `V` verb (usu. participle), `t` transitive verb, `i` intransitive verb, `A` adjective, `v` adverb, `C` conjunction, `P` preposition, `!` interjection, `r` pronoun, `D` definite article, `I` indefinite article, `o` nominative. Many entries are phrases, proper nouns, or 1990s-era; non-ASCII entries use a legacy 8-bit encoding (e.g. `a bon march\v` lost its `é`) — restrict to ASCII `[a-z]+` for a clean list. |
| faker-js `en/word/*` (MIT) | MIT | Yes | adjective 1000, noun 1000, verb 1000, adverb 325, preposition 109, conjunction 51, interjection 46 | TS string arrays |
### Word frequency lists
| Source | Licence | Verdict |
|---|---|---|
| Google Books Ngram data v2/v3 https://storage.googleapis.com/books/ngrams/books/datasetsv3.html | **CC BY 3.0** ("This compilation is licensed under a Creative Commons Attribution 3.0 Unported License") | Embeddable with attribution. Format `ngram TAB year TAB match_count TAB volume_count`; 1-gram files split by leading letter, GBs each — aggregate offline to a top-N list. |
| wordfreq https://github.com/rspeer/wordfreq | Code **Apache 2.0** (not MIT). Data: "may be redistributed under a Creative Commons Attribution-ShareAlike 4.0 license" (mix of Google Books, Wikipedia, OpenSubtitles, SUBTLEX, Leeds, ParaCrawl, Twitter); maintainer notes the CSV export "does not follow the CC-By-SA license". Project in sunset mode (data snapshot ≈2021). | **Avoid** embedding the data (ShareAlike). |
| Peter Norvig count_1w.txt https://norvig.com/ngrams/ | Page says "Code … under the MIT license"; the **data** comes from the Google Web 1T corpus distributed by LDC under LDC terms — no data licence stated. | **UNVERIFIED / avoid**. 333,333 words, `word TAB count`, lowercase. |
| hermitdave/FrequencyWords https://github.com/hermitdave/FrequencyWords | "MIT License for code. CC-by-sa-4.0 for content." | **Avoid** (ShareAlike). |
### Lorem ipsum
- Canonical paragraph (Letraset, 1966; public domain — garbled Cicero): "Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum."
- Source: Cicero, *De finibus bonorum et malorum* 1.10.32–33 (45 BC), public domain; the "dolorem ipsum" fragment begins "Neque porro quisquam est qui dolorem ipsum quia dolor sit amet, consectetur, adipisci velit…". Reference: https://en.wikipedia.org/wiki/Lorem_ipsum
- MIT word pool alternative: faker-js `en/lorem/word.ts` — 999 words.
### Quote collections
- Project Gutenberg #27889 https://www.gutenberg.org/ebooks/27889 — Bartlett, *Familiar Quotations*, **9th edition** (title page "NINTH EDITION", copyright 1875/1882/1891/1903) — not the 1919 10th. "Public domain in the USA." Plain text 3.1 MB, HTML zip 50 MB, EPUB.
- Project Gutenberg #16732 https://www.gutenberg.org/ebooks/16732 — an early (Hurst & Co. reprint) edition, no edition number given. PD in USA.
- 10th edition (1914/1919, ed. Nathan Haskell Dole) exists on archive.org (https://archive.org/details/bartlettsfamilia0000john_o4b4) — not found on Gutenberg; **UNVERIFIED** availability as clean text.
- Either Bartlett text needs parsing (quote / author / source headings); no structured file. Wikiquote is CC BY-SA — avoid.
- Validation: strip Gutenberg header/footer (`*** START OF THE PROJECT GUTENBERG EBOOK` … `*** END …`) before extraction; Gutenberg's trademark terms only bind if you keep the "Project Gutenberg" name on the output — strip it.
### Colour names
| Source | Licence | Embed | Count / format |
|---|---|---|---|
| CSS Color Module Level 4 §named colors https://www.w3.org/TR/css-color-4/#named-colors | W3C Document License (2023) https://www.w3.org/copyright/document-license-2023/ — copy/distribute allowed for any purpose; derivatives allowed only "to facilitate implementation of the technical specifications"; notice required: "Copyright © 2023 W3C®. This software or document includes material copied from or derived from [title and URI]." | Yes — the name→hex table is factual data and is copied into every browser; include the W3C notice line with the spec title + URI. | **148** names (verified: aliceblue … yellowgreen, incl. `rebeccapurple #663399`); all names lowercase ASCII, hex 6 lowercase digits. |
| XKCD colour survey https://xkcd.com/color/rgb.txt | **CC0** (file header: `# License: https://creativecommons.org/publicdomain/zero/1.0/`) | Yes, no notice needed | **949** lines (verified). Format `name TAB #rrggbb TAB` (note the trailing tab). Max name length 26; all hex valid lowercase; no duplicate names. Contains crude names ("poo", "baby poop green", "diarrhea", "vomit") — apply a blocklist. |
| faker-js `en/color/human.ts` | MIT | Yes | 31 names |
| Pantone | Proprietary | **Avoid** | — |
## 9. Dates
### IANA tz database
- Source https://www.iana.org/time-zones ; files https://data.iana.org/time-zones/tzdb/ (current version **2026d**, verified). LICENSE: "Unless specified below, all files in the tz code and data (including this LICENSE file) are in the public domain" (only `date.c`, `newstrftime.3`, `strftime.c` are BSD-3). Embed: Yes, no notice needed.
- `zone1970.tab` — **312** data rows; tab-separated; UTF-8; `#` comments. Columns: (1) comma-separated ISO 3166-1 alpha-2 codes of countries overlapping the zone, most-populous first; (2) coordinates in ISO 6709 `±DDMM±DDDMM` or `±DDMMSS±DDDMMSS` (lat then lon, no separator); (3) TZ name (`Europe/Zurich`); (4) comment, present iff the country has several zones. One row per timezone where civil time has agreed since 1970.
- `zone.tab` — **418** rows, older one-country-per-row table (column 1 is a single country code); kept for backward compatibility. Prefer `zone1970.tab`.
- `backward` — **252** `Link TARGET LINK-NAME` lines mapping old/merged names (`US/Eastern`, `Asia/Calcutta`, `Europe/Kiev`) to current ones; also a few `Zone` entries. Use it to (a) accept aliases as input and (b) never emit them as output.
- UTC offsets: do **not** embed — offsets depend on DST rules that change several times a year. Derive at runtime: `new Intl.DateTimeFormat('en-US', { timeZone, timeZoneName: 'longOffset' }).formatToParts(date)` → part `timeZoneName` = `GMT-05:00` (or `GMT` for zero) — strip `GMT`, treat bare `GMT` as `+00:00`. Validation of a generated zone name: `Intl.supportedValuesOf('timeZone')` (canonical names only; whether links like `US/Eastern` appear depends on the ICU build) or `try { new Intl.DateTimeFormat(undefined, { timeZone }) } catch { invalid }`, which accepts links too.
### ISO 8601 layouts (https://en.wikipedia.org/wiki/ISO_8601)
| Layout | Extended | Basic | Example |
|---|---|---|---|
| Calendar date | `YYYY-MM-DD` | `YYYYMMDD` | 2009-01-06 |
| Ordinal date | `YYYY-DDD` | `YYYYDDD` | 1981-095 |
| Week date | `YYYY-Www-D` | `YYYYWwwD` | 2009-W01-1 (weeks start Monday, week 1 contains 4 Jan) |
| Time | `hh:mm:ss[.fff]` | `hhmmss` | 13:47:30 (24 h; `T` prefix optional standalone) |
| UTC / offset | `Z` or `±hh:mm` | `±hhmm`, `±hh` | 2007-04-05T14:30Z, 2024-06-01T09:00-05:00 |
| Date-time | date `T` time offset | | 2007-04-05T14:30:00+02:00 |
| Duration | `PnYnMnDTnHnMnS` / `PnW` | | P3Y6M4DT12H30M5S |
| Interval | start`/`end, start`/`duration, duration`/`end | | 2007-03-01T13:00:00Z/2008-05-11T15:30:00Z |
| Recurring | `Rn/`interval | | R5/2008-03-01T13:00:00Z/P1Y2M10DT2H30M |
RFC 3339 profile (what JSON/APIs usually mean): extended forms only, `T`/`Z` may be lowercase, space allowed instead of `T`, `-00:00` = unknown offset, no durations/intervals/week/ordinal dates. Generate RFC 3339 by default.
### Unicode CLDR (locale month/weekday names, date patterns)
- Licence: **Unicode License v3** (SPDX `Unicode-3.0`), OSI-approved 2023-11-17 (https://opensource.org/license/unicode-3-0). Text https://www.unicode.org/license.txt — MIT-style: use/copy/modify/sell "provided that either (a) this copyright and permission notice appear with all copies of the Data Files or Software, or (b) this copyright and permission notice appear in associated Documentation"; no use of the Unicode name in advertising. Embed: Yes; put the notice in LICENSE/NOTICE.
- npm: `cldr-core`, `cldr-dates-full`, `cldr-numbers-full` (peer of dates), etc.; current **48.2.0** (CLDR 48, verified via npm registry, `license: Unicode-3.0`); repo https://github.com/unicode-org/cldr-json .
- Path is **`cldr-dates-full/main/<locale>/ca-gregorian.json`** (not `dates/gregorian.json`), rooted at `main.<locale>.dates.calendars.gregorian`. Keys (verified for `en`): `months`, `days`, `quarters`, `dayPeriods`, `eras`, `dateFormats`, `dateSkeletons`, `timeFormats`, `timeSkeletons`, `dateTimeFormats`, `dateTimeFormats-atTime`, `dateTimeFormats-relative`.
- `months.{format|stand-alone}.{abbreviated|narrow|wide}` keyed `"1".."12"`; `days.{format|stand-alone}.{abbreviated|narrow|short|wide}` keyed `sun..sat`.
- `dateFormats`: en = full `EEEE, MMMM d, y`, long `MMMM d, y`, medium `MMM d, y`, short `M/d/yy`.
- `timeFormats`: en = full `h:mm:ss a zzzz`, long `h:mm:ss a z`, medium `h:mm:ss a`, short `h:mm a` — **the `a` is preceded by U+202F NARROW NO-BREAK SPACE**; `*-alt-ascii` variants use a plain space. Pick one spelling (ASCII) and ignore the other.
- `dateTimeFormats.{full|long|medium|short}` = glue like `{1}, {0}` ({1} date, {0} time); `availableFormats` = skeleton→pattern map (`yMMMd` → `MMM d, y`), plus `intervalFormats`, `appendItems`.
- `dayPeriods.format.abbreviated`: `am: AM`, `pm: PM`, plus `midnight`, `noon`, `morning1`… variants.
- Pattern letters follow LDML (https://unicode.org/reports/tr35/tr35-dates.html#Date_Field_Symbol_Table): `y` year, `M`/`L` month (format/stand-alone), `d` day, `E`/`c` weekday, `h` 1–12, `H` 0–23, `a` AM/PM, `z`/`zzzz` zone name, `x`/`X` ISO offset; literal text in single quotes. Same patterns drive `Intl.DateTimeFormat` internally, so runtime `Intl` can replace embedding for month/weekday names if bundle size matters.
## Summary of licence verdicts
- Embed freely (PD/CC0): PubChem periodic table, NIST, NASA fact sheet, Wikidata, Moby POS, 12dicts (PD lists), Bartlett via Gutenberg, XKCD colours, IANA tz, Lorem ipsum/Cicero.
- Embed with notice: WordNet 3.x (WordNet License), Open English WordNet (CC BY 4.0), SCOWL/ESDB (custom MIT-like), GBIF backbone (CC BY 4.0 + citation), BIPM SI (CC BY 4.0), CSS named colours (W3C notice), CLDR (Unicode-3.0), Google Books ngrams (CC BY 3.0), faker-js lists (MIT).
- Avoid: Bowserinator periodic table (CC BY-SA 3.0), wordfreq data (CC BY-SA 4.0), hermitdave (CC BY-SA 4.0), Norvig count_1w (LDC-derived, unstated), Wikiquote (CC BY-SA), Pantone.
- UNVERIFIED: American Red Cross figures (site blocks fetches), any global blood-type table, BIPM PDF front-matter licence line, classic SCOWL per-size counts, USDA PLANTS, IAU moon counts, exact faker dog count, Bartlett 10th-edition text availability.
# Part 4 — ISO lists, commerce, books/music/food
Researched 2026-09-17. "Embed?" = can the data ship inside an MIT-licensed repo. Anything marked UNVERIFIED was not confirmed against the primary source.
## 10. ISO lists
### ISO 3166-1 countries
| Source | Licence | Embed in MIT repo? | Rows | Fields | Format |
|---|---|---|---|---|---|
| [datasets/country-codes](https://github.com/datasets/country-codes) | PDDL (README: "Public Domain Dedication and License"; GitHub API detects no licence file) | Yes, no attribution required. README caveat: ISO itself says its list is "for internal use and non-commercial purposes free of charge" — the repo argues no rights subsist in a list of facts | 249 | ISO3166-1-Alpha-2/-Alpha-3/-numeric, official_name_en, UNTERM names (ar/zh/en/fr/ru/es short+formal), Dial, TLD, Capital, Continent, Languages, ISO4217 currency code/name/numeric/minor_unit, Region/Sub-region/Intermediate names+codes (M49), FIPS, IOC, FIFA, MARC, ITU, WMO, GAUL, EDGAR, Geoname ID, wikidata_id, CLDR display name, is_independent, LDC/LLDC/SIDS flags | CSV `data/country-codes.csv` |
| [annexare/Countries](https://github.com/annexare/Countries) (npm `countries-list` 3.4.1, 2026-07-14) | MIT | Yes; keep MIT notice | 252 countries, 185 languages (incl. non-ISO entries like AC, XK) | per country: name, native, phone[] (calling codes), continent, capital, currency[], languages[], optional alias[], partOf; separate ISO 4217 currency table (name, native, symbol, numeric, decimals); languages: name, native | TS source, exported JSON/CSV/SQL |
| [Debian iso-codes](https://salsa.debian.org/iso-codes-team/iso-codes) | LGPL-2.1-or-later (Debian copyright file) | Not safely as embedded data in an MIT repo — LGPL applies to the data files; you would have to ship it as a separately licensed component with LGPL notice. Avoid unless accepting that | 3166-1: 249 | alpha_2, alpha_3, numeric, name, official_name, common_name, flag | JSON `data/iso_3166-1.json` (raw URL needs `?inline=false`) |
| [mledoze/countries](https://github.com/mledoze/countries) | ODbL-1.0 | No — share-alike + attribution; derived databases must be ODbL. Avoid | ~250 | common/official names, cca2/cca3/ccn3/cioc, tld, currencies, idd (calling codes), capital, region/subregion, languages, translations, latlng, borders, area, demonyms, flag emoji, UN member | JSON/CSV/XML/YAML |
| [restcountries](https://gitlab.com/restcountries/restcountries) | MPL-2.0 (source repo; hosted API v5 is proprietary/API-key) | File-level copyleft: a copied data file stays MPL and must carry the notice; OK to ship next to MIT code but not to relicense. Prefer PDDL/MIT sources | 250+ | name, tld, cca2, ccn3, cca3, cioc, fifa, independent, status, unMember, currencies, idd, capital, capitalInfo, altSpellings, region, subregion, continents, languages, translations, latlng, landlocked, borders, area, flag (emoji), demonyms, flags, coatOfArms, population, maps, gini, car, postalCode, startOfWeek, timezones | JSON `src/main/resources/countriesV3.1.json` |
| [lukes/ISO-3166-Countries-with-Regional-Codes](https://github.com/lukes/ISO-3166-Countries-with-Regional-Codes) | CC BY-SA 4.0 (LICENSE.md) | No — share-alike. Avoid | 249 | name, alpha-2, alpha-3, country-code, iso_3166-2, region, sub-region, intermediate-region + codes | CSV/JSON/XML (all, slim-2, slim-3) |
| [stefangabos/world_countries](https://github.com/stefangabos/world_countries) | Data: CC BY-SA 4.0 (README, GitHub licence detection); npm `world_countries_lists` package.json says LGPL-3.0-or-later — contradictory | No — share-alike either way. Avoid | 249 world / 193 UN; subdivisions 5046 | id (numeric), alpha2, alpha3, name (37 languages); subdivisions: country, code, name, name_en, type, parent | JSON/CSV/PHP/SQL/XML |
| [i18n-iso-countries](https://www.npmjs.com/package/i18n-iso-countries) 7.14.0 | MIT | Yes | 78 language files | alpha2, alpha3, numeric, localized names | JSON per language |
| Wikipedia ISO 3166-1 | CC BY-SA 4.0 | No (share-alike) | — | — | — |
| [ISO OBP](https://www.iso.org/obp) | Proprietary; ISO grants free use only for "internal use and non-commercial purposes" (quoted in datasets/country-codes README) | No | — | — | — |
- Flag emoji: derive, no data needed — for alpha-2 `XY`, emit U+1F1E6 + (X − 'A') and U+1F1E6 + (Y − 'A') (regional indicator symbols).
- Recommendation: `datasets/country-codes` (PDDL) covers every requested field (name, alpha2, alpha3, numeric, calling code, TLD, capital, currency, languages) in one CSV; flag emoji derived. `annexare` (MIT) is the fallback for native names/calling codes if PDDL's ISO caveat worries you; it lacks TLD.
- Validation rules: alpha-2 `^[A-Z]{2}$`, alpha-3 `^[A-Z]{3}$`, numeric `^\d{3}$` (zero-padded, keep as string); calling code 1–3 digits, NANP countries share `1`; TLD `^\.[a-z]{2}$` (ccTLD = lowercase alpha-2, exceptions: GB uses `.uk`).
### ISO 639 languages
| Source | Licence | Embed? | Rows | Fields | Format |
|---|---|---|---|---|---|
| [langs](https://www.npmjs.com/package/langs) 2.0.0 (2017, unmaintained) | MIT | Yes | 184 | name, local, 1 (639-1), 2, 2T, 2B, 3 | `data.js` array |
| [iso-639-1](https://www.npmjs.com/package/iso-639-1) 3.1.6 (2026-07-02) | MIT | Yes | 183 | code, name, nativeName | `src/data.js` |
| annexare/Countries languages | MIT | Yes | 185 | code (639-1), name, native | TS/JSON |
| Debian iso-codes 639-2 / 639-3 / 639-5 | LGPL-2.1+ | See above (LGPL) | 639-2: 487; 639-3: 7923 | 639-2: alpha_2, alpha_3, bibliographic, name, common_name; 639-3: + inverted_name, scope, type | JSON |
| [SIL ISO 639-3 tables](https://iso639-3.sil.org/code_tables/download_tables) | SIL terms: may incorporate into software (commercial or not) with attribution to iso639-3.sil.org, must not modify identifiers, and product "does not provide a means to redistribute the code set" | Risky — a git repo with the table IS a means to redistribute; avoid embedding the full table. Codes themselves are facts | 7927 | Id, Part2b, Part2t, Part1, Scope, Language_Type, Ref_Name, Comment | tab-delimited `iso-639-3.tab` |
| [Library of Congress ISO 639-2](https://www.loc.gov/standards/iso639-2/ascii_8bits.html) | US federal government work → public domain in the US | Yes | ~487 (matches Debian 639-2 count) | pipe-delimited: alpha3-B, alpha3-T, alpha2, English name, French name — UNVERIFIED: loc.gov is behind a Cloudflare bot check (WebFetch, curl and Playwright all blocked); field layout from memory | `ISO-639-2_utf-8.txt` |
- Recommendation: `langs` or `iso-639-1` (MIT) for 639-1 names; both are small and derived from public code tables. Anything 639-3-sized is LGPL or SIL-restricted.
- Rules: 639-1 `^[a-z]{2}$`, 639-2/3 `^[a-z]{3}$`; 639-2 has B/T pairs (e.g. `fre`/`fra`, `ger`/`deu`) — pick one consistently (T codes = 639-3).
### ISO 3166-2 subdivisions
| Source | Licence | Embed? | Rows | Fields |
|---|---|---|---|---|
| Debian iso-codes 3166-2 | LGPL-2.1+ | LGPL caveat | 5046 | code, name, type, parent |
| [olahol/iso-3166-2.json](https://github.com/olahol/iso-3166-2.json) | package.json says ISC; no LICENSE file; GitHub detects none; data scraped from eQuest xls (provenance unclear); last push 2021 | No — unclear licence and provenance | 237 countries / 3807 divisions | `{alpha2: {name, divisions: {code: name}}}` |
| [esosedi/3166](https://github.com/esosedi/3166) (npm `iso3166-2-db` 2.3.11) | MIT | Yes; data mixes GeoNames (CC BY 4.0), Wikipedia (CC BY-SA), OSM (ODbL) — the MIT label may not survive scrutiny for the Wikipedia/OSM-derived parts. UNVERIFIED how much is derived from each | "all countries and regions" (count not stated) | ISO 3166-2 + FIPS codes, names in 12 languages, admin level, GeoNames/OSM/Wikipedia/WOF ids |
| stefangabos subdivisions | CC BY-SA 4.0 | No | 5046 | country, code, name, name_en, type, parent |
| [US Census state.txt](https://www2.census.gov/geo/docs/reference/state.txt) | US government work → public domain | Yes | 57 (50 states + DC + 6 territories) | STATE (FIPS), STUSAB, STATE_NAME, STATENS |
- Recommendation: for an en_US library, ship only the US list (Census/USPS, public domain: 50 states + DC, optionally territories) — ISO 3166-2 code = `US-` + USPS abbreviation. Skip the ~5000-row world list; every open compilation is LGPL, CC BY-SA or unclear.
- Rule: `^[A-Z]{2}-[A-Z0-9]{1,3}$`.
## 11. Commerce
### Word lists (product name / adjective / material)
- [@faker-js/faker](https://github.com/faker-js/faker) 10.6.0 — MIT (LICENSE covers code and locale data; copyright Faker contributors 2022-2025 and Marak Squires 2011-2020; acknowledges Ruby Faker and Perl Data::Faker, both MIT). No separate data licence.
- Provenance policy in [CONTRIBUTING.md](https://github.com/faker-js/faker/blob/next/CONTRIBUTING.md): "Faker must not contain copyrighted materials"; facts (finite known lists) are OK; compilations must not be copied from a single source (Wikipedia, "most popular" articles, government sites) because "a compilation of facts can be copyrighted" — contributors are told to draw on multiple sources and their own judgement.
- Verdict: reusing faker's `en` commerce lists (`product_name.adjective/material/product`, `department`) is fine under MIT — keep the MIT notice with the copied data. The lists are short curated words (no external dataset). Nothing in the repo or docs restricts locale-data reuse beyond MIT.
### Garment and shoe sizes
- Letter sizes XS–XXL: EN 13402-3 letter codes with chest/bust ranges (Wikipedia [EN 13402](https://en.wikipedia.org/wiki/EN_13402), CC BY-SA — but the table is a handful of facts; the standard itself is paywalled). Men chest / women bust cm: XXS 70–78 / 66–74, XS 78–86 / 74–82, S 86–94 / 82–90, M 94–102 / 90–98, L 102–110 / 98–107, XL 110–118 / 107–119, XXL 118–129 / 119–131, 3XL 129–141 / 131–143.
- US numeric women's sizes (0–20, even) come from ASTM D5585 (paywalled); no open dataset. Just generate the even-number series.
- Shoe sizes: no open dataset found (search hits are ML demos and retailer charts). Use the formulas from Wikipedia [Shoe size](https://en.wikipedia.org/wiki/Shoe_size) (facts, not copyrightable): EU (Paris point) ≈ 1.5 × foot cm + 2 (i.e. 3/20 × mm + 2); UK adult ≈ 3 × foot in − 23; US men ≈ 3 × foot in − 22 = UK + 1; US women (common) = US men + 1.5 (FIA scale: 3 × foot in − 21); Mondopoint = foot length mm. ISO/TS 19407:2023 has the official conversion tables (paywalled). Generate from a foot length (men 24–31 cm, women 21–27 cm) in half-size steps so all systems agree.
### Check-digit rules (all verified against Wikipedia articles; algorithms are facts)
| Code | Rule |
|---|---|
| EAN-13 / UPC-A / GTIN-8/12/13/14 | Weights 3,1,3,1… starting from the rightmost data digit (the one next to the check digit); check = (10 − sum mod 10) mod 10. UPC-A = EAN-13 with a leading 0 (GS1 prefix of a 12-digit GTIN is `0` + first two digits). GTIN-14 first digit is an indicator 1–8 (packaging level) or 9 (variable measure); 0 is not valid there. |
| UPC-A number system (first digit) | 0,1,6,7,8,9 regular; 2 variable-weight; 3 drugs (NDC); 4 in-store/local; 5 coupons. |
| ISBN-10 | Σ(weights 10..2 × first 9 digits) + check ≡ 0 mod 11; check = (11 − sum mod 11) mod 11, 10 → `X`. |
| ISBN-13 | EAN-13 with prefix 978 or 979. 979 groups in use: 979-8 USA, 979-10 France, 979-11 Korea, 979-12 Italy. 978 groups 0 and 1 = English-language. For an en_US generator: `978-0…`, `978-1…`, `979-8…`. The official range file (`https://www.isbn-international.org/export_rangemessage.xml`) is XML only and the site says "You must not republish material from our site … without our written permission" — do not embed it; group/registrant length only matters for hyphenation, not validity. |
| ISSN | Format `NNNN-NNNC`; weights 8..2 over first 7 digits, sum mod 11; check = 0 if remainder 0 else 11 − remainder; 10 → `X`. EAN-13 form: `977` + 7 ISSN digits (no check) + 2 variant digits + EAN check. |
| IMEI | 15 digits: TAC 8 (RBI 2 + 6) + serial 6 + Luhn check (ISO/IEC 7812; from the right, double every second digit, sum digits, total ≡ 0 mod 10). SVN never enters the check. IMEISV = 14 digits + 2-digit SVN, no check. RBI `00` = test IMEI (Wikipedia [Reporting Body Identifier](https://en.wikipedia.org/wiki/Reporting_Body_Identifier)); live RBIs: 01 CTIA/PTCRB (US), 35 TÜV SÜD/BABT (UK), 86 TAF (China), 99 GHA. For fake data use TAC `00xxxxxx` so it can never match a real handset — UNVERIFIED: third-party pages claim the GSMA TS.06 test-TAC shape is `00 44` + 4 digits; the TS.06 PDF could not be text-extracted here. |
| ASIN | 10 chars `[A-Z0-9]{10}`; non-books start `B0` + 8 alphanumerics (`^B0[A-Z0-9]{8}$`); books reuse their ISBN-10 verbatim. No check digit (Wikipedia). |
| SKU | Not standardised (Wikipedia). Convention only: 8–12 uppercase alphanumerics, e.g. `[A-Z]{3}-\d{4}-[A-Z]{2}`; avoid leading 0 and O/I if you want scanner-safe. |
### GS1 prefixes
- Official table: [gs1.org/standards/id-keys/company-prefix](https://www.gs1.org/standards/id-keys/company-prefix) (rendered via Playwright; WebFetch gets 403). Copyright notice ([gs1.org/terms-use](https://www.gs1.org/terms-use)): reproduction allowed only "in unaltered form … for your personal, non-commercial use or use within your organisation"; "You are not permitted to re-transmit, distribute or commercialise the information or material without seeking prior written approval from GS1." → Do not copy the GS1 page. The prefix-to-country mapping is factual and also on Wikipedia ([List of GS1 country codes](https://en.wikipedia.org/wiki/List_of_GS1_country_codes), CC BY-SA 4.0 — avoid verbatim copy). Safe path: embed only the US/special ranges you need, written from the facts below.
- Ranges relevant to an en_US generator (from the GS1 page): 001–019, 030–039, 060–139 GS1 US (UPC-A compatible); 020–029 restricted circulation within a region; 040–049 restricted within a company; 050–059 GS1 US reserved; 200–299 restricted circulation (region); 952 "used for demonstrations and examples of the GS1 system"; 977 ISSN; 978–979 ISBN; 980 refund receipts; 981–983 coupons (common currency); 990–999 coupons. Prefixes do not identify country of origin. Selected others: 300–379 France, 400–440 Germany, 450–459 & 490–499 Japan, 500–509 UK, 690–699 China, 730–739 Sweden, 750 Mexico, 754–755 Canada, 760–769 Switzerland, 800–839 Italy, 840–849 Spain, 870–879 Netherlands, 880–881 South Korea, 890 India, 930–939 Australia, 940–949 New Zealand.
- Generator rule: for fake UPC/EAN use prefix `952` (GS1 demo) or a `2xx` restricted-circulation prefix — never a real company prefix; for US-looking UPC-A use number system 0–1/6–8 with a random 5-digit manufacturer code and accept that it may collide with real products.
## 12. Books, music, food
### Project Gutenberg catalog
- URL: `https://www.gutenberg.org/cache/epub/feeds/pg_catalog.csv` (20 MB) / `pg_catalog.csv.gz` (5.3 MB), regenerated weekly; also RDF (`rdf-files.tar.bz2`, 121 MB) and MARC.
- Fields: `Text#, Type, Issued, Title, Language, Authors, Subjects, LoCC, Bookshelves` (comma CSV, quoted; Authors as `Last, First, birth-death`, `;`-separated).
- Count (2026-09-17): 79 381 rows; Type: Text 78 130, Sound 1 114, Dataset 89, Image 33, other 12; Language: en 62 860, fr 4 190, fi 3 684, de 2 425, it 1 110.
- Licence: each RDF record carries `<cc:Work><cc:license rdf:resource="https://creativecommons.org/publicdomain/zero/1.0/"/>` → catalog metadata is CC0 1.0 (verified on `cache/epub/1/pg1.rdf`). The CSV itself has no header notice; treat as CC0 by the same source. Terms of use: do not hammer their servers (bulk feeds are the sanctioned path), "Project Gutenberg" is a trademark — royalties for commercial use of the *name*; do not brand the data file with it beyond a source citation.
- Embed: yes (CC0). For a fake-data library ship a filtered subset (Type=Text, Language=en, title + first author, maybe 5–10k rows) rather than 20 MB.
### Music genres and instruments
| Source | Licence | Embed? | Count | Notes |
|---|---|---|---|---|
| [MusicBrainz genre list](https://musicbrainz.org/genres) | Genre is a core entity (schema doc lists 13 core entities incl. Genre) → core data dump `mbdump.tar.bz2` is CC0. The genre→entity *associations* come via user tags (CC BY-NC-SA 3.0) — do not use those | Yes, names only | 2 202 (`/ws/2/genre/all?fmt=json`, `genre-count`) | WS API needs a descriptive User-Agent; list is a JSON array of `{id, name, disambiguation}` |
| [Discogs data dumps](https://data.discogs.com/) | CC0 | Yes | genre/style vocab is embedded in release XML (monthly dumps, GBs) — no standalone style list; UNVERIFIED count (~15 genres, ~600 styles from memory) | Impractical to extract; prefer MusicBrainz |
| ID3v1 genres | Names 0–79 from the 1999 ID3v1 spec, 80–125 Winamp, 126–191 Winamp 5.6 (2010) — a de-facto spec table, treated as public domain facts | Yes | 192 (0–191) | Wikipedia [List of ID3v1 genres](https://en.wikipedia.org/wiki/List_of_ID3v1_genres) (CC BY-SA page; the list itself is a spec enumeration). id3.org returned HTTP 500 during research — UNVERIFIED against the original spec page |
| Wikidata instruments | CC0 | Yes | 9 194 items `P31/P279* Q34379` with English labels; 2 372 with a Hornbostel–Sachs number (P1762) | Filter by P1762 for a clean orchestral/folk instrument list |
| Wikidata music genres | CC0 | Yes | 6 619 (`Q188451`) | Noisier than MusicBrainz |
### USDA FoodData Central
- Downloads: [fdc.nal.usda.gov/download-datasets](https://fdc.nal.usda.gov/download-datasets/) (April 2026 release):
- Foundation Foods: JSON 459 KB zip / 6.5 MB; CSV 3.7 MB zip / 32 MB — 394 foods (API `totalHits`, dataType=Foundation).
- SR Legacy (final, April 2018): JSON 12.3 MB zip / 205 MB; CSV 6.7 MB zip / 54 MB — 7 793 foods.
- FNDDS 2021-2023: CSV 200 MB zip / 1.6 GB — 5 432 foods.
- Branded: CSV 428 MB zip / 2.9 GB — 433 403 foods.
- Full: CSV 460 MB zip / 3.1 GB.
- Licence ([API guide](https://fdc.nal.usda.gov/api-guide/)): "USDA FoodData Central data are in the public domain and they are not copyrighted. They are published under CC0 1.0 Universal (CC0 1.0)". No permission needed; USDA *requests* listing FoodData Central as source and notifying them. Suggested citation: "U.S. Department of Agriculture, Agricultural Research Service. FoodData Central, 2019. fdc.nal.usda.gov."
- Embed: yes (CC0). Ship a derived list (SR Legacy `description` + `food_category`) not the raw dump.
- CSV structure (UNVERIFIED — the field-description PDF could not be text-extracted; from memory): `food.csv` (fdc_id, data_type, description, food_category_id, publication_date), `food_category.csv` (id, code, description), `nutrient.csv` (id, name, unit_name, nutrient_nbr, rank), `food_nutrient.csv` (id, fdc_id, nutrient_id, amount, …), `food_portion.csv`, `measure_unit.csv`, `sr_legacy_food.csv` (fdc_id, NDB_number), `foundation_food.csv`.
### Dishes, cuisines, drinks
| Source | Licence | Embed? | Count | Notes |
|---|---|---|---|---|
| Wikidata dishes (`Q746549`) | CC0 | Yes | 7 480 with English label | SPARQL `?i wdt:P31/wdt:P279* wd:Q746549` |
| Wikidata cuisines (`Q1968435`) | CC0 | Yes | 219 | |
| Wikidata cocktails (`Q134768`) | CC0 | Yes | 305 | |
| Wikidata beer (`Q44` subclass tree) / wine (`Q282`) | CC0 | Yes | 420 / 3 025 | Wine tree is mostly appellations/brands; filter by P279 depth |
| IBA official cocktails | IBA site "© IBA 2026 – All rights reserved"; Wikipedia list CC BY-SA 4.0 | Names only are facts (102 cocktails: 34 Unforgettables, 34 Contemporary Classics, 34 New Era, 2024 revision); do not copy recipes/descriptions | 102 | Safest: take the names from Wikidata (`P31 Q134768`) rather than the IBA page |
| [Open Brewery DB](https://github.com/openbrewerydb/openbrewerydb) | MIT (LICENSE, © 2025 Open Brewery DB) | Yes | 11 931 breweries (US 8 308, DE 1 445, AU 514, BE 478, CA 283, NZ 243) | `breweries.csv`: id (UUID), name, brewery_type (micro 5 906, brewpub 3 947, closed 644, planning 635, regional 240, contract 208, large 137, proprietor, taproom, bar, nano, cidery, beergarden), address_1..3, city, state_province, postal_code, country, phone, website_url, longitude, latitude. Real businesses — fine for names, but generating "fake" data with real addresses/phones may be undesirable |
| Wikipedia lists (cuisines, dishes, IBA) | CC BY-SA 4.0 | No verbatim copies | | |
## Summary of safe picks
- Countries: `datasets/country-codes` (PDDL) or `annexare/Countries` (MIT); flag emoji derived from alpha-2.
- Languages: `langs` / `iso-639-1` (MIT, 639-1 only).
- Subdivisions: US-only list from Census (public domain); skip world list.
- Commerce words: faker `en` lists (MIT). Check digits: implement from the rules above; GS1 prefix table: hand-write the few ranges needed, use `952`/`2xx` for generated barcodes; ISBN range file: do not embed.
- Books: Gutenberg `pg_catalog.csv` (CC0), filtered.
- Music: MusicBrainz genre names (CC0, 2 202), ID3v1 list (192), Wikidata instruments (CC0).
- Food: USDA FDC SR Legacy/Foundation descriptions (CC0, cite USDA), Wikidata dishes/cuisines (CC0), Open Brewery DB (MIT).
- Avoid: mledoze (ODbL), lukes (CC BY-SA), stefangabos (CC BY-SA/LGPL), Debian iso-codes (LGPL), SIL 639-3 table (no-redistribution clause), GS1/ISBN-International pages (all rights reserved), Wikipedia tables verbatim (CC BY-SA), MusicBrainz tag associations (CC BY-NC-SA).
+14 -11
View File
@@ -19,6 +19,14 @@ by a `data-import/` script, as README goal 10 asks.
- Give the address records one column set across countries: `region` and
`municipality` as columns on `geo.SE.address` too.
- Fill the 398 Swedish localities weighted 200 from SCB småorter.
- Give `url` and `email` a path that draws only domains nobody can register, keeping
the wide set as the default, per goal 12: 17 of the 40 distinct ones shipped today
sit on `.se`, `.nu`, `.io` and `.co`, which anyone may register, and only RFC 2606's
`example.com`, `.net`, `.org`, `.test`, `.example`, `.invalid` and `.localhost`
provably reach nothing.
- Audit the rest of the shipped set against goal 12 and give each a never-reaching
path where it lacks one: `phone` draws live PTS and NANP ranges, `bankgiro`,
`plusgiro` and `routing` draw live prefixes.
- Accept a middle name, and draw a shipped `personnummer` inside a *selected* sex;
both want a draw group sharing its family's pins. Stop the conflict error naming a
rewrite that returns a different value where the read it conflicts with sits inside
@@ -56,7 +64,7 @@ by a `data-import/` script, as README goal 10 asks.
| `lorem`, `hacker`, `hipster`, `catchphrase`, `buzzword`, `quote` | T/c | lorem ipsum, LLM-written | — |
| `direction`, `continent`, `ulid` | c/t | — | — |
`sv_SE` (source research: `research-sources-se.md`, in the handoff)
`sv_SE` ([source research](docs/research/research-sources-se.md))
| Category | Shape | Source | Licence |
|---|---|---|---|
@@ -73,7 +81,7 @@ by a `data-import/` script, as README goal 10 asks.
| month and weekday names, `holiday` | t+T | CLDR sv, lag 1989:253 | Unicode |
| `territory` localised names keyed by alpha2 | T | CLDR sv territory names | Unicode |
`en_US` (source research: `research-sources-world.md` part1–4, in the handoff)
`en_US` ([source research](docs/research/research-sources-world.md), and `part1`–`part4`)
| Category | Shape | Source | Licence |
|---|---|---|---|
@@ -109,20 +117,15 @@ by a `data-import/` script, as README goal 10 asks.
breaks a consumer, so it rides a 0.(x+1).0.
- Decide whether `parent: territory` stays, given a `--data-path` override of
`misc.territory` now fails `New` unless `misc.timezone` is overridden with it.
- Decide whether the shipped `url` and `email` domains are restricted to RFC 2606's
reserved names: 17 of the 40 distinct ones are not, sitting on `.se`, `.nu`, `.io`
and `.co`, which are live ccTLDs anyone can register. Restricting them costs the
locale flavour `example.se` and `.nu` were chosen for.
- Record why fejkdata ships no pop-culture catalogue. README Decisions states the
choice with no reason, because goal 10 does not supply one: it admits a hand-written
set where no register exists, and Wikidata, MusicBrainz and the Gutenberg catalog are
registers this plan already reads for `book`, `instrument` and `animal`.
### Release
- Move the README's Decisions section to `docs/decisions.md`, leaving a one-line index
of the titles in `AGENTS.md`, and drop the `AGENTS.md` line pointing decisions at the
README.
README. Carry the timezone weight's premise in with it: weighting by GeoNames city
population moved 300 seeded draws of `misc.territory[US].timezone` from 44 landing on
the four zones most Americans live in to 278, and dropped `America/Indiana/Petersburg`
(pop. 2,400) from 17 draws to 0.
- Rewrite `CHANGELOG.md`'s `[Unreleased]` as what v0.1.0 holds: nothing has shipped, so
"no longer paths", "where it used to fail" and "where it used to be an empty string"
describe versions no reader can have installed.