Read misc.language, httpstatus and mimetype from their registers #22

Merged
lilleman merged 6 commits from misc-tables into main 2026-09-18 22:31:33 +02:00
3 changed files with 10 additions and 6 deletions
Showing only changes of commit 17d1703806 - Show all commits
+7 -2
View File
@@ -163,8 +163,8 @@ Each locale carries `address`, `color`, `company`, `date`, `email`, `first-name`
`personnummer` and `samordningsnummer`, `en_US` adds `ssn` and `itin`. `misc` `personnummer` and `samordningsnummer`, `en_US` adds `ssn` and `itin`. `misc`
carries `car`, `coordinate`, `country` (ISO 3166), `creditcard` (Luhn-valid), carries `car`, `coordinate`, `country` (ISO 3166), `creditcard` (Luhn-valid),
`currency` (ISO 4217), `datetime` (RFC 3339), `emoji`, `httpstatus`, `language` `currency` (ISO 4217), `datetime` (RFC 3339), `emoji`, `httpstatus`, `language`
(ISO 639-1, with its 639-2/T code), `mac`, `mimetype`, `objectid`, `timezone` (IANA), (ISO 639-1, with its 639-2/T code), `mac`, `mimetype`, `objectid`, `timezone`
`useragent` and `uuid` (v4). Many carry sub-fields — `misc.currency.symbol`, (IANA), `useragent` and `uuid` (v4). Many carry sub-fields — `misc.currency.symbol`,
`misc.country.alpha2`, `misc.httpstatus.code` — which `--list` shows. `country`, `misc.country.alpha2`, `misc.httpstatus.code` — which `--list` shows. `country`,
`currency`, `httpstatus`, `language` and `mimetype` are [tables](#table), so `currency`, `httpstatus`, `language` and `mimetype` are [tables](#table), so
`misc.country[SE].capital` and `misc.currency[Euro].symbol` select a row; `misc.country[SE].capital` and `misc.currency[Euro].symbol` select a row;
@@ -1130,6 +1130,11 @@ App developers writing tests and fixtures, in Go and at a shell:
and VED beside VES, both Venezuela's, of which a country row names one. Linking and VED beside VES, both Venezuela's, of which a country row names one. Linking
would trade the register for the link, and `misc.country.currency` already pairs a would trade the register for the link, and `misc.country.currency` already pairs a
country with its currency in one draw. country with its currency in one draw.
- **An extension may name two media types.** `.xml`, `.rtf`, `.sub`, `.mpp` and `.ac`
each name two rows of `misc.mimetype`. Separating them would mean dropping a
registered type, or naming one by an extension that is not its own — `.mpt` is
Project's template, not its document. Selecting such a name is an error listing
both keys; select by type instead.
- **The Swedish ids draw Skatteverket's test series.** A Luhn-valid personnummer - **The Swedish ids draw Skatteverket's test series.** A Luhn-valid personnummer
over a random birth number may be a living person's; 238 and 239 after any date over a random birth number may be a living person's; 238 and 239 after any date
are blocked from assignment, so the shipped `personnummer` and are blocked from assignment, so the shipped `personnummer` and
-1
View File
@@ -16,7 +16,6 @@ SOURCE = "https://raw.githubusercontent.com/jshttp/mime-db/master/db.json"
OUT = Path(__file__).resolve().parent.parent / "data" / "misc" / "mimetype.tsv" OUT = Path(__file__).resolve().parent.parent / "data" / "misc" / "mimetype.tsv"
CACHE = Path(__file__).resolve().parent / "cache" CACHE = Path(__file__).resolve().parent / "cache"
COLUMNS = ["ext", "type"] COLUMNS = ["ext", "type"]
# mime-db lists extensions in no particular order, so where the everyday one is not first, name it.
FIXUPS = {"application/mp4": "mp4s", "audio/mpeg": "mp3", "audio/ogg": "ogg", "video/quicktime": "mov"} FIXUPS = {"application/mp4": "mp4s", "audio/mpeg": "mp3", "audio/ogg": "ogg", "video/quicktime": "mov"}
+3 -3
View File
@@ -5,14 +5,14 @@ from pathlib import Path
def write(path, columns, rows): def write(path, columns, rows):
"""Write the rows as a TSV; the table needs a row, and every cell must be non-empty and free of tabs, newlines and braces.""" """Write the rows as a TSV; the table needs two rows, and every cell must be non-empty and free of tabs, newlines and braces."""
lines = ["\t".join(columns)] lines = ["\t".join(columns)]
for row in rows: for row in rows:
cells = [str(row[c]) for c in columns] cells = [str(row[c]) for c in columns]
if not all(cells) or any(re.search(r"[\t\n{}]", c) for c in cells): if not all(cells) or any(re.search(r"[\t\n{}]", c) for c in cells):
raise ValueError(f"{path}: a cell is empty or holds a tab, newline or brace: {row}") raise ValueError(f"{path}: a cell is empty or holds a tab, newline or brace: {row}")
lines.append("\t".join(cells)) lines.append("\t".join(cells))
if len(lines) < 2: if len(lines) < 3:
raise ValueError(f"{path}: has no rows below its header; a table is at least two rows") raise ValueError(f"{path}: has {len(lines) - 1} rows below its header; a table is at least two rows")
Path(path).write_text("\n".join(lines) + "\n", encoding="utf-8") Path(path).write_text("\n".join(lines) + "\n", encoding="utf-8")
print(f"{path}: {len(lines) - 1} rows", file=sys.stderr) print(f"{path}: {len(lines) - 1} rows", file=sys.stderr)