diff --git a/CHANGELOG.md b/CHANGELOG.md index 04238bc..de8b1c8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -53,9 +53,9 @@ replacement, and each removed path, column or flag. leading dot, `.se` and `.xn--p1ai`, the register's `type` beside it, and a `unicode` column holding the form the register displays, so `misc.tld[.рф]` selects the row `.xn--p1ai` renders. A delegation the register marks "Not assigned" does not ship. -- `misc.loglevel` is RFC 5424's severity table: the eight levels a syslog PRI encodes, - keyed by the numerical `code` and rendering the `severity` the RFC spells, so - `misc.loglevel[3]` and `misc.loglevel[Error]` select one row. +- `misc.loglevel` is the eight syslog severities: the `keyword` a configuration writes, + `err` and `info`, keyed by the numerical `code` a PRI encodes and selectable by + either, with the `severity` RFC 5424 spells beside it. - `geo.SE` and `geo.US`: five linked tables per country, `region`, `municipality`, `locality`, `postal-code` and `street`, weighted by population and address counts and built from SCB, GeoNames, Trafikverket NVDB and the US Census Bureau, and an diff --git a/DATA-LICENSES.md b/DATA-LICENSES.md index 26f4bfa..48491a9 100644 --- a/DATA-LICENSES.md +++ b/DATA-LICENSES.md @@ -19,7 +19,7 @@ Every shipped dataset, its source, its licence and the attribution it asks for. | `misc/httpmethod.tsv` | [IANA HTTP Method Registry](https://www.iana.org/assignments/http-methods/) | [public domain](https://www.iana.org/help/licensing-terms) | none required | `data-import/httpmethod.py` | | `misc/httpstatus.tsv` | [IANA HTTP Status Code Registry](https://www.iana.org/assignments/http-status-codes/) | [public domain](https://www.iana.org/help/licensing-terms) | none required | `data-import/httpstatus.py` | | `misc/language.tsv` | [datasets/language-codes](https://github.com/datasets/language-codes), the [Library of Congress](https://www.loc.gov/standards/iso639-2/) ISO 639-2 register | [PDDL 1.0](https://opendatacommons.org/licenses/pddl/1-0/) | none required | `data-import/language.py` | -| `misc/loglevel.tsv` | [RFC 5424](https://www.rfc-editor.org/rfc/rfc5424.txt) Table 2, the syslog message severities | facts | none required | `data-import/loglevel.py` | +| `misc/loglevel.tsv` | [RFC 5424](https://www.rfc-editor.org/rfc/rfc5424.txt) Table 2, the syslog message severities; keywords from [POSIX](https://pubs.opengroup.org/onlinepubs/9799919799/basedefs/syslog.h.html) `` | facts | none required | `data-import/loglevel.py` | | `misc/mimetype.tsv` | [mime-db](https://github.com/jshttp/mime-db), the IANA media type registry with filename extensions | [MIT](https://github.com/jshttp/mime-db/blob/master/LICENSE) | "Copyright (c) 2014 Jonathan Ong, Copyright (c) 2015-2022 Douglas Christopher Wilson" | `data-import/mimetype.py` | | `misc/port.tsv` | [IANA Service Name and Transport Protocol Port Number Registry](https://www.iana.org/assignments/service-names-port-numbers/) | [public domain](https://www.iana.org/help/licensing-terms) | none required | `data-import/port.py` | | `misc/protocol.tsv` | [IANA Protocol Numbers](https://www.iana.org/assignments/protocol-numbers/) | [public domain](https://www.iana.org/help/licensing-terms) | none required | `data-import/protocol.py` | diff --git a/README.md b/README.md index 94bf57e..e487bf6 100644 --- a/README.md +++ b/README.md @@ -163,9 +163,9 @@ Each locale carries `address`, `color`, `company`, `date`, `email`, `first-name` `personnummer` and `samordningsnummer`, `en_US` adds `ssn` and `itin`. `misc` carries `car`, `coordinate`, `creditcard` (Luhn-valid), `currency` (ISO 4217), `datetime` (RFC 3339), `emoji`, `httpmethod`, `httpstatus`, `language` (ISO 639-1, -with its 639-2/T code), `loglevel` (RFC 5424), `mac`, `mimetype`, `objectid`, `port`, -`protocol`, `territory` (ISO 3166-1), `timezone` (IANA), `tld` (IANA root zone), -`useragent` and `uuid` (v4). Many carry sub-fields — `misc.currency.symbol`, +with its 639-2/T code), `loglevel` (syslog severity), `mac`, `mimetype`, `objectid`, +`port`, `protocol`, `territory` (ISO 3166-1), `timezone` (IANA), `tld` (IANA root +zone), `useragent` and `uuid` (v4). Many carry sub-fields — `misc.currency.symbol`, `misc.territory.alpha2`, `misc.httpstatus.code` — which `--list` shows. `car`, `currency`, `httpmethod`, `httpstatus`, `language`, `loglevel`, `mimetype`, `port`, `protocol`, `territory`, `timezone`, `tld` and `useragent` are [tables](#table), so @@ -191,9 +191,10 @@ its `unicode` the form the register displays, `.рф` for `.xn--p1ai` and the ke for an ASCII one, which selects the row too, so `misc.tld[.рф]` renders `.xn--p1ai`. A draw spans the whole zone, `.arpa` and `.zuerich` alike, so pin `misc.tld[.com]` where a fixture needs one a reader recognises. -`misc.loglevel` is RFC 5424's eight severities, keyed by the numerical code a syslog -PRI encodes and named by the severity, so `misc.loglevel[3]` and `misc.loglevel[Error]` -are one row and `misc.loglevel.code` the number beside the name. +`misc.loglevel` is the eight syslog severities, rendering the keyword a configuration +writes, `err` and `info`, keyed by the numerical code a PRI encodes and selectable by +either, so `misc.loglevel[3]` and `misc.loglevel[err]` are one row. +`misc.loglevel.severity` is the name RFC 5424 spells, `Error` beside `err`. [`DATA-LICENSES.md`](DATA-LICENSES.md) names each table's source and licence. `misc.timezone` is every zone tzdb gives a shipped territory, from one apiece for most @@ -1233,10 +1234,12 @@ the Development section below, and who ships a register the four above then draw every value names a row instead. A territory the register records no state for stands alone, which is `EH` alone, and naming one for it would be a claim fejkdata has no business making. -- **`misc.loglevel` is a table keyed by the code, not a list of names.** A syslog - message carries the severity as the number its PRI encodes, so the code and the name - are one fact and goal 1 draws them together; a flat list of names would carry - neither the code nor a selector reaching it. +- **`misc.loglevel` is a table keyed by the code, rendering POSIX's keyword.** A flat + list of names carries neither the code a PRI encodes nor a selector reaching it, and + goal 1 draws the two as one fact. RFC 5424 spells the severity `Informational`, + which no syslog, PSR-3 or `slog` checker accepts, so the shipped spelling is the one + a configuration writes, `info`, and the RFC's name stays the `severity` column. The + code and the keyword take the two selector slots, so `misc.loglevel[Error]` misses. - **`misc.tld` keys carry the leading dot, where other tables key on a bare code.** The register spells a TLD `.se` and `misc.territory.tld` already ships it so, which a bare key would make two spellings of one fact; `{/misc.tld}` also composes onto a diff --git a/data-import/loglevel.py b/data-import/loglevel.py index ed8cafb..67f2242 100644 --- a/data-import/loglevel.py +++ b/data-import/loglevel.py @@ -1,11 +1,15 @@ #!/usr/bin/env python3 -"""Rebuild data/misc/loglevel.tsv from RFC 5424's severity table. +"""Rebuild data/misc/loglevel.tsv from RFC 5424 and POSIX . - data-import/loglevel.py [--source URL_OR_FILE] [--cache DIR] [--out FILE] + data-import/loglevel.py [--source URL_OR_FILE] [--posix URL_OR_FILE] [--cache DIR] [--out FILE] -Table 2 is the whole set, read between the facility table's caption and its own. +Table 2 is the whole set, read between the facility table's caption and its own. The +keyword is POSIX's severity macro less its `LOG_` prefix, lowercased, taken in the +order POSIX lists them, which `LOG_UPTO` states is the severity order; each must be a +prefix of the severity RFC 5424 numbers alike, so the two sources pin one pairing. """ import argparse +import html import re import sys from pathlib import Path @@ -14,33 +18,51 @@ import source import tsv SOURCE = "https://www.rfc-editor.org/rfc/rfc5424.txt" +POSIX = "https://pubs.opengroup.org/onlinepubs/9799919799/basedefs/syslog.h.html" OUT = Path(__file__).resolve().parent.parent / "data" / "misc" / "loglevel.tsv" CACHE = Path(__file__).resolve().parent / "cache" -COLUMNS = ["code", "severity"] +COLUMNS = ["code", "keyword", "severity"] TABLE = re.compile(r"Table 1\.\s+Syslog Message Facilities(.*?)Table 2\.\s+Syslog Message Severities", re.S) ROW = re.compile(r"^ +(\d) +([A-Z][a-z]+): ", re.M) +SEVERITIES = re.compile(r"severity level portion of the.*?macros:(.*?)The following shall be declared as functions", re.S) +MACRO = re.compile(r"\bLOG_([A-Z]+)\b") EXPECTED = 8 -def rows(text): +def severities(text): table = TABLE.search(text) if not table: sys.exit("the RFC holds no text between the facility and the severity caption; its tables have moved") - for code, severity in ROW.findall(table.group(1)): - yield {"code": code, "severity": severity} + return ROW.findall(table.group(1)) + + +def keywords(page): + section = SEVERITIES.search(html.unescape(re.sub(r"<[^>]+>", "", page))) + if not section: + sys.exit("POSIX spells no severity list between its own caption and the function declarations; the page has moved") + return [m.lower() for m in MACRO.findall(section.group(1))] def main(): p = argparse.ArgumentParser(description=__doc__.splitlines()[0]) p.add_argument("--cache", default=str(CACHE)) p.add_argument("--source", default=SOURCE) + p.add_argument("--posix", default=POSIX) p.add_argument("--out", default=str(OUT)) a = p.parse_args() - table = sorted(rows(source.fetch(a.source, a.cache, "rfc5424.txt").decode("utf-8")), key=lambda r: r["code"]) - codes = [r["code"] for r in table] + table = sorted(severities(source.fetch(a.source, a.cache, "rfc5424.txt").decode("utf-8"))) + codes = [code for code, _ in table] if codes != [str(i) for i in range(EXPECTED)]: sys.exit(f"the severity table numbers {codes}, not 0 through {EXPECTED - 1}") - tsv.write(a.out, COLUMNS, table) + words = keywords(source.fetch(a.posix, a.cache, "posix-syslog.h.html").decode("utf-8")) + if len(words) != EXPECTED: + sys.exit(f"POSIX lists {len(words)} severity macros, not {EXPECTED}: {words}") + rows = [] + for (code, severity), keyword in zip(table, words): + if not severity.lower().startswith(keyword): + sys.exit(f"severity {code} is {severity} in the RFC and {keyword} in POSIX, which is no prefix of it") + rows.append({"code": code, "keyword": keyword, "severity": severity}) + tsv.write(a.out, COLUMNS, rows) if __name__ == "__main__": diff --git a/data/misc/loglevel.json b/data/misc/loglevel.json index 9e9d3f1..1144d26 100644 --- a/data/misc/loglevel.json +++ b/data/misc/loglevel.json @@ -1 +1 @@ -{ "format": "{severity}", "rows": "loglevel.tsv", "key": "code", "name": "severity" } +{ "format": "{keyword}", "rows": "loglevel.tsv", "key": "code", "name": "keyword" } diff --git a/data/misc/loglevel.tsv b/data/misc/loglevel.tsv index 8ba335c..922c626 100644 --- a/data/misc/loglevel.tsv +++ b/data/misc/loglevel.tsv @@ -1,9 +1,9 @@ -code severity -0 Emergency -1 Alert -2 Critical -3 Error -4 Warning -5 Notice -6 Informational -7 Debug +code keyword severity +0 emerg Emergency +1 alert Alert +2 crit Critical +3 err Error +4 warning Warning +5 notice Notice +6 info Informational +7 debug Debug