From fc92bf8bf56f5787d622fcec314d24347ad82857 Mon Sep 17 00:00:00 2001 From: Lilleman auf Larv Date: Mon, 24 Aug 2026 11:40:11 +0200 Subject: [PATCH] Corpus by contract kind: pin the fence and escaping rules, hold two questions --- corpus/README.md | 17 ++++-- corpus/canonical-form/ordered-list.json | 53 ------------------- corpus/canonical-form/ordered-list.md | 3 -- .../commonmark-subset}/blockquote.json | 0 .../commonmark-subset}/blockquote.md | 0 .../commonmark-subset}/bullet-list.json | 0 .../commonmark-subset}/bullet-list.md | 0 .../code-block-backtick-run-mid-line.json | 18 +++++++ .../code-block-backtick-run-mid-line.md | 3 ++ .../code-block-backtick-run.json | 3 ++ .../code-block-backtick-run.md | 2 +- .../commonmark-subset}/code-block.json | 0 .../commonmark-subset}/code-block.md | 0 .../commonmark-subset}/code-span.json | 0 .../commonmark-subset}/code-span.md | 0 .../commonmark-subset/document-empty.json | 4 ++ .../commonmark-subset/document-empty.md | 0 .../commonmark-subset}/em-strong-strike.json | 0 .../commonmark-subset}/em-strong-strike.md | 0 .../commonmark-subset}/hard-break.json | 0 .../commonmark-subset}/hard-break.md | 0 .../commonmark-subset}/heading.json | 0 .../commonmark-subset}/heading.md | 0 .../commonmark-subset}/link.json | 0 .../commonmark-subset}/link.md | 0 .../ordered-list-order.json | 0 .../commonmark-subset}/ordered-list-order.md | 0 .../commonmark-subset}/paragraph-empty.json | 0 .../commonmark-subset}/paragraph-empty.md | 0 .../commonmark-subset}/paragraph.json | 0 .../commonmark-subset}/paragraph.md | 0 .../commonmark-subset}/rule.json | 0 .../commonmark-subset}/rule.md | 0 .../commonmark-subset}/text-escaping.json | 9 ++++ .../commonmark-subset}/text-escaping.md | 2 + spec/flavour.md | 10 ++-- todo.md | 35 ++++++++---- 37 files changed, 84 insertions(+), 75 deletions(-) delete mode 100644 corpus/canonical-form/ordered-list.json delete mode 100644 corpus/canonical-form/ordered-list.md rename corpus/{canonical-form => round-trip/commonmark-subset}/blockquote.json (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/blockquote.md (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/bullet-list.json (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/bullet-list.md (100%) create mode 100644 corpus/round-trip/commonmark-subset/code-block-backtick-run-mid-line.json create mode 100644 corpus/round-trip/commonmark-subset/code-block-backtick-run-mid-line.md rename corpus/{canonical-form => round-trip/commonmark-subset}/code-block-backtick-run.json (78%) rename corpus/{canonical-form => round-trip/commonmark-subset}/code-block-backtick-run.md (58%) rename corpus/{canonical-form => round-trip/commonmark-subset}/code-block.json (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/code-block.md (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/code-span.json (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/code-span.md (100%) create mode 100644 corpus/round-trip/commonmark-subset/document-empty.json create mode 100644 corpus/round-trip/commonmark-subset/document-empty.md rename corpus/{canonical-form => round-trip/commonmark-subset}/em-strong-strike.json (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/em-strong-strike.md (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/hard-break.json (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/hard-break.md (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/heading.json (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/heading.md (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/link.json (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/link.md (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/ordered-list-order.json (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/ordered-list-order.md (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/paragraph-empty.json (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/paragraph-empty.md (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/paragraph.json (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/paragraph.md (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/rule.json (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/rule.md (100%) rename corpus/{canonical-form => round-trip/commonmark-subset}/text-escaping.json (81%) rename corpus/{canonical-form => round-trip/commonmark-subset}/text-escaping.md (79%) diff --git a/corpus/README.md b/corpus/README.md index 4122780..5510e21 100644 --- a/corpus/README.md +++ b/corpus/README.md @@ -1,8 +1,15 @@ # The corpus -`.json` is an ADF document; `.md` is the markdown `adfToMarkdown` must emit for it, -byte for byte including the trailing newline, and that `markdownToAdf` must read back to that same -document (AGENTS.md §2). Directories mirror `spec/flavour.md`'s sections. +One directory per contract kind: -The JSON is editor-normal — empty `attrs`, `marks` and `content` as the absent key, adjacent -identical-mark text nodes merged — two-space indent, keys sorted. +- `round-trip/` — `.json` + `.md`: the markdown `adfToMarkdown` must emit for that + document, byte for byte, and that `markdownToAdf` must read back to it (AGENTS.md §2). Grouped + by node family: `commonmark-subset/`, `block-nodes/`, `inline-nodes/`, `opaque-carry/`, + `combinations/`. +- `normalization/` — `.md` + `.json`: markdown input, and the document + `markdownToAdf` must build from it. One-way; the markdown is not canonical. +- `errors/` — `.md` + `.error`: markdown input that must not convert. +- `real-payloads/` — `.json`: sanitized live ADF, round-tripped ADF→markdown→ADF. No + expected markdown. + +JSON is editor-normal (AGENTS.md §2), two-space indent, keys sorted. diff --git a/corpus/canonical-form/ordered-list.json b/corpus/canonical-form/ordered-list.json deleted file mode 100644 index a39528a..0000000 --- a/corpus/canonical-form/ordered-list.json +++ /dev/null @@ -1,53 +0,0 @@ -{ - "content": [ - { - "content": [ - { - "content": [ - { - "content": [ - { - "text": "Bolt M8", - "type": "text" - } - ], - "type": "paragraph" - } - ], - "type": "listItem" - }, - { - "content": [ - { - "content": [ - { - "text": "Nut M8", - "type": "text" - } - ], - "type": "paragraph" - } - ], - "type": "listItem" - }, - { - "content": [ - { - "content": [ - { - "text": "Washer M8", - "type": "text" - } - ], - "type": "paragraph" - } - ], - "type": "listItem" - } - ], - "type": "orderedList" - } - ], - "type": "doc", - "version": 1 -} diff --git a/corpus/canonical-form/ordered-list.md b/corpus/canonical-form/ordered-list.md deleted file mode 100644 index 3e512e7..0000000 --- a/corpus/canonical-form/ordered-list.md +++ /dev/null @@ -1,3 +0,0 @@ -1. Bolt M8 -2. Nut M8 -3. Washer M8 diff --git a/corpus/canonical-form/blockquote.json b/corpus/round-trip/commonmark-subset/blockquote.json similarity index 100% rename from corpus/canonical-form/blockquote.json rename to corpus/round-trip/commonmark-subset/blockquote.json diff --git a/corpus/canonical-form/blockquote.md b/corpus/round-trip/commonmark-subset/blockquote.md similarity index 100% rename from corpus/canonical-form/blockquote.md rename to corpus/round-trip/commonmark-subset/blockquote.md diff --git a/corpus/canonical-form/bullet-list.json b/corpus/round-trip/commonmark-subset/bullet-list.json similarity index 100% rename from corpus/canonical-form/bullet-list.json rename to corpus/round-trip/commonmark-subset/bullet-list.json diff --git a/corpus/canonical-form/bullet-list.md b/corpus/round-trip/commonmark-subset/bullet-list.md similarity index 100% rename from corpus/canonical-form/bullet-list.md rename to corpus/round-trip/commonmark-subset/bullet-list.md diff --git a/corpus/round-trip/commonmark-subset/code-block-backtick-run-mid-line.json b/corpus/round-trip/commonmark-subset/code-block-backtick-run-mid-line.json new file mode 100644 index 0000000..82391ee --- /dev/null +++ b/corpus/round-trip/commonmark-subset/code-block-backtick-run-mid-line.json @@ -0,0 +1,18 @@ +{ + "content": [ + { + "attrs": { + "language": "bash" + }, + "content": [ + { + "text": "echo \"```\" >> notes.md", + "type": "text" + } + ], + "type": "codeBlock" + } + ], + "type": "doc", + "version": 1 +} diff --git a/corpus/round-trip/commonmark-subset/code-block-backtick-run-mid-line.md b/corpus/round-trip/commonmark-subset/code-block-backtick-run-mid-line.md new file mode 100644 index 0000000..cd3a3b3 --- /dev/null +++ b/corpus/round-trip/commonmark-subset/code-block-backtick-run-mid-line.md @@ -0,0 +1,3 @@ +````bash +echo "```" >> notes.md +```` diff --git a/corpus/canonical-form/code-block-backtick-run.json b/corpus/round-trip/commonmark-subset/code-block-backtick-run.json similarity index 78% rename from corpus/canonical-form/code-block-backtick-run.json rename to corpus/round-trip/commonmark-subset/code-block-backtick-run.json index 012ffd4..a6452fb 100644 --- a/corpus/canonical-form/code-block-backtick-run.json +++ b/corpus/round-trip/commonmark-subset/code-block-backtick-run.json @@ -1,6 +1,9 @@ { "content": [ { + "attrs": { + "language": "markdown" + }, "content": [ { "text": "```\ncode\n```", diff --git a/corpus/canonical-form/code-block-backtick-run.md b/corpus/round-trip/commonmark-subset/code-block-backtick-run.md similarity index 58% rename from corpus/canonical-form/code-block-backtick-run.md rename to corpus/round-trip/commonmark-subset/code-block-backtick-run.md index 436503d..721c874 100644 --- a/corpus/canonical-form/code-block-backtick-run.md +++ b/corpus/round-trip/commonmark-subset/code-block-backtick-run.md @@ -1,4 +1,4 @@ -```` +````markdown ``` code ``` diff --git a/corpus/canonical-form/code-block.json b/corpus/round-trip/commonmark-subset/code-block.json similarity index 100% rename from corpus/canonical-form/code-block.json rename to corpus/round-trip/commonmark-subset/code-block.json diff --git a/corpus/canonical-form/code-block.md b/corpus/round-trip/commonmark-subset/code-block.md similarity index 100% rename from corpus/canonical-form/code-block.md rename to corpus/round-trip/commonmark-subset/code-block.md diff --git a/corpus/canonical-form/code-span.json b/corpus/round-trip/commonmark-subset/code-span.json similarity index 100% rename from corpus/canonical-form/code-span.json rename to corpus/round-trip/commonmark-subset/code-span.json diff --git a/corpus/canonical-form/code-span.md b/corpus/round-trip/commonmark-subset/code-span.md similarity index 100% rename from corpus/canonical-form/code-span.md rename to corpus/round-trip/commonmark-subset/code-span.md diff --git a/corpus/round-trip/commonmark-subset/document-empty.json b/corpus/round-trip/commonmark-subset/document-empty.json new file mode 100644 index 0000000..6591ef5 --- /dev/null +++ b/corpus/round-trip/commonmark-subset/document-empty.json @@ -0,0 +1,4 @@ +{ + "type": "doc", + "version": 1 +} diff --git a/corpus/round-trip/commonmark-subset/document-empty.md b/corpus/round-trip/commonmark-subset/document-empty.md new file mode 100644 index 0000000..e69de29 diff --git a/corpus/canonical-form/em-strong-strike.json b/corpus/round-trip/commonmark-subset/em-strong-strike.json similarity index 100% rename from corpus/canonical-form/em-strong-strike.json rename to corpus/round-trip/commonmark-subset/em-strong-strike.json diff --git a/corpus/canonical-form/em-strong-strike.md b/corpus/round-trip/commonmark-subset/em-strong-strike.md similarity index 100% rename from corpus/canonical-form/em-strong-strike.md rename to corpus/round-trip/commonmark-subset/em-strong-strike.md diff --git a/corpus/canonical-form/hard-break.json b/corpus/round-trip/commonmark-subset/hard-break.json similarity index 100% rename from corpus/canonical-form/hard-break.json rename to corpus/round-trip/commonmark-subset/hard-break.json diff --git a/corpus/canonical-form/hard-break.md b/corpus/round-trip/commonmark-subset/hard-break.md similarity index 100% rename from corpus/canonical-form/hard-break.md rename to corpus/round-trip/commonmark-subset/hard-break.md diff --git a/corpus/canonical-form/heading.json b/corpus/round-trip/commonmark-subset/heading.json similarity index 100% rename from corpus/canonical-form/heading.json rename to corpus/round-trip/commonmark-subset/heading.json diff --git a/corpus/canonical-form/heading.md b/corpus/round-trip/commonmark-subset/heading.md similarity index 100% rename from corpus/canonical-form/heading.md rename to corpus/round-trip/commonmark-subset/heading.md diff --git a/corpus/canonical-form/link.json b/corpus/round-trip/commonmark-subset/link.json similarity index 100% rename from corpus/canonical-form/link.json rename to corpus/round-trip/commonmark-subset/link.json diff --git a/corpus/canonical-form/link.md b/corpus/round-trip/commonmark-subset/link.md similarity index 100% rename from corpus/canonical-form/link.md rename to corpus/round-trip/commonmark-subset/link.md diff --git a/corpus/canonical-form/ordered-list-order.json b/corpus/round-trip/commonmark-subset/ordered-list-order.json similarity index 100% rename from corpus/canonical-form/ordered-list-order.json rename to corpus/round-trip/commonmark-subset/ordered-list-order.json diff --git a/corpus/canonical-form/ordered-list-order.md b/corpus/round-trip/commonmark-subset/ordered-list-order.md similarity index 100% rename from corpus/canonical-form/ordered-list-order.md rename to corpus/round-trip/commonmark-subset/ordered-list-order.md diff --git a/corpus/canonical-form/paragraph-empty.json b/corpus/round-trip/commonmark-subset/paragraph-empty.json similarity index 100% rename from corpus/canonical-form/paragraph-empty.json rename to corpus/round-trip/commonmark-subset/paragraph-empty.json diff --git a/corpus/canonical-form/paragraph-empty.md b/corpus/round-trip/commonmark-subset/paragraph-empty.md similarity index 100% rename from corpus/canonical-form/paragraph-empty.md rename to corpus/round-trip/commonmark-subset/paragraph-empty.md diff --git a/corpus/canonical-form/paragraph.json b/corpus/round-trip/commonmark-subset/paragraph.json similarity index 100% rename from corpus/canonical-form/paragraph.json rename to corpus/round-trip/commonmark-subset/paragraph.json diff --git a/corpus/canonical-form/paragraph.md b/corpus/round-trip/commonmark-subset/paragraph.md similarity index 100% rename from corpus/canonical-form/paragraph.md rename to corpus/round-trip/commonmark-subset/paragraph.md diff --git a/corpus/canonical-form/rule.json b/corpus/round-trip/commonmark-subset/rule.json similarity index 100% rename from corpus/canonical-form/rule.json rename to corpus/round-trip/commonmark-subset/rule.json diff --git a/corpus/canonical-form/rule.md b/corpus/round-trip/commonmark-subset/rule.md similarity index 100% rename from corpus/canonical-form/rule.md rename to corpus/round-trip/commonmark-subset/rule.md diff --git a/corpus/canonical-form/text-escaping.json b/corpus/round-trip/commonmark-subset/text-escaping.json similarity index 81% rename from corpus/canonical-form/text-escaping.json rename to corpus/round-trip/commonmark-subset/text-escaping.json index 1465d19..5481877 100644 --- a/corpus/canonical-form/text-escaping.json +++ b/corpus/round-trip/commonmark-subset/text-escaping.json @@ -35,6 +35,15 @@ } ], "type": "paragraph" + }, + { + "content": [ + { + "text": "*not emphasis*", + "type": "text" + } + ], + "type": "paragraph" } ], "type": "doc", diff --git a/corpus/canonical-form/text-escaping.md b/corpus/round-trip/commonmark-subset/text-escaping.md similarity index 79% rename from corpus/canonical-form/text-escaping.md rename to corpus/round-trip/commonmark-subset/text-escaping.md index de71c9c..c30602b 100644 --- a/corpus/canonical-form/text-escaping.md +++ b/corpus/round-trip/commonmark-subset/text-escaping.md @@ -5,3 +5,5 @@ 2 * 3 * 4 = 24 snake_case_name + +\*not emphasis* diff --git a/spec/flavour.md b/spec/flavour.md index 0a92971..72ddebd 100644 --- a/spec/flavour.md +++ b/spec/flavour.md @@ -23,7 +23,8 @@ normalizes to it through the round-trip. - Blockquotes prefix lines with `> `; a blank line inside a blockquote is a bare `>`. - ATX headings (`#` … `######`); setext input normalizes to ATX. - Code fences ``` with the node's language as info string, the fence lengthened past any backtick - run in the content; indented-code input normalizes to fences. + run in the content — the longest run anywhere plus one, counting mid-line runs no closing fence + could match; indented-code input normalizes to fences. - Code spans: a backtick string one longer than the longest backtick run in the text, the text padded with one space on each side where it begins or ends with a backtick, or begins and ends with a space without being all spaces. The content is literal — inline parsing does not see @@ -37,8 +38,11 @@ normalizes to it through the round-trip. CommonMark autolink (absolute URI). - Paragraphs on one line — no soft wrapping; a soft line break in input becomes a single space. - Entity references in input decode to their characters; output backslash-escapes only where text - would otherwise parse as syntax. -- Blocks separated by one blank line, no trailing whitespace, single trailing newline. + would otherwise parse as syntax: escape the leading delimiter of a construct that would otherwise + open, re-scan from there, and repeat — with the opener literal the closer parses as text, so + `*not emphasis*` is `\*not emphasis*`, one backslash. +- Blocks separated by one blank line, no trailing whitespace, single trailing newline; a document + with no blocks is the empty string. ## Directives diff --git a/todo.md b/todo.md index 1e57210..6d7fd5b 100644 --- a/todo.md +++ b/todo.md @@ -20,11 +20,18 @@ detail is settled at its own milestone. escape-based, never literal, since pipe cells trim and pad. At `mediaInline`, check real payloads for external-URL support — if it exists, revisit the media section's mid-text-image error and its "no slot" ground. -- [ ] **1d — Corpus start** (§10): checked-in ADF ↔ canonical-markdown fixture pairs per spec'd - node, in `corpus/`, one directory per `spec/flavour.md` section. - - [ ] **1d1 — Canonical form**: the plain-CommonMark subset — blockquote, bulletList, - codeBlock, heading, orderedList, paragraph, rule, listItem, hardBreak, text, code spans, - and the `code`, `em`, `link`, `strike` and `strong` marks. +- [ ] **1d — Corpus start** (§10): checked-in fixtures per spec'd node, in `corpus/`, one + directory per contract kind (`corpus/README.md`). + **Blocked on the maintainer** (§15), not to be guessed: Canonical form has no totality + guard — a `paragraph`, `heading`, `blockquote` or `codeBlock` carrying a `localId`, or a + `blockquote` carrying marks, has no spelling that keeps it, and picking one (directive + sections for the six CommonMark block nodes, or the opaque carry) is a permanent format + decision (§8). Its three collision sites stay out of the corpus until then: an `orderedList` + starting at 1, a `codeBlock` whose info string is empty, and `media` with an empty `alt` — + each a choice between the absent attribute and the empty value. + - [ ] **1d1 — The CommonMark subset**: blockquote, bulletList, codeBlock, heading, orderedList, + paragraph, rule, listItem, hardBreak, text, code spans, and the `code`, `em`, `link`, + `strike` and `strong` marks — one mark per text node; nesting is 1d5's. - [ ] **1d2 — Block nodes**: panel, expand/nestedExpand, the media family and the CommonMark image shape, both table forms, task and decision lists, layout, extensions, syncBlock — with the reserved `marks` attribute and the fence lengths nesting forces. @@ -38,14 +45,22 @@ detail is settled at its own milestone. combining nodes rather than isolating one. - [ ] **1d6 — Input normalization**: one-way markdown→ADF fixtures, not pairs — setext headings, indented code, loose lists, `*`/`+` bullets, entity references, soft wraps. - - [ ] **1d7 — Error input**: also one-way, a markdown input per named error. Waits on milestone - 3 naming them; 1d's pairs are valid documents only. -- [ ] **2 — `adfToMarkdown`.** First real code — decide here where §10's coverage check lives. + - [ ] **1d7 — Error input**: also one-way, a markdown input per named error, asserting only + that conversion fails — malformed directives, the image gap, a claimed pipe-table line + that does not parse, the content slot, raw HTML with no mapping. Which error each returns + is pinned at milestone 3, where they are named. +- [ ] **2 — `adfToMarkdown`.** First real code — decide here where §10's coverage check lives, and + gate that every `corpus/**/*.json` re-serializes to itself under the library's own canonical + serializer: one implementation, keys sorted, two spellings — two-space indent for the corpus + files and the block carry's body, compact for the inline carry. - [ ] **3 — `markdownToAdf`.** The CommonMark parser is the largest single component. The raw-HTML element mapping is empty until milestone 6, so at `0.1.0` every raw-HTML construct in input - is an error result. + is an error result. The CommonMark spec suite runs against it from here (§10), and + `corpus/errors/` gains the error each fixture must return (1d7). - [ ] **4 — Round-trip property tests** over the corpus, both ways — the thing that proves 2 and - 3. Generators emit editor-normal ADF (§2). + 3. Generators emit editor-normal ADF (§2). Real sanitized ADF from live Atlassian APIs lands + here too (§10), in `corpus/real-payloads/`: an ADF→markdown→ADF check with no expected + markdown, the payloads supplied by the maintainer. - [ ] **5 — Release pipeline, ship `0.1.0`.** Publish-on-version-change (§9), `NPM_TOKEN` secret, the repo made public first (§6). `0.1.0` is the markdown round-trip: both markdown directions, the types, `isAdfDocument`. The build lands here: a build tsconfig emitting JS