Merge pull request 'Corpus 1d1: the CommonMark subset pairs, and the 1d split' (#6) from corpus-start into main
CI / gate (push) Successful in 4s

This commit was merged in pull request #6.
This commit is contained in:
2026-08-24 12:54:01 +02:00
37 changed files with 985 additions and 8 deletions
+15
View File
@@ -0,0 +1,15 @@
# The corpus
One directory per contract kind:
- `round-trip/``<name>.json` + `<name>.md`: the markdown `adfToMarkdown` must emit for that
document, byte for byte, and that `markdownToAdf` must read back to it (AGENTS.md §2). Grouped
by node family.
- `normalization/``<name>.md` + `<name>.json`: markdown input, and the document
`markdownToAdf` must build from it. One-way; the markdown is not canonical.
- `errors/``<name>.md`: markdown input that must not convert. A `<name>.error` beside it
pins which error.
- `real-payloads/``<name>.json`: sanitized live ADF, round-tripped ADF→markdown→ADF. No
expected markdown.
JSON is editor-normal (AGENTS.md §2), two-space indent, keys sorted.
@@ -0,0 +1,29 @@
{
"content": [
{
"content": [
{
"content": [
{
"text": "Ship it.",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "Then tell them.",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "blockquote"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,3 @@
> Ship it.
>
> Then tell them.
@@ -0,0 +1,72 @@
{
"content": [
{
"content": [
{
"content": [
{
"content": [
{
"text": "Bolt M8",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "listItem"
},
{
"content": [
{
"content": [
{
"text": "Nut M8",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "listItem"
},
{
"content": [
{
"content": [
{
"text": "Washer M8",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"content": [
{
"content": [
{
"text": "Fibre",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "listItem"
}
],
"type": "bulletList"
}
],
"type": "listItem"
}
],
"type": "bulletList"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,4 @@
- Bolt M8
- Nut M8
- Washer M8
- Fibre
@@ -0,0 +1,18 @@
{
"content": [
{
"attrs": {
"language": "bash"
},
"content": [
{
"text": "echo \"```\" >> notes.md",
"type": "text"
}
],
"type": "codeBlock"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,3 @@
````bash
echo "```" >> notes.md
````
@@ -0,0 +1,18 @@
{
"content": [
{
"attrs": {
"language": "markdown"
},
"content": [
{
"text": "```\ncode\n```",
"type": "text"
}
],
"type": "codeBlock"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,5 @@
````markdown
```
code
```
````
@@ -0,0 +1,18 @@
{
"content": [
{
"attrs": {
"language": "markdown"
},
"content": [
{
"text": "Line one \nLine two",
"type": "text"
}
],
"type": "codeBlock"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,4 @@
```markdown
Line one
Line two
```
@@ -0,0 +1,18 @@
{
"content": [
{
"attrs": {
"language": "sql"
},
"content": [
{
"text": "SELECT id\nFROM part\nWHERE qty > 0;",
"type": "text"
}
],
"type": "codeBlock"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,5 @@
```sql
SELECT id
FROM part
WHERE qty > 0;
```
@@ -0,0 +1,118 @@
{
"content": [
{
"content": [
{
"text": "Run ",
"type": "text"
},
{
"marks": [
{
"type": "code"
}
],
"text": "npm ci",
"type": "text"
},
{
"text": " first.",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "Nested backticks: ",
"type": "text"
},
{
"marks": [
{
"type": "code"
}
],
"text": "run `date` twice",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "A span holding ",
"type": "text"
},
{
"marks": [
{
"type": "code"
}
],
"text": "`code`",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "Held: ",
"type": "text"
},
{
"marks": [
{
"type": "code"
}
],
"text": " spaced ",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "Three: ",
"type": "text"
},
{
"marks": [
{
"type": "code"
}
],
"text": " ",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "Literal: ",
"type": "text"
},
{
"marks": [
{
"type": "code"
}
],
"text": "~~not strike~~",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,11 @@
Run `npm ci` first.
Nested backticks: ``run `date` twice``
A span holding `` `code` ``
Held: ` spaced `
Three: ` `
Literal: `~~not strike~~`
@@ -0,0 +1,4 @@
{
"type": "doc",
"version": 1
}
@@ -0,0 +1,76 @@
{
"content": [
{
"content": [
{
"text": "Read the ",
"type": "text"
},
{
"marks": [
{
"type": "em"
}
],
"text": "manual",
"type": "text"
},
{
"text": " before ",
"type": "text"
},
{
"marks": [
{
"type": "strong"
}
],
"text": "wiring",
"type": "text"
},
{
"text": " the ",
"type": "text"
},
{
"marks": [
{
"type": "strike"
}
],
"text": "relay",
"type": "text"
},
{
"text": ".",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "An un",
"type": "text"
},
{
"marks": [
{
"type": "em"
}
],
"text": "real",
"type": "text"
},
{
"text": "istic goal.",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,3 @@
Read the _manual_ before **wiring** the ~~relay~~.
An un*real*istic goal.
@@ -0,0 +1,53 @@
{
"content": [
{
"content": [
{
"text": "Line one",
"type": "text"
},
{
"type": "hardBreak"
},
{
"text": "Line two",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "Trailing",
"type": "text"
},
{
"type": "hardBreak"
}
],
"type": "paragraph"
},
{
"attrs": {
"level": 2
},
"content": [
{
"text": "Two",
"type": "text"
},
{
"type": "hardBreak"
},
{
"text": "lines",
"type": "text"
}
],
"type": "heading"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,6 @@
Line one\
Line two
Trailing:hardBreak{}
## Two:hardBreak{}lines
@@ -0,0 +1,78 @@
{
"content": [
{
"attrs": {
"level": 1
},
"content": [
{
"text": "Assembly",
"type": "text"
}
],
"type": "heading"
},
{
"attrs": {
"level": 2
},
"content": [
{
"text": "Parts",
"type": "text"
}
],
"type": "heading"
},
{
"attrs": {
"level": 3
},
"content": [
{
"text": "Fasteners",
"type": "text"
}
],
"type": "heading"
},
{
"attrs": {
"level": 4
},
"content": [
{
"text": "Bolts",
"type": "text"
}
],
"type": "heading"
},
{
"attrs": {
"level": 5
},
"content": [
{
"text": "Sizes",
"type": "text"
}
],
"type": "heading"
},
{
"attrs": {
"level": 6
},
"content": [
{
"text": "M8",
"type": "text"
}
],
"type": "heading"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,11 @@
# Assembly
## Parts
### Fasteners
#### Bolts
##### Sizes
###### M8
@@ -0,0 +1,116 @@
{
"content": [
{
"content": [
{
"text": "See ",
"type": "text"
},
{
"marks": [
{
"attrs": {
"href": "https://example.com/changelog"
},
"type": "link"
}
],
"text": "the changelog",
"type": "text"
},
{
"text": ".",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"marks": [
{
"attrs": {
"href": "https://example.com/"
},
"type": "link"
}
],
"text": "https://example.com/",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "Read ",
"type": "text"
},
{
"marks": [
{
"attrs": {
"href": "https://example.com/guide",
"title": "Setup guide"
},
"type": "link"
}
],
"text": "the guide",
"type": "text"
},
{
"text": ".",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "Open ",
"type": "text"
},
{
"marks": [
{
"attrs": {
"href": "https://example.com/the plan.pdf"
},
"type": "link"
}
],
"text": "the plan",
"type": "text"
},
{
"text": ".",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"marks": [
{
"attrs": {
"href": "/parts/m8"
},
"type": "link"
}
],
"text": "/parts/m8",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,9 @@
See [the changelog](https://example.com/changelog).
<https://example.com/>
Read [the guide](https://example.com/guide "Setup guide").
Open [the plan](<https://example.com/the plan.pdf>).
[/parts/m8](/parts/m8)
@@ -0,0 +1,74 @@
{
"content": [
{
"attrs": {
"order": 9
},
"content": [
{
"content": [
{
"content": [
{
"text": "Bolt M8",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "Grade 8.8.",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "listItem"
},
{
"content": [
{
"content": [
{
"text": "Nut M8",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "Zinc-plated.",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "listItem"
},
{
"content": [
{
"content": [
{
"text": "Washer M8",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "listItem"
}
],
"type": "orderedList"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,7 @@
9. Bolt M8
Grade 8.8.
10. Nut M8
Zinc-plated.
11. Washer M8
@@ -0,0 +1,27 @@
{
"content": [
{
"content": [
{
"text": "First.",
"type": "text"
}
],
"type": "paragraph"
},
{
"type": "paragraph"
},
{
"content": [
{
"text": "Second.",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,5 @@
First.
::paragraph
Second.
@@ -0,0 +1,24 @@
{
"content": [
{
"content": [
{
"text": "The converter emits every paragraph on a single line, however long it runs, because a soft line break in input carries no meaning ADF can hold.",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "Blocks are separated by exactly one blank line.",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,3 @@
The converter emits every paragraph on a single line, however long it runs, because a soft line break in input carries no meaning ADF can hold.
Blocks are separated by exactly one blank line.
@@ -0,0 +1,27 @@
{
"content": [
{
"content": [
{
"text": "Before.",
"type": "text"
}
],
"type": "paragraph"
},
{
"type": "rule"
},
{
"content": [
{
"text": "After.",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,5 @@
Before.
---
After.
@@ -0,0 +1,51 @@
{
"content": [
{
"content": [
{
"text": "# Not a heading",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "- not a bullet",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "2 * 3 * 4 = 24",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "snake_case_name",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"text": "*not emphasis*",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,9 @@
\# Not a heading
\- not a bullet
2 * 3 * 4 = 24
snake_case_name
\*not emphasis*
+8 -3
View File
@@ -23,7 +23,8 @@ normalizes to it through the round-trip.
- Blockquotes prefix lines with `> `; a blank line inside a blockquote is a bare `>`. - Blockquotes prefix lines with `> `; a blank line inside a blockquote is a bare `>`.
- ATX headings (`#``######`); setext input normalizes to ATX. - ATX headings (`#``######`); setext input normalizes to ATX.
- Code fences ``` with the node's language as info string, the fence lengthened past any backtick - Code fences ``` with the node's language as info string, the fence lengthened past any backtick
run in the content; indented-code input normalizes to fences. run in the content — the longest run anywhere plus one, at least three, counting mid-line runs
no closing fence could match; indented-code input normalizes to fences.
- Code spans: a backtick string one longer than the longest backtick run in the text, the text - Code spans: a backtick string one longer than the longest backtick run in the text, the text
padded with one space on each side where it begins or ends with a backtick, or begins and ends padded with one space on each side where it begins or ends with a backtick, or begins and ends
with a space without being all spaces. The content is literal — inline parsing does not see with a space without being all spaces. The content is literal — inline parsing does not see
@@ -37,8 +38,12 @@ normalizes to it through the round-trip.
CommonMark autolink (absolute URI). CommonMark autolink (absolute URI).
- Paragraphs on one line — no soft wrapping; a soft line break in input becomes a single space. - Paragraphs on one line — no soft wrapping; a soft line break in input becomes a single space.
- Entity references in input decode to their characters; output backslash-escapes only where text - Entity references in input decode to their characters; output backslash-escapes only where text
would otherwise parse as syntax. would otherwise parse as syntax, scanning the assembled line rather than each text node: escape
- Blocks separated by one blank line, no trailing whitespace, single trailing newline. the leading delimiter of a construct that would otherwise open, re-scan from there, and repeat —
with the opener literal the closer parses as text, so `*not emphasis*` is `\*not emphasis*`, one
backslash.
- Blocks separated by one blank line, no trailing whitespace outside a code block's content,
single trailing newline; a document with no blocks is the empty string.
## Directives ## Directives
+48 -5
View File
@@ -20,14 +20,57 @@ detail is settled at its own milestone.
escape-based, never literal, since pipe cells trim and pad. At `mediaInline`, check real escape-based, never literal, since pipe cells trim and pad. At `mediaInline`, check real
payloads for external-URL support — if it exists, revisit the media section's payloads for external-URL support — if it exists, revisit the media section's
mid-text-image error and its "no slot" ground. mid-text-image error and its "no slot" ground.
- [ ] **1d — Corpus start** (§10): checked-in ADF ↔ canonical-markdown fixture pairs per spec'd - [ ] **1d — Corpus start** (§10): checked-in fixtures per spec'd node, in `corpus/`, one
node. directory per contract kind (`corpus/README.md`).
- [ ] **2 — `adfToMarkdown`.** First real code — decide here where §10's coverage check lives. **Blocked on the maintainer** (§15), not to be guessed: Canonical form has no totality
guard. Per `@atlaskit/adf-schema` 57.1.0 every block node it spells — `blockquote`,
`bulletList`, `codeBlock`, `heading`, `listItem`, `orderedList`, `paragraph`, `rule`
carries a `localId` with no spelling, `codeBlock` also `hideLineNumbers`, `uniqueId` and
`wrap`, `blockquote` also marks, and `hardBreak` `text` and `localId` with no section for
the carry fallback to reach. Picking one (directive sections for those nodes, or the opaque
carry) is a permanent format decision (§8). Three collision sites are held out of the corpus
meanwhile, each a choice between the absent attribute and the empty value: a `codeBlock`
whose info string is empty and `media` with an empty `alt`, which one "exactly that shape"
rule — as the CommonMark image already uses — could settle together, and an `orderedList`
starting at 1, independent of the totality answer since `order: 9` keeps the markdown form
either way. **Also blocked**: the link rule covers destination spaces only, so two shapes
break §2 silently — href `https://example.com/a)b` emits `[t](https://example.com/a)b)`,
read back as href `…/a` plus literal `b)`; title `He said "hi"` emits
`[t](u "He said "hi"")`, which holds no title. Two defensible spellings each — angle
brackets or a backslash escape, and for titles `'…'` or `(…)` besides — so §8 leaves the
pick here.
- [ ] **1d1 — The CommonMark subset**: blockquote, bulletList, codeBlock, heading, orderedList,
paragraph, rule, listItem, hardBreak, text, code spans, and the `code`, `em`, `link`,
`strike` and `strong` marks — one mark per text node; nesting is 1d5's.
- [ ] **1d2 — Block nodes**: panel, expand/nestedExpand, the media family and the CommonMark
image shape, both table forms, task and decision lists, layout, extensions, syncBlock —
with the reserved `marks` attribute and the fence lengths nesting forces.
- [ ] **1d3 — Inline nodes and marks**: date, emoji, inlineCard, mediaInline, mention, status;
border, subsup, textColor, underline; the content slot's `text` attribute and the
`:text{text="…"}` whitespace spelling.
- [ ] **1d4 — Opaque carry** (§3): an unknown node in both positions, the reserved `adf` info
string, and the `codeBlock` whose language is `adf`.
- [ ] **1d5 — Carve-outs and combinations**: the three carve-outs and their escapes, mark
nesting order and the runs a carry breaks, attribute canonicalization, and documents
combining nodes rather than isolating one.
- [ ] **1d6 — Input normalization**: one-way markdown→ADF fixtures, not pairs — setext
headings, indented code, loose lists, `*`/`+` bullets, entity references, soft wraps.
- [ ] **1d7 — Error input**: also one-way, a markdown input per named error, asserting only
that conversion fails — malformed directives, the image gap, a claimed pipe-table line
that does not parse, the content slot, raw HTML with no mapping. Which error each returns
is pinned at milestone 3, where they are named.
- [ ] **2 — `adfToMarkdown`.** First real code — decide here where §10's coverage check lives, and
gate that every `corpus/**/*.json` re-serializes to itself under the library's own canonical
serializer: one implementation, keys sorted, two spellings — two-space indent for the corpus
files and the block carry's body, compact for the inline carry.
- [ ] **3 — `markdownToAdf`.** The CommonMark parser is the largest single component. The raw-HTML - [ ] **3 — `markdownToAdf`.** The CommonMark parser is the largest single component. The raw-HTML
element mapping is empty until milestone 6, so at `0.1.0` every raw-HTML construct in input element mapping is empty until milestone 6, so at `0.1.0` every raw-HTML construct in input
is an error result. is an error result. The CommonMark spec suite runs against it from here (§10), and
`corpus/errors/` gains the error each fixture must return (1d7).
- [ ] **4 — Round-trip property tests** over the corpus, both ways — the thing that proves 2 and - [ ] **4 — Round-trip property tests** over the corpus, both ways — the thing that proves 2 and
3. Generators emit editor-normal ADF (§2). 3. Generators emit editor-normal ADF (§2). Real sanitized ADF from live Atlassian APIs lands
here too (§10), in `corpus/real-payloads/`: an ADF→markdown→ADF check with no expected
markdown, the payloads supplied by the maintainer.
- [ ] **5 — Release pipeline, ship `0.1.0`.** Publish-on-version-change (§9), `NPM_TOKEN` secret, - [ ] **5 — Release pipeline, ship `0.1.0`.** Publish-on-version-change (§9), `NPM_TOKEN` secret,
the repo made public first (§6). `0.1.0` is the markdown round-trip: both markdown the repo made public first (§6). `0.1.0` is the markdown round-trip: both markdown
directions, the types, `isAdfDocument`. The build lands here: a build tsconfig emitting JS directions, the types, `isAdfDocument`. The build lands here: a build tsconfig emitting JS