11 KiB
Todo
The plan. Design questions are settled in AGENTS.md; remaining spec detail is settled at its own
milestone. A done item shrinks to its title here; its full text moves to todo-history.md.
Milestones
Shipping order: 3h, 3i, 3j, 5a, 5b, 5c, 5d, 5 → 0.1.0; 4b and 4c → 0.1.1; 4, 3k → 0.2.0;
6, 7 → 0.3.0.
The numbering is the order the work was planned in, not the order it ships.
- 0 — Scaffold.
- 1a — The directive grammar.
- 1b — Block node syntaxes.
- 1c — Inline node syntaxes and marks.
- 1d — Corpus start.
- 1d1 — The CommonMark subset.
- 1d2 — Block nodes.
- 1d3 — Inline nodes and marks.
- 2 —
adfToMarkdown.- 2a — The runner and the CommonMark subset.
- 2b — Block nodes.
- 2c — Inline nodes and marks.
- 2d — The opaque carry.
- 2e — Carve-outs and combinations.
- 2e1 — The carve-outs and the claimed line.
- 2e2 — Mark runs and the runs a carry breaks.
- 2e3 — Attribute canonicalization and the quoted value's escape.
- 2e4 — The carry's fallback triggers.
- 2e5 — Combined documents and the collision property.
- 2f — The attributes CommonMark cannot hold.
- 3 —
markdownToAdf(0.1.0). Each sub-item lands the fixtures its own code reads, and the runner grows a parse half as they do: readers forcorpus/normalization/(setext, indented code, loose lists,*/+bullets, entity references, soft wraps — one-way, the markdown not canonical) andcorpus/errors/(a markdown input per named error, the code in a.errorbeside it) with the first fixture each.commonmark-subset/cannot be the first to green —::paragraphand:hardBreak{}sit in it — so 3b through 3f answer to their own tests and the one-way fixtures they land, and 3g is where the first directory reads back. The raw-HTML element mapping is empty until milestone 6, so at0.1.0every raw-HTML construct in input — a block, an inline tag, a comment, a processing instruction — is a named error. Input is where unbounded nesting actually arrives, so §11's 500 binds all three of the emitter's guards here: block depth at 3c and again at 3f's container fences, inline and mark depth at 3f and 3i, a carried value's JSON at 3j, whereisJsonValuealready bounds it.- 3a — The hierarchy.
- 3b — The leaf blocks.
- 3c — The container blocks.
- 3d — Inline text.
- 3e — Emphasis and links.
- 3f — The directive grammar.
- 3g — The node tables read backwards.
- 3h — The block nodes.
- 3i — The inline nodes and the marks.
- 3j — The carry and the combinations.
- 3k — The CommonMark spec suite (
0.2.0). Checked in atcorpus/commonmark-spec/, pinned to the version it ships — the onehtml-blocks.tsnames for its start conditions —corpus/README.mdgaining the kind. Settled (the maintainer, 2026-08-27): three checks an example must pass, the reference HTML each ships read as corpus data — which adds no format and no direction (§1). §2's canonical fixpoint: a named error, or markdown that parses and emits to itself byte for byte. That HTML's text, tags stripped and entities decoded, against the parsed document's concatenatedtext. And a count of the dozen elements the CommonMark subset covers against the marks and nodes they map to — counting distinct mark types per text node, since 3e collapses a spelling nested inside its own kind and*(*a*)*is two<em>against oneem. The fixpoint alone is self-consistency a parser returning the empty document passes, and the text alone one dropping every emphasis; the counts close both. The exception list stays the maintainer's, and one entry is owed already: 3h continues a list across the marker change CommonMark splits on, so an example the reference HTML gives two<ul>counts onebulletList. One outcome is no exception and must not be filed as one: valid CommonMark parsing to a documentadfToMarkdownrefuses is a §2 hole, which is whatcorpus/unspellable/held until 3c, 3e and 3h landed their answers and emptied it.
- 4 — Round-trip property tests (
0.2.0), widening 3j's corpus round-trip past the documents a human wrote — the thing that proves 2 and 3 beyond them. Editor-normal (§2) is finished here, on 3i's merging —toEditorNormal(doc)and the equality the round-trip asserts, which over normalized input is the canonical serializer's compact spelling — rather than staying spelled inline as?? []at every reader. The reading half isnodeContent/nodeAttrs/nodeMarksover the ~28 sites spelling it inline today, which also lifts the branch floor §10 keeps below 100 for exactly those halves. Generators emit editor-normal ADF (§2). Real sanitized ADF from live Atlassian APIs lands here too (§10), incorpus/real-payloads/: an ADF→markdown→ADF check with no expected markdown, the payloads supplied by the maintainer. This subsumes 2e5's collision property — a document that round-trips proves no other document shares its spelling — so decide here whether that gate stays as the parser-free, faster-failing signal or goes; the half holding no fixture duplicates is hygiene rather than a round-trip claim, and stays either way. - 4b — The block walk's retry (
0.1.1).emitBlockwalks a subtree twice whereverreadableBlockreads it whole and then gives up — a list item whose first line reads back as a thematic break — and the walk below does the same, so the cost doubles per level: 3.4kB of nested lists takes half a second, depth 20 about eight, depth 24 minutes. It predates 3g on both directions, and 3g'scommonMarkSpellinggave it a second entry point. The README's bot and pipeline personas feed markdown nobody typed, so this ships as a hang on a small input; §11's scanning rule is the same argument one shape further in. The retry is what to remove — one walk answering both the readable question and the directive fallback. MemoizingemitBlockis the shortcut, and the node reference is the wrong key: a caller may hold one node object at two positions, where the cached depth and path are another node's.0.1.0ships with the retry in it, so a deep document is slow rather than wrong until the patch.adfDocumentFaultis the second site to look at:isNodeArrayreads every node and attribute value, thennestingFaultreads them again, so the emit entry the export persona runs in bulk walks the document twice. Both walks are linear, so this is a constant factor rather than 4b's class change, and the parting is what gives depth its own code (§8) — measure before joining them back. - 4c — The scanning rule's remaining sites (
0.1.1). A trailing-anchored regex re-walks its run from every start position, so an interior whitespace run costs quadratic time rather than linear — 3h measured 80k spaces inside an ATX heading at 11.3s, and 3ms once the walk replaced the regex. Three sites the same sweep did not reach:normalizeLabelinlink-syntax.ts, whose shortcut-reference input isscan.source.slice(...)rather than the 999-cappedreadLabelvalue, and two inemit/inline-line.ts. The fix is the one 3h used — an index walk,trimTrailingSpacewhere the ends match. A fourth of another shape joins them:readNestedDirectiverestarts its depth counter per level, so each parse level re-scans the region below it and nested inline directives cost O(depth × content) — 3f's cost, which 3i's slot parse doubles rather than changes in class, bounded by the 500-level guard. §11's scanning rule is the whole argument; the pipeline persona feeds documents nobody typed. - 5 — Ship
0.1.0. Only the maintainer's own acts are left (§15): make the Gitea repo public (§6), create theNPM_TOKENsecret, and open the bump PR that setsversionto0.1.0and dropsprivate: true, the guard against any earlier publish.0.1.0is the markdown round-trip: both markdown directions, the types,isAdfDocument, proved over the checked-in corpus. Settled (the maintainer, 2026-09-01): the round-trip proved over the checked-in corpus is what0.1.0ships on, and the open-ended proof work follows it rather than gating it — 3k's spec suite and 4's generators and maintainer-supplied payloads are0.2.0, 4b's retry0.1.1. A consumer using the library is worth more than a wider proof nobody has needed yet, and §8's pre-1.0 rules cover what the wider proof then finds. - 5a — Rename to
@larvit/adf-codec. - 5b — The consumer's error surface.
- 5b1 — The error's source position.
- 5b2 — The error messages.
- 5b3 — The code list and the flavour's gaps.
- 5b4 — The README's consumer surface.
- 5c — The build and the release pipeline.
- 5d — The browser leg.
- 6 — The HTML dialect spec (
0.3.0). Element-by-element mapping, thedata-*fidelity scheme, the opaque-carry form, and the documented foreign-element sethtmlToAdfaccepts. - 7 — HTML, ship
0.3.0.adfToHtml,htmlToAdf, the composedmarkdownToHtml/htmlToMarkdown. CommonMark spec suite runs againstmarkdownToHtmlfrom here (§10). - 8 — CLI. A later goal, shaped around the personas once the library exists.
The ADF inventory to cover
From Atlassian's structure
reference — not the
whole schema: real payloads also carry taskList/taskItem, decisionList/decisionItem,
layoutSection/layoutColumn, blockCard/embedCard, extension/bodiedExtension/inlineExtension
and placeholder, none documented there. The documented set is the floor: the floor gets designed
syntax, the rest rides the opaque carry (§3) until it does too.
| Top-level block | blockquote bodiedSyncBlock bulletList codeBlock expand heading mediaGroup mediaSingle multiBodiedExtension orderedList panel paragraph rule syncBlock table |
| Child block | blockTaskItem extensionFrame listItem media nestedExpand tableCell tableHeader tableRow |
| Inline | date emoji hardBreak inlineCard mediaInline mention status text |
| Marks | border code em link strike strong subsup textColor underline |
Plain markdown covers blockquote, bulletList, codeBlock, heading, orderedList,
paragraph, rule, listItem, hardBreak, text, and the code, em, link and strong
marks; strike is the flavour's ~~ carve-out. Everything else is what the flavour is for.