diff --git a/AGENTS.md b/AGENTS.md index bcae506..f2a449b 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -148,7 +148,7 @@ not take — is `unsupported-node-shape`, the emitter's code for the same mismatch read the other way — one code across both directions for good, since the call site knows which direction it called and parting them after `0.1.0` is MAJOR. `unmappable-html` names the version rather than the element: this one -converts no raw HTML, so at `0.3.0` the mapped elements stop erroring and the code stays for what +converts no raw HTML, so at `0.2.0` the mapped elements stop erroring and the code stays for what no ADF node carries. A refusal found before its path is known — the block walk's, a directive reader's — is a `ConvertFault`, the code and message alone; the node walk attaches the path as it descends, so a document reports its first error in document order. `not-an-adf-document` carries @@ -177,7 +177,7 @@ A parse names a position for every refusal it returns, so the type says so rathe `Result`, and a direction reading a source returns `Result` — `ConvertError` with `position` required. An optional field a direction always fills is a branch a consumer cannot take, and the `!` §11 bans is how they take it anyway. -`htmlToAdf` inherits this at `0.3.0`; the composed `markdownToHtml` and `htmlToMarkdown` keep the +`htmlToAdf` inherits this at `0.2.0`; the composed `markdownToHtml` and `htmlToMarkdown` keep the wide `Result`, since half their refusals come from an emit stage that read no source. ## 9. Release automation @@ -298,10 +298,17 @@ someone spells it or pins it. has — for the fallback to refuse at zero rather than walk again; counting every list twice halved the list limit, counting the directive form once doubled the parser's frames per level (the maintainer, 2026-09-18). +- Nothing spreads an unbounded array into a call — a node's siblings, a code block's held lines, a + mark run's segments: the argument list caps near 125k and throws a `RangeError` where a `Result` + is owed. A walk pushes one at a time. A literal spread (`[...value]`) is not the same thing and + is fine (4c). - A reader takes the text and an index — a sticky regex whose `lastIndex` the caller sets on the line before it reads, `indexOf` — never a fresh slice per character, and a per-character walk hoists the scan that does not vary with the character. The pipeline persona feeds documents - nobody typed, and a megabyte through a quadratic walk is a minute rather than a millisecond. + nobody typed, and a megabyte through a quadratic walk is a minute rather than a millisecond. A + scan may keep what it read for a later walk of the same text, and the fallback where it kept + nothing must be the same reader over the same text at the same index, so the two cannot disagree + — which is what makes the kept value a memo rather than a second spelling (4c). - No casts: `as`, `as unknown as`, non-null `!`. A boundary owes a type guard validating the fields it claims (`isAdfDocument`); past it everything is typed. Make invalid states unrepresentable. diff --git a/README.md b/README.md index 299928f..bcb8170 100644 --- a/README.md +++ b/README.md @@ -4,7 +4,7 @@ Lossless conversion between **Atlassian Document Format** (ADF), an extended mar an HTML dialect. **Status: published — the markdown round-trip (`adfToMarkdown`, `markdownToAdf`); HTML at -`0.3.0`.** +`0.2.0`.** Plan: `todo.md`. Decisions: `AGENTS.md`. The flavour's grammar: [`spec/flavour.md`](https://gitea.larvit.se/larvit/adf-codec/src/branch/main/spec/flavour.md). Upgrading from `0.1.0`: [convert your markdown first](https://gitea.larvit.se/larvit/adf-codec/src/branch/main/MIGRATION.md). @@ -44,10 +44,10 @@ adfToMarkdown(doc: AdfDocument): Result markdownToAdf(markdown: string): Result isAdfDocument(v: unknown): v is AdfDocument -adfToHtml(doc: AdfDocument): Result // 0.3.0 -htmlToAdf(html: string): Result // 0.3.0 -markdownToHtml(markdown: string): Result // 0.3.0, via ADF -htmlToMarkdown(html: string): Result // 0.3.0, via ADF +adfToHtml(doc: AdfDocument): Result // 0.2.0 +htmlToAdf(html: string): Result // 0.2.0 +markdownToHtml(markdown: string): Result // 0.2.0, via ADF +htmlToMarkdown(html: string): Result // 0.2.0, via ADF ``` `Result` is `{ ok: true; value: T } | { ok: false; error: ConvertError }` — nothing throws. @@ -75,17 +75,17 @@ UTF-16 code unit, a JavaScript string index rather than a codepoint or a byte of or before the refusal — currently the start of the line the enclosing block begins on; a later minor may narrow that, never widen it. -Parsing — `markdownToAdf`, and `htmlToAdf` at `0.3.0`: +Parsing — `markdownToAdf`, and `htmlToAdf` at `0.2.0`: | Code | Fires when | What you can do | | --- | --- | --- | | `malformed-directive` | an `!adf:` the grammar cannot read — a prefix completing no directive, an unclosed container, `[content]` or `{attrs}`, a closer with no container of its name open, a leaf given a body, `{attrs}` out of order or duplicated, invalid JSON in a `carry` | write the spelling the message names, or escape the prefix — `\!adf:`, block and inline alike — to keep it literal text | | `malformed-pipe-table` | a pipe row that is no pipe table — a missing or ragged `---` delimiter row, an alignment colon in it, or a row not opening with a pipe | open every row with a pipe and give the delimiter row the header's cell count; to keep the lines literal text instead, escape the leading pipe of every one — escaping a single row leaves the next to open a fresh table and fail the same way | | `unknown-directive-name` | a directive whose name is no node or mark this version spells | check the name in `spec/flavour.md`, or escape the prefix as `\!adf:`; the spelling itself is well formed, so a later minor may give the name meaning | -| `unmappable-html` | the markdown holds a raw HTML tag, comment or processing instruction | remove it or write it in the flavour — ADF holds no raw-HTML node, and the element mapping lands at `0.3.0` | +| `unmappable-html` | the markdown holds a raw HTML tag, comment or processing instruction | remove it or write it in the flavour — ADF holds no raw-HTML node, and the element mapping lands at `0.2.0` | | `unmappable-image` | an image sits inside other content, or carries a title | give the image a paragraph of its own and drop the title | -Emitting — `adfToMarkdown`, and `adfToHtml` at `0.3.0`: +Emitting — `adfToMarkdown`, and `adfToHtml` at `0.2.0`: | Code | Fires when | What you can do | | --- | --- | --- | @@ -121,7 +121,7 @@ emit refuses: reference matching its definition only under Unicode case folding stays unresolved. Each is pinned `pending` in `corpus/commonmark-spec/exceptions.json`. - Raw HTML in markdown input is an error result, never a silent drop — a tag, a comment and a - processing instruction alike. ADF holds no raw-HTML node; the element mapping ships at `0.3.0`. + processing instruction alike. ADF holds no raw-HTML node; the element mapping ships at `0.2.0`. - Not every document converts back: `adfToMarkdown` is partial on valid ADF — a text node holding a carriage return, or a paragraph line beginning with a code span whose backticks read back as a fence. Show the refusal and keep the document read-only; saving markdown you could not produce @@ -134,7 +134,7 @@ emit refuses: is the `taskList` directive. - A document nested deeper than 500 levels is an error result, not a stack overflow. - The emitted formats are semver surface (AGENTS.md §8). -- **`0.3.0`** — `htmlToAdf(adfToHtml(doc))` equals `doc`; fidelity HTML cannot express rides +- **`0.2.0`** — `htmlToAdf(adfToHtml(doc))` equals `doc`; fidelity HTML cannot express rides `data-*` attributes. Foreign HTML maps a documented element set, an unmappable element is an error, and well-formed HTML only — no tag-soup recovery. diff --git a/src/adf/document.test.ts b/src/adf/document.test.ts index b72bfc9..4efd5f9 100644 --- a/src/adf/document.test.ts +++ b/src/adf/document.test.ts @@ -77,3 +77,9 @@ test('accepts the JSON values an attribute may hold', () => { assert.equal(isAdfDocument({ content: [{ attrs: { a: [1, 'x', null, true, { b: 2 }] }, type: 'paragraph' }], type: 'doc', version: 1 }), true) assert.equal(isAdfDocument({ content: [{ attrs: { a: [() => 1] }, type: 'paragraph' }], type: 'doc', version: 1 }), false) }) + +test("reads a node's siblings as a walk rather than as one call's arguments", () => { + const wide = { content: [{ content: Array.from({ length: 200000 }, () => ({ type: 'rule' })), type: 'blockquote' }], type: 'doc', version: 1 } + assert.equal(isAdfDocument(wide), true) + assert.equal(fault(wide), 'accepted') +}) diff --git a/src/adf/document.ts b/src/adf/document.ts index bedca8c..765ab9b 100644 --- a/src/adf/document.ts +++ b/src/adf/document.ts @@ -96,7 +96,7 @@ function isNodeArray(value: readonly unknown[]): value is readonly AdfNode[] { if ('content' in node) { const content = node['content'] if (!Array.isArray(content)) return false - pending.push(...content) + for (const child of content) pending.push(child) } } return true @@ -109,7 +109,7 @@ function nestingFault(nodes: readonly AdfNode[]): ConvertFault | undefined { if (node === undefined) continue const fault = attributesFault(nodeAttrs(node), node.type) ?? marksFault(nodeMarks(node)) if (fault !== undefined) return fault - pending.push(...nodeContent(node)) + for (const child of nodeContent(node)) pending.push(child) } return undefined } diff --git a/src/markdown/commonmark-grammar.ts b/src/markdown/commonmark-grammar.ts index 19b73d4..53c9f1a 100644 --- a/src/markdown/commonmark-grammar.ts +++ b/src/markdown/commonmark-grammar.ts @@ -2,6 +2,9 @@ import { readEntityReference, replacementCharacter } from './entity-references.t export type LinePosition = 'first' | 'later' +// The suffix lengths of a text that spell a thematic break, and the marker each of them opens with. +export type ThematicBreakTail = { longest: number; marker: string; shortest: number } + type OpenHtmlBlock = { closer: RegExp | undefined; construct: string } type HtmlBlockCondition = { closer: RegExp | undefined; construct: string | undefined; interrupts: boolean; start: RegExp } @@ -52,6 +55,7 @@ const htmlBlockConditions: HtmlBlockCondition[] = [ const asciiPunctuation = /[!"#$%&'()*+,\-./:;<=>?@[\\\]^_`{|}~]/ const atxHeadingOpener = /^(#{1,6})(?:[ \t]|$)/ const blankLine = /^[ \t]*$/ +const breakMarkers = '*-_' const codeFenceOpener = /^(`{3,}|~{3,})/ const pipeClaim = /^\|/ const bulletListOpener = /^[*+-](?:[ \t]|$)/ @@ -63,7 +67,6 @@ const emailLabelSource = '[A-Za-z0-9](?:[A-Za-z0-9-]{0,61}[A-Za-z0-9])?' const emailAutolink = new RegExp(`<${emailNameSource}@${emailLabelSource}(?:\\.${emailLabelSource})*>`, 'y') const orderedListOpener = /^(\d{1,9})(?:[.)])(?:[ \t]|$)/ const setextUnderline = /^(=+|-+)[ \t]*$/ -const thematicBreak = /^(?:(?:\*[ \t]*){3,}|(?:-[ \t]*){3,}|(?:_[ \t]*){3,})$/ const unicodeWhitespace = /[\t\n\f\r \p{Zs}]/u export function atxHeading(line: string): { level: number; text: string } | undefined { @@ -118,7 +121,7 @@ export function decodeTextEscapes(text: string): string { export function escapesLineClaim(line: string, offset: number, position: LinePosition): boolean { if (offset === 0) { - if (firstCharacterOpeners.some((opener) => opener.test(line)) || thematicBreak.test(line)) return true + if (firstCharacterOpeners.some((opener) => opener.test(line)) || isThematicBreak(line)) return true if (openingHtmlBlock(line, position === 'later') !== undefined) return true return position === 'later' && setextUnderline.test(line) } @@ -168,7 +171,34 @@ export function isBlankLine(line: string): boolean { } export function isThematicBreak(line: string): boolean { - return thematicBreak.test(line) + return holdsThematicBreak(thematicBreakTail(line), line) +} + +export function thematicBreakTail(text: string): ThematicBreakTail | undefined { + let marker: string | undefined + let markers = 0 + let first = text.length + let third = text.length + let cursor = text.length + while (cursor > 0) { + const character = text.charAt(cursor - 1) + if (marker === undefined ? breakMarkers.includes(character) : character === marker) { + marker = character + markers += 1 + first = cursor - 1 + if (markers === 3) third = cursor - 1 + } else if (!spaceOrTab(character)) break + cursor -= 1 + } + if (marker === undefined || markers < 3) return undefined + return { longest: text.length - first, marker, shortest: text.length - third } +} + +// `text` is a suffix of what `tail` was read from, or opens with a space: a length alone names a +// suffix, and the marker check is what refuses one starting mid-run — `--- ---` holds ` ---`. +export function holdsThematicBreak(tail: ThematicBreakTail | undefined, text: string): boolean { + if (tail === undefined) return false + return text.length >= tail.shortest && text.length <= tail.longest && text.charAt(0) === tail.marker } export function isUnicodeWhitespace(character: string): boolean { diff --git a/src/markdown/directive-syntax.ts b/src/markdown/directive-syntax.ts index af3a669..5ae76af 100644 --- a/src/markdown/directive-syntax.ts +++ b/src/markdown/directive-syntax.ts @@ -17,7 +17,10 @@ export type DirectiveLine = | { argument: string | undefined; attributes: DirectiveAttributes; kind: 'opener'; name: string } | { kind: 'closer'; name: string } -export type DirectiveSpan = { attributes: DirectiveAttributes; content: string | undefined; length: number; name: string } +// spans keys index content, so the two travel together: separate them and every offset is wrong. +export type DirectiveSpan = { attributes: DirectiveAttributes; content: string | undefined; length: number; name: string; spans: NestedSpans } + +export type NestedSpans = ReadonlyMap export type Read = { fault: ConvertFault; value?: undefined } | { fault?: undefined; value: T } @@ -25,7 +28,9 @@ type Attributes = { attributes: DirectiveAttributes; end: number } type AttributePair = { end: number; key: string; value: DirectiveValue } -type Content = { content: string | undefined; end: number } +type Content = { content: string | undefined; end: number; spans: NestedSpans } + +type DirectiveContent = { end: number; spans: NestedSpans } export const directivePrefix = '!adf:' @@ -41,6 +46,8 @@ const quotedEscapes = new RegExp(reservedSource, 'g') const rawReserved = new RegExp(reservedSource) const noAttributes: DirectiveAttributes = new Map() +export const noSpans: NestedSpans = new Map() + export const directiveEscape = `\\${directivePrefix} keeps the prefix literal` const closerFault = `a closer carries nothing after its name: this one does; ${directiveEscape}` @@ -213,14 +220,14 @@ function readNestedDirective(text: string, index: number, depth: number): Read { - if (text.charAt(index) !== '[') return { value: { content: undefined, end: index } } + if (text.charAt(index) !== '[') return { value: { content: undefined, end: index, spans: noSpans } } const close = readDirectiveContent(text, index + 1, depth) if (close.fault !== undefined) return { fault: close.fault } - return { value: { content: text.slice(index + 1, close.value), end: close.value + 1 } } + return { value: { content: text.slice(index + 1, close.value.end), end: close.value.end + 1, spans: close.value.spans } } } function readAttributesAt(text: string, index: number, braceClaims: boolean): Read { @@ -231,7 +238,8 @@ function readAttributesAt(text: string, index: number, braceClaims: boolean): Re } // A code span, an escape and a nested directive each bind before the content's own closing bracket. -function readDirectiveContent(text: string, start: number, depth: number): Read { +function readDirectiveContent(text: string, start: number, depth: number): Read { + const spans = new Map() let brackets = 0 let cursor = start while (cursor < text.length && text.charAt(cursor) !== '\n') { @@ -247,12 +255,13 @@ function readDirectiveContent(text: string, start: number, depth: number): Read< continue } const nested = readNestedDirective(text, cursor, depth + 1) + if (nested?.fault !== undefined) return { fault: nested.fault } if (nested !== undefined) { - if (nested.fault !== undefined) return { fault: nested.fault } + spans.set(cursor - start, nested.value) cursor += nested.value.length continue } - if (character === ']' && brackets === 0) return { value: cursor } + if (character === ']' && brackets === 0) return { value: { end: cursor, spans } } if (character === '[') brackets += 1 if (character === ']') brackets -= 1 cursor += 1 diff --git a/src/markdown/emit/adf-to-markdown.test.ts b/src/markdown/emit/adf-to-markdown.test.ts index ca07e0a..e5a2969 100644 --- a/src/markdown/emit/adf-to-markdown.test.ts +++ b/src/markdown/emit/adf-to-markdown.test.ts @@ -736,3 +736,8 @@ test('carries whitespace CommonMark strips in the reserved text directive', () = // CommonMark strips spaces and tabs alone, so the whitespace beside them is plain text. assert.equal(emitted({ text: '\va\f', type: 'text' }), '\va\f\n') }) + +test("joins a mark run's segments as a walk rather than as one call's arguments", () => { + const run = Array.from({ length: 200000 }, (): AdfNode => ({ marks: [{ type: 'strong' }], text: 'a', type: 'text' })) + assert.equal(markdown(adfToMarkdown(document({ content: run, type: 'paragraph' }))), `**${'a'.repeat(200000)}**\n`) +}) diff --git a/src/markdown/emit/inline-line.ts b/src/markdown/emit/inline-line.ts index 06bfe66..0cddd44 100644 --- a/src/markdown/emit/inline-line.ts +++ b/src/markdown/emit/inline-line.ts @@ -2,7 +2,7 @@ import type { AdfMark, AdfNode } from '../../adf/document.ts' import type { InlineDirective } from '../../adf/inline-directives.ts' import { assembleInlineLine, isSyntax, type InlineEscaping, type InlineSegment, type LineContainer, type NodeRange } from './line-escaping.ts' import { carriedInline } from '../opaque-carry.ts' -import { claimsLine, holdsNullCharacter } from '../commonmark-grammar.ts' +import { claimsLine, holdsNullCharacter, trimTrailingSpace } from '../commonmark-grammar.ts' import { commonMarkLink, markSpelling, spellMarkAttributes } from '../mark-spellings.ts' import { escapeUnbalanced, spellDestination } from '../link-syntax.ts' import { failure, faulted, success, type ConvertErrorPath, type Result } from '../../result.ts' @@ -129,8 +129,8 @@ function carryEdges(segment: InlineSegment, leading: boolean, trailing: boolean) if (segment.escaping !== 'backslash' && segment.escaping !== 'bracketed') return [segment] const head = leading ? (/^[ \t]+/.exec(segment.text)?.[0] ?? '') : '' const body = segment.text.slice(head.length) - const tail = trailing ? (/[ \t]+$/.exec(body)?.[0] ?? '') : '' - const middle = body.slice(0, body.length - tail.length) + const middle = trailing ? trimTrailingSpace(body) : body + const tail = body.slice(middle.length) const edges: InlineSegment[] = [] if (head !== '') edges.push(carriedText(head)) if (middle !== '') edges.push({ escaping: segment.escaping, text: middle }) @@ -166,7 +166,7 @@ function emitRun(nodes: readonly AdfNode[], depth: number, firstIndex: number, c const emitted = run.kind === 'plain' ? emitLeaf(run.node, runContext, run.index) : emitMarkedRun(run.nodes, run.mark, depth, run.index, runContext) if (!emitted.ok) return emitted if (emitted.value.carry !== undefined) return emitted - segments.push(...emitted.value.segments) + for (const segment of emitted.value.segments) segments.push(segment) } return success({ segments }) } diff --git a/src/markdown/link-syntax.ts b/src/markdown/link-syntax.ts index 0210f9d..c34e58b 100644 --- a/src/markdown/link-syntax.ts +++ b/src/markdown/link-syntax.ts @@ -17,10 +17,10 @@ export function readLabel(text: string, offset: number): LinkPart | undefined { } export function normalizeLabel(raw: string): string { - return raw - .replace(/^[ \t\n]+|[ \t\n]+$/g, '') - .replace(/[ \t\n]+/g, ' ') - .toLowerCase() + const collapsed = raw.replace(/[ \t\n]+/g, ' ') + const start = collapsed.startsWith(' ') ? 1 : 0 + const end = collapsed.endsWith(' ') ? collapsed.length - 1 : collapsed.length + return collapsed.slice(start, end).toLowerCase() } export function readDestination(text: string, offset: number): LinkPart | undefined { diff --git a/src/markdown/parse/blocks.test.ts b/src/markdown/parse/blocks.test.ts index f4061be..2866051 100644 --- a/src/markdown/parse/blocks.test.ts +++ b/src/markdown/parse/blocks.test.ts @@ -118,3 +118,7 @@ test('closes the containers a closer names past as unclosed, and crosses no list assert.deepEqual(kinds('!adf:rule {localId=a-1}\nPart.\n!adf:/rule\n'), ['fault', 'paragraph']) assert.deepEqual(kinds('!adf:rule {localId=a-1}\n!adf:/rule\n!adf:/rule\n'), ['fault', 'fault']) }) + +test("releases an indented code block's held blank lines as a walk rather than as one call's arguments", () => { + assert.deepEqual(kinds(` a\n${'\n'.repeat(200000)} b\n`), ['code']) +}) diff --git a/src/markdown/parse/blocks.ts b/src/markdown/parse/blocks.ts index 4983ae1..b011452 100644 --- a/src/markdown/parse/blocks.ts +++ b/src/markdown/parse/blocks.ts @@ -6,6 +6,7 @@ import { claimsPipeLine, closingCodeFence, decodeTextEscapes, + holdsThematicBreak, isBlankLine, isThematicBreak, listMarker, @@ -14,6 +15,8 @@ import { openingHtmlBlock, replaceNullCharacters, setextHeadingLevel, + thematicBreakTail, + type ThematicBreakTail, } from '../commonmark-grammar.ts' import { blockDirectiveForm } from '../block-directive-forms.ts' import { directiveEscape, malformedDirective, readDirectiveLine, spellDirectiveCloser } from '../directive-syntax.ts' @@ -143,10 +146,11 @@ function blockquoteRest(opener: Line): Line | undefined { function openContainers(walk: Walk, line: Line, paragraphOpen: boolean, depth: number): { opened: boolean; rest: Line } { const unmatched = walk.stack[depth] + const tail = thematicBreakTail(line.text) let opened = false let rest = line while (leadingColumns(rest) < indentedCodeColumns) { - const start = containerStart(rest, opened ? false : paragraphOpen, opened ? undefined : unmatched, walk.position) + const start = containerStart(rest, tail, opened ? false : paragraphOpen, opened ? undefined : unmatched, walk.position) if (start === undefined) break if (!opened) closeContainers(walk, depth) opened = true @@ -156,11 +160,17 @@ function openContainers(walk: Walk, line: Line, paragraphOpen: boolean, depth: n return { opened, rest } } -function containerStart(line: Line, paragraphOpen: boolean, enclosing: OpenContainer | undefined, position: SourcePosition): ContainerStart | undefined { +function containerStart( + line: Line, + tail: ThematicBreakTail | undefined, + paragraphOpen: boolean, + enclosing: OpenContainer | undefined, + position: SourcePosition, +): ContainerStart | undefined { const opener = removeColumns(line, largestOpenerIndentation) const blockquote = blockquoteRest(opener) if (blockquote !== undefined) return { kind: 'blockquote', rest: blockquote } - if (isThematicBreak(opener.text)) return undefined + if (holdsThematicBreak(tail, opener.text)) return undefined return itemStart(line, opener, paragraphOpen, enclosing, position) } @@ -355,7 +365,8 @@ function readIndentedCodeLine(leaf: Extract return true } if (leadingColumns(line) < indentedCodeColumns) return false - leaf.lines.push(...leaf.held, removeColumns(line, indentedCodeColumns).text) + for (const held of leaf.held) leaf.lines.push(held) + leaf.lines.push(removeColumns(line, indentedCodeColumns).text) leaf.held.length = 0 return true } diff --git a/src/markdown/parse/inline-content.ts b/src/markdown/parse/inline-content.ts index 25ccfdf..e969b79 100644 --- a/src/markdown/parse/inline-content.ts +++ b/src/markdown/parse/inline-content.ts @@ -1,5 +1,5 @@ import type { AdfMark, AdfNode } from '../../adf/document.ts' -import type { DirectiveSpan } from '../directive-syntax.ts' +import type { DirectiveSpan, NestedSpans } from '../directive-syntax.ts' import type { EmphasisPairing } from '../emphasis-matching.ts' import type { LineContainer } from '../emit/line-escaping.ts' import type { LinkDefinition } from '../link-syntax.ts' @@ -15,7 +15,7 @@ import { normalizeLabel, readInlineTarget, readLabel } from '../link-syntax.ts' import { openingLinkTakesDirective } from '../emit/inline-line.ts' import { readCarriedInline } from '../opaque-carry.ts' import { readDirectiveMark } from './directive-marks.ts' -import { readInlineDirective } from '../directive-syntax.ts' +import { noSpans, readInlineDirective } from '../directive-syntax.ts' import { readInlineDirectiveNode } from './directive-nodes.ts' import { readTextDirective } from '../text-directive.ts' @@ -45,6 +45,7 @@ type Scan = { pending: string pieces: Piece[] source: string + spans: NestedSpans } type SlotContent = { carry: boolean; nodes: AdfNode[] } @@ -54,11 +55,11 @@ const imageAlone = 'an image fits only as a paragraph of its own: this one sits const spellableLink = 'link takes the directive form only where CommonMark cannot spell it: this one it can, as [text](url "title") or ' export function parseInlineContent(source: string, definitions: LinkDefinitions, path: ConvertErrorPath, container: LineContainer): Result { - return parseInline(source, definitions, path, container) + return parseInline(source, definitions, path, container, noSpans) } -function parseInline(source: string, definitions: LinkDefinitions, path: ConvertErrorPath, container: LineContainer | undefined): Result { - const scan: Scan = { container, definitions, openingSpellableLink: false, path, pending: '', pieces: [], source } +function parseInline(source: string, definitions: LinkDefinitions, path: ConvertErrorPath, container: LineContainer | undefined, spans: NestedSpans): Result { + const scan: Scan = { container, definitions, openingSpellableLink: false, path, pending: '', pieces: [], source, spans } let index = 0 while (index < source.length) { switch (source.charAt(index)) { @@ -168,14 +169,20 @@ function openBracket(scan: Scan, index: number): number { } function readDirective(scan: Scan, index: number): Result | undefined { + const held = scan.spans.get(index) + if (held !== undefined) return pushDirective(scan, held, index) const directive = readInlineDirective(scan.source, index) if (directive === undefined) return undefined if (directive.fault !== undefined) return faulted(directive.fault, scan.path) - const piece = directivePiece(scan, directive.value, index) + return pushDirective(scan, directive.value, index) +} + +function pushDirective(scan: Scan, span: DirectiveSpan, index: number): Result { + const piece = directivePiece(scan, span, index) if (!piece.ok) return piece flush(scan, false) scan.pieces.push(piece.value) - return success(index + directive.value.length) + return success(index + span.length) } function directivePiece(scan: Scan, span: DirectiveSpan, index: number): Result { @@ -187,7 +194,7 @@ function directivePiece(scan: Scan, span: DirectiveSpan, index: number): Result< const text = readTextDirective(span) if (text?.fault !== undefined) return faulted(text.fault, scan.path) if (text !== undefined) return success({ kind: 'nodes', nodes: [{ text: text.value, type: 'text' }] }) - const slot = slotContent(scan, span.content) + const slot = slotContent(scan, span) if (!slot.ok) return slot const mark = readDirectiveMark(span.name, span.attributes, scan.path) if (mark !== undefined) return mark.ok ? directiveMarkPiece(scan, span.name, mark.value, slot.value, index) : mark @@ -216,9 +223,9 @@ function refuseSpellableLink(scan: Scan, mark: AdfMark, nodes: readonly AdfNode[ return undefined } -function slotContent(scan: Scan, content: string | undefined): Result { - if (content === undefined) return success(undefined) - const parsed = parseInline(content, scan.definitions, scan.path, undefined) +function slotContent(scan: Scan, span: DirectiveSpan): Result { + if (span.content === undefined) return success(undefined) + const parsed = parseInline(span.content, scan.definitions, scan.path, undefined, span.spans) if (!parsed.ok) return parsed if (parsed.value.image !== undefined) return failure('unmappable-image', imageAlone, scan.path) return success(parsed.value) diff --git a/todo-history.md b/todo-history.md index 5560327..9847d6b 100644 --- a/todo-history.md +++ b/todo-history.md @@ -504,6 +504,48 @@ The done `todo.md` items in full, as they were written. `todo.md` keeps a one-li counting the directive form once for doubling the parser's frames per level. **Measured** (2026-09-18): `adfDocumentFault` walks a 9 MB document in 52 ms against 314 ms for the emit, so its two walks stay parted. +- [x] **4c — The scanning rule's remaining sites (`0.2.0`).** A trailing-anchored regex re-walks + its run from every start position, so an interior whitespace run costs quadratic time rather + than linear — 3h measured 80k spaces inside an ATX heading at 11.3s, and 3ms once the walk + replaced the regex. The sites the same sweep did not reach: `normalizeLabel` in + `link-syntax.ts`, whose shortcut-reference input is `scan.source.slice(...)` rather than the + 999-capped `readLabel` value, and `carryEdges` in `emit/inline-line.ts`. A third of another + shape joins them: `readNestedDirective` restarts its depth counter per level, so each parse + level re-scans the region below it and nested inline directives cost O(depth × content), + bounded by the 500-level guard. A fourth predates 12c: the list-item walk re-scans the rest + of a line once per item level — `isThematicBreak` in `containerStart` on an opener line, + `isBlankLine` and `leadingColumns` in `continuesContainer` on a continuation line, and a + blank line continues every open item without consuming input; 30000 nested items take 4.4s + at 59 KB (the stability-reviewer, 2026-09-16). §11's scanning rule is the whole argument; the + pipeline persona feeds documents nobody typed. A fifth is a throw rather than a cost: + `adfDocumentFault` pushes a node's content with a spread, so past about 125k sibling nodes + the guard throws a `RangeError` where §11 owes a `Result` (the stability-reviewer and the + maintainer, 2026-09-18). + **Settled** (the maintainer, 2026-09-18): the five sites land in one PR rather than split + into sub-items, and a behaviour-preserving cost fix is accepted on the suite staying green + with no fixture output changed, plus the measurement below — §14 promises no figure, so + nothing times the gate. The guard's spread is the one behavioural fix and carries a test. + **Corrected** (2026-09-18): the entry filed two sites in `emit/inline-line.ts` on 2026-09-01 + and the file has changed since — `tryImageLine`'s alternation measures linear (3.4 / 1.9 / + 5.3 ms over 10k / 20k / 40k spaces), leaving `carryEdges`' trailing trim the only one. + **Measured** (2026-09-18), each at the size its filing named: `normalizeLabel` 1026 ms → 5 ms + at 40k interior spaces, `carryEdges` 1024 ms → 7 ms (its heading path 978 ms → 5 ms), + `readNestedDirective` 434 ms → 10 ms at 397 kB and 200 levels, the list-item walk 4196 ms → + 39 ms at 30000 items, and the document guard a `RangeError` → 42 ms at 200k siblings under + one node. The list-item walk's mixed-marker shape, which the fix had to answer too, reads + 38 ms where the tail scan alone would have left it quadratic. + The sixth site the sweep found went to 18 rather than landing here (the maintainer, + 2026-09-18). + **Left as is** (the stability-reviewer, 2026-09-18): of the list-item walk's three re-scans + only `containerStart`'s is fixed. `continuesContainer`'s pair costs the same either way — 400 + levels at 627 kB read 469 ms before and 448 ms after, linear in the line count and only + mildly superlinear in a depth the 500-level guard bounds — so it is measured and left rather + than made an item. + **Widened** (the systems-architect, 2026-09-18): the guard's spread was a class rather than a + site, and two more threw out of the public API — `readIndentedCodeLine` releasing the blank + lines an indented code block held (200k of them at 200 kB), and `emitRun` joining a mark + run's segments (200k nodes under one mark). Both fixed here with the same loop and a test + each, and §11 gained the rule so the spelling cannot walk back in. - [x] **5a — Rename to `@larvit/adf-codec` (`0.1.0`).** Before the first publish, the name being the published identity: `package.json` `name` and `repository`, the Gitea repo and its remote, the README title, §6's published-as line, the checkout directory. diff --git a/todo.md b/todo.md index 6e9a25a..6d51374 100644 --- a/todo.md +++ b/todo.md @@ -9,18 +9,26 @@ Start a session with: `Read AGENTS.md and todo.md, then do what todo.md's "Next 1. The first unchecked item in shipping order, per AGENTS.md §15 — or, where that item has no release, the planning chunk §15 describes. -2. In flight: 4c has uncommitted work in `.claude/worktrees/scanning-rule-sites` (branch - `4c-scanning-rule-sites`) and a stray `bench-4c.ts` that does not ship; continue from the diff. -3. Before stopping, rewrite this section: the in-flight line, and the prompt itself wherever the +2. In flight: nothing. +3. `git fetch origin` and branch off `origin/main`, not the worktree left behind: `main` moved + under 4c mid-chunk and the branch needed a rebase before it could merge. +4. Before stopping, rewrite this section: the in-flight line, and the prompt itself wherever the session found it wrong or short. ## Milestones -Shipping order: 3h, 3i, 3j, 5a, 5b, 5c, 5d, 5 → `0.1.0` (shipped 2026-09-05); 3k, 11, 4, 12, 13, 4b, 4c, 14, 15, 16, 10, 5g → `0.2.0`; -4d, 5f → `0.2.1`; 6, 7 → `0.3.0`; 9, 17 → TBD; 5e last. -The numbering is the order the work was planned in, not the order it ships. `0.2.0`'s order is settled -(the maintainer, 2026-09-13): 11 makes the tables 4 generates from answer to Atlassian's schema, 4 -proves 12, 13 spells 11's gaps in 12's grammar, and 12 rewrites code 4b and 4c change. +Shipping order: 3h, 3i, 3j, 5a, 5b, 5c, 5d, 5 → `0.1.0` (shipped 2026-09-05); 3k, 11, 4, 12, 13, 4b, +4c, 14, 15, 16, 18, 4d, 17, 10, 6, 7, 5f, 5g → `0.2.0`; 8, 9 → TBD; 5e last. +The numbering is the order the work was planned in, not the order it ships. Everything known and +shaped ships in one release rather than a string of them: nothing waits on a version, and no +consumer is served by the churn (the maintainer, 2026-09-18). So `0.2.0` completes §1's three +formats, and `0.2.1` and `0.3.0` are gone. `8` and `9` stay out as the two goals nothing has shaped +yet. `0.2.0`'s order is settled (the maintainer, 2026-09-13, extended 2026-09-18): 11 makes the +tables 4 generates from answer to Atlassian's schema, 4 proves 12, 13 spells 11's gaps in 12's +grammar, and 12 rewrites code 4b and 4c change; then 14 moves the files 15, 16 and 10 edit and HTML +is written against that layout, 4d marks the gate legs before 17 adds one, 17 puts the complexity +guardrail under the largest body of new code, and 5f and 5g read last because 7 is what changes the +bundle size and the tagline. - [x] **0 — Scaffold.** - [x] **1a — The directive grammar.** @@ -60,33 +68,8 @@ proves 12, 13 spells 11's gaps in 12's grammar, and 12 rewrites code 4b and 4c c - [x] **4.3 — The markdown property.** - [x] **4.4 — The real payloads.** - [x] **4b — The block walk's retry (`0.2.0`).** -- [ ] **4c — The scanning rule's remaining sites (`0.2.0`).** A trailing-anchored regex re-walks - its run from every start position, so an interior whitespace run costs quadratic time rather - than linear — 3h measured 80k spaces inside an ATX heading at 11.3s, and 3ms once the walk - replaced the regex. Three sites the same sweep did not reach: `normalizeLabel` in - `link-syntax.ts`, whose shortcut-reference input is `scan.source.slice(...)` rather than the - 999-capped `readLabel` value, and two in `emit/inline-line.ts`. The fix is the one 3h used — - an index walk, `trimTrailingSpace` where the ends match. A fourth of another shape joins - them: `readNestedDirective` restarts its depth counter per level, so each parse level - re-scans the region below it and nested inline directives cost O(depth × content) — 3f's - cost, which 3i's slot parse doubles rather than changes in class, bounded by the 500-level - guard. A fifth predates 12c: the list-item walk re-scans the rest of a line once per item - level — `isThematicBreak` in `containerStart` on an opener line, `isBlankLine` and - `leadingColumns` in `continuesContainer` on a continuation line, and a blank line - continues every open item without consuming input; 30000 nested items take 4.4s at 59 KB - (the stability-reviewer, 2026-09-16). §11's scanning rule is - the whole argument; the pipeline persona feeds documents nobody typed. - `readDirectiveContent`'s scan splits into named steps with that fix rather than keeping its - complexity (the maintainer, 2026-09-16). A sixth 4b leaves behind: the parser asks - `commonMarkSpelling` at every directive-spelled list it reads, and the answer spells the whole - subtree below, itself quadratic in the depth left, so nested directive lists cost about the - cube of their depth — 250 rule-first levels parse in 1.3 s at 16.5 kB, 1.5 MB of that shape - at 250 levels in 0.6 s — bounded by the depth guard like `readNestedDirective` (the - maintainer, 2026-09-18). A seventh is a throw rather than a cost: `adfDocumentFault` pushes a - node's content with a spread, so past about 125k sibling nodes the guard throws a - `RangeError` where §11 owes a `Result` — a loop over the content closes it (the - stability-reviewer and the maintainer, 2026-09-18). -- [ ] **4d — What the gate says while it runs (`0.2.1`).** `ci.sh` runs nine legs and announces +- [x] **4c — The scanning rule's remaining sites (`0.2.0`).** +- [ ] **4d — What the gate says while it runs (`0.2.0`).** `ci.sh` runs nine legs and announces none of them, so five minutes of a Gitea run read as silence and a hang cannot be told from a slow pull — the maintainer hit exactly this on the `0.1.0` release. Three causes, each its own fix. The legs need markers: `plainpages`' `ci.sh` prints a `step()` header per leg and @@ -113,7 +96,7 @@ proves 12, 13 spells 11's gaps in 12's grammar, and 12 rewrites code 4b and 4c c and is the trade to weigh rather than discover on a red release run. **Settled** (the maintainer, 2026-09-13): last of the known work, clear of `0.2.0`, placed there knowing the cutoff may land before `0.2.0` ships. -- [ ] **5f — Publish the bundle size (`0.2.1`).** Measure the shipped artifact and put the number in the +- [ ] **5f — Publish the bundle size (`0.2.0`).** Measure the shipped artifact and put the number in the README, kept honest by the release pipeline rather than by a human re-reading it. The quantity is what a consumer downloads and loads: the tarball `npm pack` produces, its unpacked `dist`, and the built JavaScript minified + gzipped — the figure the competitors @@ -131,10 +114,10 @@ proves 12, 13 spells 11's gaps in 12's grammar, and 12 rewrites code 4b and 4c c "why" note left. The top follows the package-README order: an npm version badge and the Gitea Actions badge, a tagline that is also `package.json`'s `description`, a feature list and a one-line table of contents, then install and the shortest runnable example; a table of - everything exported sits near the bottom. The HTML directions are one aside line under the API - until `0.3.0` ships them, the `// 0.3.0` signatures and the `0.3.0` guarantee going until then. - The tagline and `description` read "Lossless conversion between Atlassian Document Format and - extended markdown" until 7 restores HTML. + everything exported sits near the bottom. The HTML directions were to stay an aside until a + later release shipped them; 7 now ships in this one and reads ahead of this item, so the + README documents HTML as it documents markdown, the tagline and `description` naming both + (the maintainer, 2026-09-13, revised 2026-09-18). - [x] **5a — Rename to `@larvit/adf-codec`.** - [x] **5b — The consumer's error surface.** - [x] **5b1 — The error's source position.** @@ -143,11 +126,11 @@ proves 12, 13 spells 11's gaps in 12's grammar, and 12 rewrites code 4b and 4c c - [x] **5b4 — The README's consumer surface.** - [x] **5c — The build and the release pipeline.** - [x] **5d — The browser leg.** -- [ ] **6 — The HTML dialect spec (`0.3.0`).** Element-by-element mapping, the `data-*` fidelity +- [ ] **6 — The HTML dialect spec (`0.2.0`).** Element-by-element mapping, the `data-*` fidelity scheme, the opaque-carry form, and the documented foreign-element set `htmlToAdf` accepts. -- [ ] **7 — HTML, ship `0.3.0`.** `adfToHtml`, `htmlToAdf`, the composed `markdownToHtml` / - `htmlToMarkdown`. CommonMark spec suite runs against `markdownToHtml` from here (§10). The - README's tagline and `package.json`'s `description` regain HTML (5g). +- [ ] **7 — HTML, the third format (`0.2.0`).** `adfToHtml`, `htmlToAdf`, the composed + `markdownToHtml` / `htmlToMarkdown`. CommonMark spec suite runs against `markdownToHtml` from + here (§10). The README's tagline and `package.json`'s `description` regain HTML (5g). - [ ] **8 — CLI.** A later goal, shaped around the personas once the library exists. - [ ] **9 — The online sandbox.** A web page with two textboxes converting back and forth between ADF and markdown, powered by the library's browser build. - [ ] **10 — Lossy conversion (`0.2.0`).** Markdown other tools render readably, to and from ADF, @@ -242,10 +225,30 @@ proves 12, 13 spells 11's gaps in 12's grammar, and 12 rewrites code 4b and 4c c brackets stay literal text, CommonMark's rule that no link holds another — rather than dropping the outer link silently as `closeLink`'s `applyMark` does today, with a normalization fixture per shape (the stability-reviewer, 2026-09-16; the maintainer, 2026-09-17). -- [ ] **17 — A machine-enforced size guardrail.** Add a per-function complexity check to the gate — - branch count or size — so the fits-in-your-head guardrail fails the build rather than - waiting for a review to catch it (the systems-architect, 2026-09-16); placed after `0.3.0` - (the maintainer, 2026-09-17). +- [ ] **17 — A machine-enforced size guardrail (`0.2.0`).** Add a per-function complexity check to + the gate — branch count or size — so the fits-in-your-head guardrail fails the build rather + than waiting for a review to catch it (the systems-architect, 2026-09-16). It reads ahead of + 6, 7 and 10 so the largest body of new code is written under it, which is also what decides + the threshold: today's worst is `readDirectiveContent`, 27 lines and about 12 decision points + over four concerns in one loop — escape, code span, nested directive, bracket balance — which + 4c left half-split and this item either passes or forces apart (the systems-architect and the + maintainer, 2026-09-18). +- [ ] **18 — The subtree the directive spelling asks about (`0.2.0`).** The parser asks + `commonMarkSpelling` at every directive-spelled block and the answer emits the whole subtree + below, so a node at depth d is spelled d times: three nested rule-first directive lists cost + 18 asks over 10 nodes, and 250 levels parse in 1.2 s at 16.4 kB, 4.9 s at 261 kB with a + kilobyte of content per level. The depth guard bounds the levels at about 250, never the + content, so this is the pipeline persona's hang on an input nobody typed (§11). Keeping each + child's emitted result for its parent's ask is not a straight handover: the same node object + is asked at different depths — 4, 3 and 2 for the innermost list of three — because the + parser counts a list and its item as two levels where the emitter's readable list counts one + (4b), and `headroom` is that guard's slack. The parts that survive the measurement: the paths + agree, `text` and `spelling` carry no depth, `headroom` is affine in it, and the parser asks + first at the deepest of them, so a kept result rebases by the difference. Either rebase and + record that argument in `AGENTS.md`, or give both directions one list accounting so a node + has one depth and nothing needs rebasing — which reopens 4b. A single post-build walk was + rejected: it reports the outer offender where the build reports the inner one (the + maintainer, 2026-09-18). ## The ADF inventory to cover