Model CommonMark's emphasis matching, and escape the delimiter that only closes #25

Merged
lilleman merged 3 commits from tick-2e5 into main 2026-08-27 13:17:36 +02:00
18 changed files with 1016 additions and 71 deletions
+6
View File
@@ -119,6 +119,12 @@ live Atlassian APIs; property-generated ADF trees; the CommonMark spec suite aga
the spot. the spot.
- Only the hard break's inline segment holds a raw newline — every other spelling escapes one or - Only the hard break's inline segment holds a raw newline — every other spelling escapes one or
refuses it — which is how the whitespace carry finds a line edge. refuses it — which is how the whitespace carry finds a line edge.
- Emphasis is spelled against CommonMark's matching, never flanking alone: a delimiter run in text
escapes wherever CommonMark could open or close with it, leaving the emitter's own delimiters the
only ones in play, and a pair that matching hands to another delimiter rides the carry instead.
`matchEmphasis` transcribes the reference `process_emphasis` line for line, and its closer walk and
opener search stay whole: broken into named steps they drift from the algorithm being faithful is
the whole point of.
- A readable spelling tried ahead of a general one — the image, the pipe table, a pipe cell — - A readable spelling tried ahead of a general one — the image, the pipe table, a pipe cell —
returns `string | undefined`, never a `Result`: any failure is the fallback signal, and the returns `string | undefined`, never a `Result`: any failure is the fallback signal, and the
general form owns the refusal. Refusing there refuses a document the general form spells. general form owns the refusal. Refusing there refuses a document the general form spells.
@@ -1,6 +1,6 @@
\~~not strike~~ \~~not strike\~~
~~gone~~ but \~~kept~~ ~~gone~~ but \~~kept\~~
A ~ B ~ C A ~ B ~ C
@@ -0,0 +1,161 @@
{
"content": [
{
"content": [
{
"attrs": {
"width": 50
},
"content": [
{
"attrs": {
"panelType": "info"
},
"content": [
{
"content": [
{
"text": "Check the collation before importing.",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "panel"
}
],
"type": "layoutColumn"
},
{
"attrs": {
"width": 50
},
"content": [
{
"attrs": {
"localId": "0198f3a2-7c41-7f2e-9b3a-4d8e2c1a6b90"
},
"content": [
{
"attrs": {
"localId": "0198f3a2-8d52-70b1-8c4f-5e9f3d2b7ca1",
"state": "DONE"
},
"content": [
{
"text": "Write the spec",
"type": "text"
}
],
"type": "taskItem"
},
{
"attrs": {
"localId": "0198f3a2-9e63-7d80-a15b-6fa04e3c8db2",
"state": "TODO"
},
"content": [
{
"text": "Ship it",
"type": "text"
}
],
"type": "taskItem"
}
],
"type": "taskList"
}
],
"type": "layoutColumn"
}
],
"type": "layoutSection"
},
{
"attrs": {
"layout": "center",
"width": 50
},
"content": [
{
"attrs": {
"collection": "MediaServicesSample",
"id": "4478e39c-cf9b-41d1-ba92-68589487cd75",
"type": "file"
},
"type": "media"
},
{
"content": [
{
"text": "The moon, at night.",
"type": "text"
}
],
"type": "caption"
}
],
"type": "mediaSingle"
},
{
"attrs": {
"title": "Full build log"
},
"content": [
{
"attrs": {
"isNumberColumnEnabled": true
},
"content": [
{
"content": [
{
"content": [
{
"content": [
{
"text": "Step",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "tableHeader"
}
],
"type": "tableRow"
},
{
"content": [
{
"attrs": {
"background": "#deebff"
},
"content": [
{
"content": [
{
"text": "Compile",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "tableCell"
}
],
"type": "tableRow"
}
],
"type": "table"
}
],
"type": "expand"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,39 @@
::::::layoutSection
::::layoutColumn {width=50}
:::panel info
Check the collation before importing.
:::
::::
:::::layoutColumn {width=50}
::::taskList {localId=0198f3a2-7c41-7f2e-9b3a-4d8e2c1a6b90}
:::taskItem DONE {localId=0198f3a2-8d52-70b1-8c4f-5e9f3d2b7ca1}
Write the spec
:::
:::taskItem TODO {localId=0198f3a2-9e63-7d80-a15b-6fa04e3c8db2}
Ship it
:::
::::
:::::
::::::
::::mediaSingle {layout=center width=50}
::media {collection=MediaServicesSample id=4478e39c-cf9b-41d1-ba92-68589487cd75 type=file}
:::caption
The moon, at night.
:::
::::
::::::expand {title="Full build log"}
:::::table {isNumberColumnEnabled=true}
::::tableRow
:::tableHeader
Step
:::
::::
::::tableRow
:::tableCell {background="#deebff"}
Compile
:::
::::
:::::
::::::
@@ -0,0 +1,312 @@
{
"content": [
{
"attrs": {
"level": 1
},
"content": [
{
"text": "Release 2.4",
"type": "text"
}
],
"type": "heading"
},
{
"content": [
{
"text": "Shipped ",
"type": "text"
},
{
"attrs": {
"shortName": ":rocket:",
"text": "🚀"
},
"type": "emoji"
},
{
"text": " on ",
"type": "text"
},
{
"attrs": {
"timestamp": "1756080000000"
},
"type": "date"
},
{
"text": " — ",
"type": "text"
},
{
"attrs": {
"id": "01a032c3-7a7c-775f-a730-2d79351338b4",
"text": "@Mikael"
},
"type": "mention"
},
{
"text": " owns the rollout, status ",
"type": "text"
},
{
"attrs": {
"color": "yellow",
"text": "In review"
},
"type": "status"
},
{
"text": ".",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"content": [
{
"content": [
{
"content": [
{
"text": "Part",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "tableHeader"
},
{
"content": [
{
"content": [
{
"text": "Qty",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "tableHeader"
}
],
"type": "tableRow"
},
{
"content": [
{
"content": [
{
"content": [
{
"text": "Bolt M8",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "tableCell"
},
{
"content": [
{
"content": [
{
"text": "40",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "tableCell"
}
],
"type": "tableRow"
},
{
"content": [
{
"content": [
{
"content": [
{
"marks": [
{
"attrs": {
"href": "https://example.com/washer"
},
"type": "link"
}
],
"text": "Washer",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "tableCell"
},
{
"content": [
{
"content": [
{
"text": "12",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "tableCell"
}
],
"type": "tableRow"
}
],
"type": "table"
},
{
"content": [
{
"content": [
{
"content": [
{
"text": "Torque the ",
"type": "text"
},
{
"marks": [
{
"type": "strong"
}
],
"text": "M8 bolt",
"type": "text"
},
{
"text": " to ",
"type": "text"
},
{
"marks": [
{
"type": "em"
}
],
"text": "25 Nm",
"type": "text"
},
{
"text": ".",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "listItem"
},
{
"content": [
{
"content": [
{
"text": "Check the collation:",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"content": [
{
"content": [
{
"marks": [
{
"type": "code"
}
],
"text": "mysqldump --default-character-set=utf8mb4",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "listItem"
}
],
"type": "bulletList"
}
],
"type": "listItem"
}
],
"type": "bulletList"
},
{
"content": [
{
"content": [
{
"text": "Rolled back once, see the ",
"type": "text"
},
{
"marks": [
{
"type": "underline"
}
],
"text": "postmortem",
"type": "text"
},
{
"text": ".",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "blockquote"
},
{
"type": "rule"
},
{
"attrs": {
"panelType": "warning"
},
"content": [
{
"content": [
{
"text": "Do not skip the pre-flight.",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "panel"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,20 @@
# Release 2.4
Shipped :emoji[🚀]{shortName=":rocket:"} on :date{timestamp=1756080000000} — :mention[@Mikael]{id=01a032c3-7a7c-775f-a730-2d79351338b4} owns the rollout, status :status[In review]{color=yellow}.
| Part | Qty |
| --- | --- |
| Bolt M8 | 40 |
| [Washer](https://example.com/washer) | 12 |
- Torque the **M8 bolt** to _25 Nm_.
- Check the collation:
- `mysqldump --default-character-set=utf8mb4`
> Rolled back once, see the :underline[postmortem].
---
:::panel warning
Do not skip the pre-flight.
:::
@@ -0,0 +1,56 @@
{
"content": [
{
"content": [
{
"text": "un",
"type": "text"
},
{
"marks": [
{
"type": "em"
}
],
"text": "a* b",
"type": "text"
},
{
"text": "istic",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"marks": [
{
"type": "strike"
}
],
"text": "a~~ b",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"marks": [
{
"type": "em"
}
],
"text": "a_ b",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,5 @@
un*a\* b*istic
~~a\~~ b~~
_a\_ b_
@@ -0,0 +1,149 @@
{
"content": [
{
"content": [
{
"text": "un",
"type": "text"
},
{
"marks": [
{
"type": "em"
}
],
"text": "a",
"type": "text"
},
{
"marks": [
{
"type": "em"
},
{
"type": "strong"
}
],
"text": "b",
"type": "text"
},
{
"marks": [
{
"type": "strong"
}
],
"text": "c",
"type": "text"
},
{
"text": "istic",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"marks": [
{
"type": "strong"
},
{
"type": "em"
}
],
"text": "a",
"type": "text"
},
{
"marks": [
{
"type": "strong"
}
],
"text": "b",
"type": "text"
},
{
"marks": [
{
"type": "strong"
},
{
"type": "em"
}
],
"text": "c",
"type": "text"
}
],
"type": "paragraph"
},
{
"content": [
{
"marks": [
{
"type": "strike"
}
],
"text": "un",
"type": "text"
},
{
"marks": [
{
"type": "strike"
},
{
"type": "em"
}
],
"text": "a",
"type": "text"
},
{
"marks": [
{
"type": "strike"
},
{
"type": "em"
},
{
"type": "strong"
}
],
"text": "b",
"type": "text"
},
{
"marks": [
{
"type": "strike"
},
{
"type": "strong"
}
],
"text": "c",
"type": "text"
},
{
"marks": [
{
"type": "strike"
}
],
"text": "istic",
"type": "text"
}
],
"type": "paragraph"
}
],
"type": "doc",
"version": 1
}
@@ -0,0 +1,5 @@
un:adf{json="{\"marks\":[{\"type\":\"em\"}],\"text\":\"a\",\"type\":\"text\"}"}:adf{json="{\"marks\":[{\"type\":\"em\"},{\"type\":\"strong\"}],\"text\":\"b\",\"type\":\"text\"}"}**c**istic
***a*b**:adf{json="{\"marks\":[{\"type\":\"strong\"},{\"type\":\"em\"}],\"text\":\"c\",\"type\":\"text\"}"}
~~un~~:adf{json="{\"marks\":[{\"type\":\"strike\"},{\"type\":\"em\"}],\"text\":\"a\",\"type\":\"text\"}"}:adf{json="{\"marks\":[{\"type\":\"strike\"},{\"type\":\"em\"},{\"type\":\"strong\"}],\"text\":\"b\",\"type\":\"text\"}"}~~**c**istic~~
@@ -6,4 +6,4 @@
snake_case_name snake_case_name
\*not emphasis* \*not emphasis\*
+7 -4
View File
@@ -39,9 +39,10 @@ normalizes to it through the round-trip.
- Paragraphs on one line — no soft wrapping; a soft line break in input becomes a single space. - Paragraphs on one line — no soft wrapping; a soft line break in input becomes a single space.
- Entity references in input decode to their characters; output backslash-escapes only where text - Entity references in input decode to their characters; output backslash-escapes only where text
would otherwise parse as syntax, scanning the assembled line rather than each text node: escape would otherwise parse as syntax, scanning the assembled line rather than each text node: escape
the leading delimiter of a construct that would otherwise open, re-scan from there, and the leading delimiter of a construct that would otherwise open, re-scan from there, and repeat.
repeat — with the opener literal the closer parses as text, so `*not emphasis*` is An emphasis delimiter run in text escapes where CommonMark can open **or** close with it, so
`\*not emphasis*`, one backslash. `*not emphasis*` is `\*not emphasis\*` — no delimiter the emitter did not write reaches the
matching below, which is what lets the emitter decide its own pairings.
- Blocks separated by one blank line at document level, inside a blockquote and between CommonMark - Blocks separated by one blank line at document level, inside a blockquote and between CommonMark
blocks; two directive blocks inside a container take none. No trailing whitespace outside a code blocks; two directive blocks inside a container take none. No trailing whitespace outside a code
block's block's
@@ -390,7 +391,9 @@ An inline node whose marks no nesting spells — a mark type not listed here, an
spelling does not list, a value that is not the spelling's type, an attribute the spelling needs spelling does not list, a value that is not the spelling's type, an attribute the spelling needs
and the mark lacks, an order putting a code span outside another mark, `code` over anything but a and the mark lacks, an order putting a code span outside another mark, `code` over anything but a
text node or over text holding a newline, or a spelling CommonMark's flanking rules cannot open or text node or over text holding a newline, or a spelling CommonMark's flanking rules cannot open or
close where the run sits (`un**-real**istic`) — rides the inline carry whole. A value the spelling close where the run sits (`un**-real**istic`), or one CommonMark's matching pairs elsewhere — the
intra-word `*` runs together with a neighbouring `**`, and the multiple-of-3 rule can leave the
merged run's pairing to another delimiter — rides the inline carry whole. A value the spelling
holds but CommonMark cannot write — a link destination or title — is a named error instead. An holds but CommonMark cannot write — a link destination or title — is a named error instead. An
opaque carry inside a mark spelling is a named error in input: the carry restores its node opaque carry inside a mark spelling is a named error in input: the carry restores its node
exactly, marks included (AGENTS.md §3). exactly, marks included (AGENTS.md §3).
+23 -4
View File
@@ -197,12 +197,12 @@ test('escapes only text that would otherwise open a construct', () => {
assert.equal(emitted(':::panel info'), '\\:::panel info\n') assert.equal(emitted(':::panel info'), '\\:::panel info\n')
assert.equal(emitted('10:30 tomorrow'), '10:30 tomorrow\n') assert.equal(emitted('10:30 tomorrow'), '10:30 tomorrow\n')
assert.equal(emitted('[a](b)'), '\\[a](b)\n') assert.equal(emitted('[a](b)'), '\\[a](b)\n')
assert.equal(emitted('**bold**'), '\\*\\*bold**\n') assert.equal(emitted('**bold**'), '\\*\\*bold\\*\\*\n')
assert.equal(emitted('a `x` b'), 'a \\`x` b\n') assert.equal(emitted('a `x` b'), 'a \\`x` b\n')
assert.equal(emitted('~~struck~~'), '\\~~struck~~\n') assert.equal(emitted('~~struck~~'), '\\~~struck\\~~\n')
assert.equal(emitted('a \\* b'), 'a \\\\* b\n') assert.equal(emitted('a \\* b'), 'a \\\\\\* b\n')
assert.equal(emitted('1. not a list'), '1\\. not a list\n') assert.equal(emitted('1. not a list'), '1\\. not a list\n')
assert.equal(emitted('*"quoted"*'), '\\*"quoted"*\n') assert.equal(emitted('*"quoted"*'), '\\*"quoted"\\*\n')
assert.equal(emitted('x"_y"'), 'x"\\_y"\n') assert.equal(emitted('x"_y"'), 'x"\\_y"\n')
}) })
@@ -297,6 +297,25 @@ test('escapes a literal delimiter that would merge with an emitted one', () => {
assert.equal(emitted(marked('x', { attrs: { href: 'https://example.com/' }, type: 'link' }), { text: '{}', type: 'text' }), '[x](https://example.com/){}\n') assert.equal(emitted(marked('x', { attrs: { href: 'https://example.com/' }, type: 'link' }), { text: '{}', type: 'text' }), '[x](https://example.com/){}\n')
}) })
test('escapes a literal delimiter run that only closes', () => {
const emitted = (text: string): string => markdown(adfToMarkdown(document(paragraph({ text, type: 'text' }))))
assert.equal(emitted('a* b'), 'a\\* b\n')
assert.equal(emitted('2 * 3'), '2 * 3\n')
})
test("spells a mark run CommonMark's matching pairs as written", () => {
const marked = (text: string, ...marks: AdfMark[]): AdfNode => ({ marks, text, type: 'text' })
const emitted = (...content: AdfNode[]): string => markdown(adfToMarkdown(document(paragraph(...content))))
const em: AdfMark = { type: 'em' }
const strong: AdfMark = { type: 'strong' }
assert.equal(emitted({ text: 'un', type: 'text' }, marked('a', em, strong), { text: 'istic', type: 'text' }), 'un***a***istic\n')
assert.equal(emitted({ text: 're', type: 'text' }, marked('structure', strong), { text: ' the code', type: 'text' }), 're**structure** the code\n')
assert.equal(
emitted({ text: 'un', type: 'text' }, marked('a', em), marked('b', em, strong), marked('c', em), { text: 'istic', type: 'text' }),
'un*a**b**c*istic\n',
)
})
test('escapes a hyphen underline a hard break would expose', () => { test('escapes a hyphen underline a hard break would expose', () => {
const line = (text: string): string => markdown(adfToMarkdown(document(paragraph({ text: 'foo', type: 'text' }, { type: 'hardBreak' }, { text, type: 'text' })))) const line = (text: string): string => markdown(adfToMarkdown(document(paragraph({ text: 'foo', type: 'text' }, { type: 'hardBreak' }, { text, type: 'text' }))))
assert.equal(line('--'), 'foo\\\n\\--\n') assert.equal(line('--'), 'foo\\\n\\--\n')
+29
View File
@@ -66,6 +66,35 @@ for (const directory of emittingDirectories) {
} }
} }
function roundTripFixtures(): { name: string; path: string }[] {
return emittingDirectories.flatMap((directory) =>
fixtureNames(directory, '.json').map((name) => ({ name: `${directory}/${name}`, path: join(roundTripRoot, directory, `${name}.json`) })),
)
}
test('no two round-trip documents share one markdown spelling', () => {
const spellings = new Map<string, string>()
for (const fixture of roundTripFixtures()) {
const parsed: unknown = JSON.parse(readFileSync(fixture.path, 'utf8'))
assert.ok(isAdfDocument(parsed), `${fixture.name} is not an ADF document`)
const result = adfToMarkdown(parsed)
assert.ok(result.ok, result.ok ? '' : `${result.error.code}: ${result.error.message}`)
assert.equal(spellings.get(result.value), undefined, `${fixture.name} and ${spellings.get(result.value)} share one markdown spelling`)
spellings.set(result.value, fixture.name)
}
})
test('no round-trip fixture repeats the document another holds', () => {
const documents = new Map<string, string>()
for (const fixture of roundTripFixtures()) {
const parsed: unknown = JSON.parse(readFileSync(fixture.path, 'utf8'))
assert.ok(isJsonValue(parsed), `${fixture.name} does not hold a JSON value`)
const document = serializeCanonicalJson(parsed, 'compact')
assert.equal(documents.get(document), undefined, `${fixture.name} repeats the document ${documents.get(document)} holds`)
documents.set(document, fixture.name)
}
})
// spec/flavour.md, Directives: the container fence rule, checked against the emitted bytes. // spec/flavour.md, Directives: the container fence rule, checked against the emitted bytes.
function fenceNestingFault(markdown: string): string | undefined { function fenceNestingFault(markdown: string): string | undefined {
const open: number[] = [] const open: number[] = []
+116
View File
@@ -0,0 +1,116 @@
import { isUnicodeWhitespace } from './commonmark-grammar.ts'
type DelimiterRun = { canClose: boolean; canOpen: boolean; character: string; length: number }
type EmphasisPairing<Run> = { closer: Run; closerOffset: number; opener: Run; openerOffset: number; used: number }
type Candidate<Run> = {
head: number
next: Candidate<Run> | undefined
original: number
previous: Candidate<Run> | undefined
remaining: number
run: Run
tail: number
}
const unicodePunctuation = /[\p{P}\p{S}]/u
export function delimiterFlags(character: string, before: string, after: string): { canClose: boolean; canOpen: boolean } {
const left = isLeftFlanking(before, after)
const right = isRightFlanking(before, after)
if (character !== '_') return { canClose: right, canOpen: left }
return { canClose: right && (!left || isPunctuation(after)), canOpen: left && (!right || isPunctuation(before)) }
}
export function isWordCharacter(character: string): boolean {
return character !== '' && !isWhitespace(character) && !isPunctuation(character)
}
export function matchEmphasis<Run extends DelimiterRun>(runs: readonly Run[]): EmphasisPairing<Run>[] {
const pairings: EmphasisPairing<Run>[] = []
const bottoms = new Map<string, Candidate<Run> | undefined>()
let closer = candidates(runs)
while (closer !== undefined) {
if (!closer.run.canClose) {
closer = closer.next
continue
}
const key = `${closer.run.character}${closer.run.canOpen}${closer.original % 3}`
const bottom = bottoms.get(key)
let opener = closer.previous
while (opener !== undefined && opener !== bottom && !pairs(opener, closer)) opener = opener.previous
if (opener === undefined || opener === bottom) {
bottoms.set(key, closer.previous)
const following = closer.next
if (!closer.run.canOpen) unlink(closer)
closer = following
continue
}
const used = closer.remaining >= 2 && opener.remaining >= 2 ? 2 : 1
opener.remaining -= used
opener.tail -= used
pairings.push({ closer: closer.run, closerOffset: closer.head, opener: opener.run, openerOffset: opener.tail, used })
closer.head += used
closer.remaining -= used
opener.next = closer
closer.previous = opener
if (opener.remaining === 0) unlink(opener)
if (closer.remaining > 0) continue
const following = closer.next
unlink(closer)
closer = following
}
return pairings
}
function candidates<Run extends DelimiterRun>(runs: readonly Run[]): Candidate<Run> | undefined {
let first: Candidate<Run> | undefined
let previous: Candidate<Run> | undefined
for (const run of runs) {
const candidate: Candidate<Run> = {
head: 0,
next: undefined,
original: run.length,
previous,
remaining: run.length,
run,
tail: run.length,
}
if (previous === undefined) first = candidate
else previous.next = candidate
previous = candidate
}
return first
}
function pairs<Run extends DelimiterRun>(opener: Candidate<Run>, closer: Candidate<Run>): boolean {
if (!opener.run.canOpen || opener.run.character !== closer.run.character) return false
const odd = (closer.run.canOpen || opener.run.canClose) && closer.original % 3 !== 0 && (opener.original + closer.original) % 3 === 0
return !odd
}
function unlink<Run>(candidate: Candidate<Run>): void {
if (candidate.previous !== undefined) candidate.previous.next = candidate.next
if (candidate.next !== undefined) candidate.next.previous = candidate.previous
}
function isLeftFlanking(before: string, after: string): boolean {
if (isWhitespace(after)) return false
if (!isPunctuation(after)) return true
return isWhitespace(before) || isPunctuation(before)
}
function isRightFlanking(before: string, after: string): boolean {
if (isWhitespace(before)) return false
if (!isPunctuation(before)) return true
return isWhitespace(after) || isPunctuation(after)
}
function isPunctuation(character: string): boolean {
return character !== '' && unicodePunctuation.test(character)
}
function isWhitespace(character: string): boolean {
return character === '' || isUnicodeWhitespace(character)
}
+64 -51
View File
@@ -1,4 +1,5 @@
import { escapesLineClaim, isUnicodeWhitespace, opensBracketedAutolink, startsEntityReference, type LinePosition } from './commonmark-grammar.ts' import { escapesLineClaim, opensBracketedAutolink, startsEntityReference, type LinePosition } from './commonmark-grammar.ts'
import { delimiterFlags, isWordCharacter, matchEmphasis } from './emphasis-matching.ts'
export type EmphasisRole = 'close' | 'open' export type EmphasisRole = 'close' | 'open'
@@ -14,7 +15,9 @@ export type AssembledLine = { line: string; unspellableRun: NodeRange | undefine
export type LineContainer = 'heading' | 'paragraph' | 'table-cell' export type LineContainer = 'heading' | 'paragraph' | 'table-cell'
type DelimiterRun = { character: string; closeNodes: NodeRange | undefined; end: number; openNodes: NodeRange | undefined; start: number } type EmittedDelimiter = { closes: boolean; offset: number; pair: number; width: number }
type EmittedRun = { canClose: boolean; canOpen: boolean; character: string; delimiters: EmittedDelimiter[]; length: number; start: number }
const delimiters = ['*', '_', '`', '~'] const delimiters = ['*', '_', '`', '~']
@@ -22,7 +25,6 @@ const asciiPunctuation = /[!"#$%&'()*+,\-./:;<=>?@[\\\]^_`{|}~]/
const htmlConstructs = [/^<[!?]/, /^<\/?[A-Za-z][A-Za-z0-9-]*(?:[\s/>]|$)/, /^<[^\s<>@]+@[^\s<>@]+>/] const htmlConstructs = [/^<[!?]/, /^<\/?[A-Za-z][A-Za-z0-9-]*(?:[\s/>]|$)/, /^<[^\s<>@]+@[^\s<>@]+>/]
const inlineDirectiveOpener = /^:[a-z][A-Za-z0-9]*[[{]/ const inlineDirectiveOpener = /^:[a-z][A-Za-z0-9]*[[{]/
const followsLinkText = /[([:]/ const followsLinkText = /[([:]/
const unicodePunctuation = /[\p{P}\p{S}]/u
export function assembleInlineLine(segments: readonly InlineSegment[], container: LineContainer): AssembledLine { export function assembleInlineLine(segments: readonly InlineSegment[], container: LineContainer): AssembledLine {
return escape(resolveEmphasis(segments), container) return escape(resolveEmphasis(segments), container)
@@ -78,40 +80,78 @@ function escape(segments: readonly InlineSegment[], container: LineContainer): A
} }
function unspellableRun(segments: readonly InlineSegment[], output: string, placements: readonly number[]): NodeRange | undefined { function unspellableRun(segments: readonly InlineSegment[], output: string, placements: readonly number[]): NodeRange | undefined {
for (const run of delimiterRuns(segments, placements)) { const { nodes, runs } = emittedRuns(segments, placements, output)
const before = charAt(output, run.start - 1) const pair = misflanked(runs) ?? unpaired(runs)
const after = output.charAt(run.end) return pair === undefined ? undefined : nodes[pair]
if (run.openNodes !== undefined && !isLeftFlanking(before, after)) return run.openNodes }
if (run.closeNodes !== undefined && !isRightFlanking(before, after)) return run.closeNodes
function misflanked(runs: readonly EmittedRun[]): number | undefined {
for (const run of runs) {
for (const delimiter of run.delimiters) {
if (!(delimiter.closes ? run.canClose : run.canOpen)) return delimiter.pair
}
} }
return undefined return undefined
} }
function delimiterRuns(segments: readonly InlineSegment[], placements: readonly number[]): DelimiterRun[] { function unpaired(runs: readonly EmittedRun[]): number | undefined {
const runs: DelimiterRun[] = [] const matched = new Set<number>()
for (const pairing of matchEmphasis(runs)) {
const opened = delimiterAt(pairing.opener, false, pairing.openerOffset, pairing.used)
const closed = delimiterAt(pairing.closer, true, pairing.closerOffset, pairing.used)
if (opened !== undefined && closed !== undefined && opened.pair === closed.pair) matched.add(opened.pair)
}
// The last opener left unpaired is the innermost: the smallest carry that changes the line.
let innermost: number | undefined
for (const run of runs) {
for (const delimiter of run.delimiters) {
if (!delimiter.closes && !matched.has(delimiter.pair)) innermost = delimiter.pair
}
}
return innermost
}
function delimiterAt(run: EmittedRun, closes: boolean, offset: number, width: number): EmittedDelimiter | undefined {
return run.delimiters.find((delimiter) => delimiter.closes === closes && delimiter.offset === offset && delimiter.width === width)
}
function emittedRuns(segments: readonly InlineSegment[], placements: readonly number[], output: string): { nodes: NodeRange[]; runs: EmittedRun[] } {
const runs: EmittedRun[] = []
const nodes: NodeRange[] = []
const open: number[] = []
let cursor = 0 let cursor = 0
for (const segment of segments) { for (const segment of segments) {
const start = placements[cursor] ?? 0 const start = placements[cursor] ?? 0
cursor += segment.text.length cursor += segment.text.length
if (segment.emphasis === undefined) continue if (segment.emphasis === undefined) continue
const closes = segment.emphasis === 'close' const closes = segment.emphasis === 'close'
const end = start + segment.text.length const pair = closes ? (open.pop() ?? nodes.length) : nodes.length
if (!closes) {
nodes.push(segment.nodes)
open.push(pair)
}
const width = segment.text.length
const previous = runs[runs.length - 1] const previous = runs[runs.length - 1]
if (previous !== undefined && previous.end === start && previous.character === segment.text.charAt(0)) { if (previous !== undefined && previous.start + previous.length === start && previous.character === segment.text.charAt(0)) {
previous.closeNodes = previous.closeNodes ?? (closes ? segment.nodes : undefined) previous.delimiters.push({ closes, offset: start - previous.start, pair, width })
previous.end = end previous.length += width
previous.openNodes = previous.openNodes ?? (closes ? undefined : segment.nodes)
continue continue
} }
runs.push({ runs.push({
canClose: false,
canOpen: false,
character: segment.text.charAt(0), character: segment.text.charAt(0),
closeNodes: closes ? segment.nodes : undefined, delimiters: [{ closes, offset: 0, pair, width }],
end, length: width,
openNodes: closes ? undefined : segment.nodes,
start, start,
}) })
} }
return runs for (const run of runs) {
const flags = delimiterFlags(run.character, charAt(output, run.start - 1), output.charAt(run.start + run.length))
run.canClose = flags.canClose
run.canOpen = flags.canOpen
}
return { nodes, runs }
} }
function mergesWithSyntax(scan: string, escapings: readonly (InlineEscaping | undefined)[], index: number): boolean { function mergesWithSyntax(scan: string, escapings: readonly (InlineEscaping | undefined)[], index: number): boolean {
@@ -177,7 +217,7 @@ function claimsCharacter(
if (character === ':') return inlineDirectiveOpener.test(rest) if (character === ':') return inlineDirectiveOpener.test(rest)
if (character === '[') return opensLink(scan, escapings, index) if (character === '[') return opensLink(scan, escapings, index)
if (character === '`') return opensCodeSpan(scan, index, escaped) if (character === '`') return opensCodeSpan(scan, index, escaped)
if (character === '*' || character === '_' || character === '~') return opensEmphasis(scan, index, escaped) if (character === '*' || character === '_' || character === '~') return claimsEmphasis(scan, index, escaped)
return false return false
} }
@@ -196,16 +236,13 @@ function opensCodeSpan(scan: string, index: number, escaped: ReadonlySet<number>
return new RegExp('(?<!`)`{' + length + '}(?!`)').test(scan.slice(index + length)) return new RegExp('(?<!`)`{' + length + '}(?!`)').test(scan.slice(index + length))
} }
function opensEmphasis(scan: string, index: number, escaped: ReadonlySet<number>): boolean { function claimsEmphasis(scan: string, index: number, escaped: ReadonlySet<number>): boolean {
if (!startsRun(scan, index, escaped)) return false if (!startsRun(scan, index, escaped)) return false
const character = scan.charAt(index) const character = scan.charAt(index)
const length = runLength(scan, index) const length = runLength(scan, index)
const before = index === 0 ? '' : scan.charAt(index - 1) if (character === '~' && length !== 2) return false
const after = scan.charAt(index + length) const flags = delimiterFlags(character, index === 0 ? '' : scan.charAt(index - 1), scan.charAt(index + length))
if (character === '~') return length === 2 && isLeftFlanking(before, after) return flags.canClose || flags.canOpen
if (!isLeftFlanking(before, after)) return false
if (character === '*') return true
return !isRightFlanking(before, after) || isPunctuation(before)
} }
function startsRun(scan: string, index: number, escaped: ReadonlySet<number>): boolean { function startsRun(scan: string, index: number, escaped: ReadonlySet<number>): boolean {
@@ -220,30 +257,6 @@ function runLength(scan: string, index: number): number {
return length return length
} }
function isLeftFlanking(before: string, after: string): boolean {
if (isWhitespace(after)) return false
if (!isPunctuation(after)) return true
return isWhitespace(before) || isPunctuation(before)
}
function isRightFlanking(before: string, after: string): boolean {
if (isWhitespace(before)) return false
if (!isPunctuation(before)) return true
return isWhitespace(after) || isPunctuation(after)
}
function isPunctuation(character: string): boolean {
return character !== '' && unicodePunctuation.test(character)
}
function isWhitespace(character: string): boolean {
return character === '' || isUnicodeWhitespace(character)
}
function isWordCharacter(character: string): boolean {
return character !== '' && !isUnicodeWhitespace(character) && !unicodePunctuation.test(character)
}
function charAt(text: string, index: number): string { function charAt(text: string, index: number): string {
return index < 0 ? '' : text.charAt(index) return index < 0 ? '' : text.charAt(index)
} }
+1 -1
View File
@@ -48,7 +48,7 @@ export function tryImageLine(alt: string | undefined, href: string, path: Conver
return attempt.ok ? attempt.value.line : undefined return attempt.ok ? attempt.value.line : undefined
} }
// A demand names a run no spelling holds, and a carried node joins no run, so every pass carries at least one more node. // A carried node joins no run, so every pass carries at least one more node.
function emitLine(nodes: readonly AdfNode[], container: LineContainer, path: ConvertErrorPath): Result<EmittedLine> { function emitLine(nodes: readonly AdfNode[], container: LineContainer, path: ConvertErrorPath): Result<EmittedLine> {
const carried = new Set<number>() const carried = new Set<number>()
for (;;) { for (;;) {
+20 -8
View File
@@ -110,14 +110,20 @@ detail is settled at its own milestone.
added to that list is the odd one out: `unspellableMark` finds it after assembly and added to that list is the odd one out: `unspellableMark` finds it after assembly and
names a mark type against the line's path, so the failing run needs identifying before names a mark type against the line's path, so the failing run needs identifying before
the carry can replace the refusal `mark-inside-word` pinned. the carry can replace the refusal `mark-inside-word` pinned.
- [ ] **2e5 — Combined documents and the collision property.** Documents combining nodes rather - [x] **2e5 — Combined documents and the collision property.** Documents combining nodes rather
than isolating one, and the gate's collision property: no two corpus documents may emit than isolating one, and the gate's collision property: no two corpus documents may emit
the same bytes — one spelling for two documents is a round-trip break no parser can undo, the same bytes — one spelling for two documents is a round-trip break no parser can undo,
and it is provable without one. It also settles the emitter's one known approximation: and it is provable without one.
delimiter flanking is exact, but CommonMark's *matching* — the multiple-of-3 rule and the **Settled** (the maintainer, 2026-08-27): the approximation this item inherited — flanking
way a run splits across several openers — is not modelled. No reachable violation has been exact, CommonMark's *matching* unmodelled — had two round-trip breaks reachable by hand,
found by hand; the property test is what decides it, and 2e1's `carve-out-strike` pins a so the emitter now models the matching. `process_emphasis` runs over the runs the emitter
second backslash only flanking-without-matching emits. wrote (`emphasis-matching.ts`) and a pair it hands to another delimiter rides the carry,
which is what the multiple-of-3 rule did to the em in `un*a**b*****c**istic`. A delimiter
run in text now escapes wherever CommonMark could open or close with it, not only open:
one that could only close stole the spelling around it (`un*a* b*istic`), and escaping
both ways keeps every delimiter the emitter did not write out of the matching. The
canonical form gained a backslash where a run only closes — `\*not emphasis\*`, and
2e1's `carve-out-strike` a third and fourth.
- [ ] **2f — The attributes CommonMark cannot hold.** 1d's settled answer: the block nodes - [ ] **2f — The attributes CommonMark cannot hold.** 1d's settled answer: the block nodes
CommonMark spells — `blockquote`, `bulletList`, `codeBlock`, `heading`, `listItem`, CommonMark spells — `blockquote`, `bulletList`, `codeBlock`, `heading`, `listItem`,
`orderedList`, `paragraph`, `rule` — get directive sections in `spec/flavour.md` carrying `orderedList`, `paragraph`, `rule` — get directive sections in `spec/flavour.md` carrying
@@ -156,7 +162,10 @@ detail is settled at its own milestone.
stay above all of it — the vocabulary a string-typed attribute grammar needs, which is why stay above all of it — the vocabulary a string-typed attribute grammar needs, which is why
HTML will want them too, not a markdown spelling. Both node tables are a second copy of HTML will want them too, not a markdown spelling. Both node tables are a second copy of
`spec/flavour.md`'s prose with no drift guard, and a mistyped attribute name degrades into a `spec/flavour.md`'s prose with no drift guard, and a mistyped attribute name degrades into a
false refusal no test catches. false refusal no test catches. The parser reuses `emphasis-matching.ts` whole and lands it
beside the grammar module: `delimiterFlags` and `matchEmphasis` take CommonMark's own run
vocabulary rather than the emitter's, so no second `process_emphasis` exists to drift from
the first.
- [ ] **4 — Round-trip property tests** over the corpus, both ways — the thing that proves 2 and - [ ] **4 — Round-trip property tests** over the corpus, both ways — the thing that proves 2 and
3. Editor-normal (§2) gets its implementation here — `toEditorNormal(doc)` and the equality 3. Editor-normal (§2) gets its implementation here — `toEditorNormal(doc)` and the equality
the round-trip asserts, which over normalized input is the canonical serializer's compact the round-trip asserts, which over normalized input is the canonical serializer's compact
@@ -165,7 +174,10 @@ detail is settled at its own milestone.
lifts the branch floor §10 keeps below 100 for exactly those halves. lifts the branch floor §10 keeps below 100 for exactly those halves.
Generators emit editor-normal ADF (§2). Real sanitized ADF from live Atlassian APIs lands Generators emit editor-normal ADF (§2). Real sanitized ADF from live Atlassian APIs lands
here too (§10), in `corpus/real-payloads/`: an ADF→markdown→ADF check with no expected here too (§10), in `corpus/real-payloads/`: an ADF→markdown→ADF check with no expected
markdown, the payloads supplied by the maintainer. markdown, the payloads supplied by the maintainer. This subsumes 2e5's collision property —
a document that round-trips proves no other document shares its spelling — so decide here
whether that gate stays as the parser-free, faster-failing signal or goes; the half holding
no fixture duplicates is hygiene rather than a round-trip claim, and stays either way.
- [ ] **5 — Release pipeline, ship `0.1.0`.** Publish-on-version-change (§9), `NPM_TOKEN` secret, - [ ] **5 — Release pipeline, ship `0.1.0`.** Publish-on-version-change (§9), `NPM_TOKEN` secret,
the repo made public first (§6). The `ConvertErrorCode` freeze (§8) is checkable here: every the repo made public first (§6). The `ConvertErrorCode` freeze (§8) is checkable here: every
`corpus/unspellable/` document is a decision or a deferred trigger this file names, so the `corpus/unspellable/` document is a decision or a deferred trigger this file names, so the