diff --git a/AGENTS.md b/AGENTS.md index aa7dde3..bff6957 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,69 +1,125 @@ # Working in this repo -Decisions that a reader would otherwise relitigate. Everything about *using* the library is in -`README.md`; what is still to build, and what is still undecided, is in `todo.md`. +Decisions a reader would otherwise relitigate, and the rules for every collaborator, human or +agent. Using the library: `README.md`. What is still to build: `todo.md`. -## 1. Two formats, never three +## 1. Three formats, ADF is the hub -ADF and one markdown flavour. **No HTML** — not as an output, not as an intermediate, not as a -convenience export. A consumer that wants HTML renders the markdown itself, with its own escaping -and its own stylesheet; a consumer that wants neither shows the markdown verbatim, which is what the -first one does. - -Three formats would mean six directions to keep lossless instead of two. +ADF, one markdown flavour, one HTML dialect. Six directions exposed, but markdown↔HTML compose +through ADF: four conversions exist to keep correct — never write a fifth. No fourth format, ever; +each one doubles the directions. ## 2. The round-trip is the product -`markdownToAdf(adfToMarkdown(doc))` must equal `doc`. Anything less and a consumer that lets someone -edit a ticket destroys what it could not represent — a panel, a mention, an attachment — in a -document it did not author. +`markdownToAdf(adfToMarkdown(doc))` and `htmlToAdf(adfToHtml(doc))` must equal `doc` — anything +less silently destroys content an editor could not represent, in a document it did not author. +When losslessness and readability conflict, losslessness wins. -That is why the flavour is *extended*: markdown has no syntax for most of what ADF holds, so the -flavour invents it. Designing that syntax is the first real task, and it is open (`todo.md`). +The other direction is a canonical fixpoint, not byte-identity: human markdown normalizes, the way +back yields the library's canonical spelling, and that spelling round-trips byte-identically. -Two consequences to settle before any node is implemented, not after: +Round-trip equality is a property tested over a corpus, not a claim made in prose. -- **What happens to a node the library does not know.** The documented ADF node set is not the whole - schema, and Atlassian adds to it. Whether an unknown node is carried opaquely, refused, or dropped - is a correctness decision for the whole library, and it decides the return shape of both functions. -- **Whether a lossless document must stay readable to a plain markdown reader.** Anything the flavour - invents is noise to a reader that does not know it. How much noise is acceptable bounds the syntax. +## 3. Unknown input policy -Round-trip equality is a property to test over a corpus, not a claim to make in prose. +- Unknown ADF node: carried opaquely — raw JSON rides a dedicated syntax in both formats, restored + byte-for-byte. The round-trip holds for documents newer than the library. +- Unmappable foreign HTML element: error result naming the element — never a silent drop. +- Bare `@name` / `:smile:` in typed text: stays a text node. Only directives produce + mention/emoji/media nodes; resolving names to ids needs I/O, which is the consumer's job. -## 3. Zero runtime dependencies +## 4. The flavour -Nothing in `dependencies`, ever. TypeScript and whatever the tests need are `devDependencies`, and -they never reach a consumer. A markdown parser is exactly the dependency this rule exists to refuse: -the flavour is not CommonMark, so a general parser would have to be extended into one anyway. +- Directives, one grammar for everything markdown lacks: `:::panel info` … `:::` blocks, + `:mention[@Mikael]{id=5b10a2}` inline. Prior art: CommonMark's generic-directives proposal. +- Plain CommonMark is a subset: the flavour adds syntax, never changes CommonMark meaning. +- Tables: one header row plus plain inline cells → pipe table; anything richer → directive form. +- Identity-bearing nodes carry their ids in attributes; a document is only portable within its + site — accepted. +- The HTML dialect mirrors this: semantic elements, stable `adf-*` classes, `data-*` for what HTML + cannot express, text always escaped. No stylesheet ships. -## 4. The package contract +## 5. Dependencies -- **ESM only.** No CommonJS build, no dual-package hazard. -- **Two entrypoints.** The built JavaScript for ordinary consumers, and the TypeScript source for - consumers that run TypeScript directly through Node's type stripping — the first consumer is one, - which is why this exists. -- **Types for both.** The JavaScript entrypoint ships `.d.ts` beside it; the TypeScript entrypoint is - its own types. -- **Published to public npmjs as `@larvit/atlassian-adf-converter`**, matching `@larvit/log`. Public - means the source is public: the Gitea repo starts private, and going public — with the LICENSE in - place — is a step before the first publish, not after it. -- **Exact versions.** `save-exact=true` in `.npmrc`, as in every other repo here. +`dependencies` is empty. A runtime dependency enters only through a decision entry here stating +why ~20 lines of own code cannot do the job, who maintains it, and what auditing it costs. So the +CommonMark and HTML parsers are written in this repo. `devDependencies`: few, each earning its +keep; they never reach a consumer. -## 5. Nothing about any consumer +## 6. The package contract -No Jira, no HTTP, no REST response shapes, no plainpages, no issue keys. The library takes a document -tree and returns a string, or the reverse. A consumer's concern that leaks in here is a seam nobody -declared — and the reason this is a library at all rather than a file in the client that needed it. +- ESM only — no CommonJS build, no dual-package hazard. +- Two entrypoints: built JavaScript, and TypeScript source for Node's type stripping. Types for + both (`.d.ts` beside the JavaScript). +- Published to public npmjs as `@larvit/atlassian-adf-converter`. Public source: the Gitea repo + goes public, LICENSE in place, before the first publish. +- Exact versions: `save-exact=true` in `.npmrc`. -## 6. Tests first, in Docker +## 7. Nothing about any consumer -Write the test for the behaviour wanted, then implement until it passes. `node --test`, beside the -code. Node, tsc and npm never run on the host — a compose service or a `docker run` against a -**full patch version** image tag (`node:24.19.0-alpine3.24`, never `node:24`), so the same commit -builds the same thing on a different day. +No Jira client, no HTTP, no REST shapes, no issue keys, no actual consumer named anywhere. Design +against the README's personas. -## 7. Style +## 8. Semver: the formats are API -Two-space indent, alphabetically sorted object keys, strict TypeScript. Failures are values, not -exceptions: a function that both returns a result and throws for some inputs has two error channels. +The emitted markdown and HTML are contracts. After 1.0: previously-emitted output parsing +differently, or not at all, is MAJOR; new syntax while old output still round-trips is MINOR. +Pre-1.0, normal 0.x rules. + +## 9. Release automation + +- `package.json` version on `main` is the source of truth. CI on `main`: tests green and version + differs from npm → publish and tag `vX.Y.Z`. No bump, no deploy; the bump is each shipping PR's + deliberate semver judgment. +- Renovate watches devDependencies, Docker pins and action tags; automerges everything on green CI. +- Docker images pin the full patch version (`node:24.19.0-alpine3.24`, never `node:24`); actions + pin semver tags. + +## 10. Tests first, in Docker + +Test for the behaviour wanted first, then implement until green. `node --test`, beside the code. +Node, tsc and npm never run on the host — only via the pinned images (§9). Tests are independent, +coverage does not decline, containers are torn down after a run. + +The corpus, all checked in: hand-built fixtures per node and combination; real sanitized ADF from +live Atlassian APIs; property-generated ADF trees; the CommonMark spec suite against +`markdownToAdf` and `markdownToHtml`. + +## 11. Code rules + +- Two-space indent, strict TypeScript, English everywhere. Alphabetical order wherever order + carries no meaning. +- Failures are values: everything returns + `Result` — `{ ok: true; value } | { ok: false; error: ConvertError }` — nothing throws. + `try/catch` only wrapped tightly around a call that genuinely throws, converted to a result on + the spot. +- No casts: `as`, `as unknown as`, non-null `!`. A boundary owes a type guard validating the + fields it claims (`isAdfDocument`); past it everything is typed. Make invalid states + unrepresentable. +- Explicit over implicit; descriptive names; no catch-all files (`utils`, `helpers`, `misc`). +- Reuse before adding; the smallest sufficient diff is the benchmark; no speculative generality — + a second consumer, or it goes. + +## 12. Prose to a minimum + +Applies everywhere: comments, every markdown file in this repo (this one included), PR text. + +- Default is no comment. One earns its single line only by naming an invariant, footgun or + external constraint the code cannot show — never restatement, history, absence or arrangement. + A second line belongs in the commit message or a decision entry here. +- Every prose comment in a diff is a review question; the default answer is delete. +- A doc paragraph says what the repo cannot say for itself, or it goes. The fix for a redundant + one is deletion, not trimming. A false claim in any doc is a bug, fixed where found. +- Published text — npm README, error messages, API docs — never references internal systems, + tickets or repos. + +## 13. Commits and PRs + +One-line commit messages and PR titles; short PR summaries. No AI-attribution markers, ever. + +## 14. Non-goals + +No wiki markup (§1), no network or filesystem I/O, no name→id resolution (§3), no ADF schema +validation or exported validator, no shipped CSS (§4), no streaming APIs, no performance budget — +conversions are O(n), real documents are kilobytes. A CLI is a later goal (`todo.md`), not a +non-goal. diff --git a/README.md b/README.md index 11071ae..2e50f6f 100644 --- a/README.md +++ b/README.md @@ -1,42 +1,59 @@ # @larvit/atlassian-adf-converter -Lossless conversion between **Atlassian Document Format** (ADF) and an extended markdown flavour -that can carry the nodes plain markdown has no syntax for. +Lossless conversion between **Atlassian Document Format** (ADF), an extended markdown flavour, and +an HTML dialect. -**Status: specification only. No code is implemented yet.** `todo.md` holds the plan and the design -questions still open; `AGENTS.md` holds the decisions already made. +**Status: specification only, no code yet.** Plan: `todo.md`. Decisions: `AGENTS.md`. ## What it is for -Jira Cloud's REST v3 API hands out issue descriptions and comment bodies as ADF — a JSON node tree, -ProseMirror-shaped — and takes them back the same way. There is no Atlassian endpoint that converts -it: `pf-editor-service/convert` was decommissioned and -[JRACLOUD-77436](https://jira.atlassian.com/browse/JRACLOUD-77436) is still an open request. The npm -ecosystem covers one direction each, drops what markdown cannot express, and none of it round-trips. +Atlassian Cloud REST APIs hand out rich text — issue descriptions, comments, pages — as ADF, a +ProseMirror-shaped JSON tree, and take it back the same way. No Atlassian endpoint converts it +(`pf-editor-service/convert` is decommissioned, +[JRACLOUD-77436](https://jira.atlassian.com/browse/JRACLOUD-77436) open), and the npm ecosystem is +one-directional and lossy. A consumer that shows a document and lets someone edit it needs both +directions lossless — otherwise saving destroys the panels, mentions and attachments it could not +represent. -A client that shows a ticket and lets someone edit it needs both directions, and needs them -lossless — otherwise saving an edit silently destroys the panels, mentions and attachments that were -in someone else's ticket. That is what this library is. +## The shape -## The intended shape - -Two pure functions and their types. No I/O, no network, no configuration: +Pure functions, no I/O, no configuration. ADF is the hub: markdown↔HTML compose through it. ```ts -adfToMarkdown(document: AdfDocument): string -markdownToAdf(markdown: string): AdfDocument +adfToMarkdown(doc: AdfDocument): Result +markdownToAdf(markdown: string): Result +adfToHtml(doc: AdfDocument): Result +htmlToAdf(html: string): Result +markdownToHtml(markdown: string): Result // via ADF +htmlToMarkdown(html: string): Result // via ADF +isAdfDocument(v: unknown): v is AdfDocument ``` -The published package is ESM only, has **no runtime dependencies**, and offers two entrypoints — the -built JavaScript for ordinary consumers, and the TypeScript source for consumers that run TypeScript -directly (Node's type stripping), with exported types either way. `AGENTS.md` §4 has the contract. +`Result` is `{ ok: true; value: T } | { ok: false; error: ConvertError }` — nothing throws. -## The first consumer +## The guarantees -[`plainpages-plugin-fastjira`](https://gitea.larvit.se/larvit/plainpages-plugin-fastjira) — a -server-rendered Jira client. Its read-only ticket view shows the markdown this library produces -verbatim, with no HTML rendering anywhere; its later write paths post back what this library -converts the other way. **That view is blocked on `0.1.0`,** and it needs `adfToMarkdown` first. +- `markdownToAdf(adfToMarkdown(doc))` equals `doc` — unknown node types included, carried opaquely + (AGENTS.md §3). +- `htmlToAdf(adfToHtml(doc))` equals `doc` — fidelity HTML cannot express rides `data-*` + attributes. +- Plain CommonMark is valid input to `markdownToAdf`; converting back yields the library's + canonical spelling, which round-trips byte-identically. +- Foreign HTML maps a documented element set; an unmappable element is an error, never a silent + drop. Well-formed HTML only — no tag-soup recovery. +- The emitted formats are semver surface (AGENTS.md §8). -The library knows nothing about that consumer. No Jira, no HTTP, no REST shapes, no plainpages — a -document tree in, a string out, and the reverse. +## Who it is for + +Personas, never named consumers (AGENTS.md §7): + +- **Viewer/editor app** — shows a document, lets a human edit, posts back. Losslessness above all. +- **Bot posting content** — converts generated markdown to ADF; needs the CommonMark promise. +- **Export/indexing tool** — bulk ADF→markdown/HTML; needs readable output. +- **LLM/agent pipeline** — documents to a model as markdown, edits back; needs the round-trip and + markdown legible to a reader that half-knows the flavour. + +## The package + +ESM only, no runtime dependencies, public npmjs. Two entrypoints — built JavaScript, and +TypeScript source for Node's type stripping — with types either way. Contract: `AGENTS.md` §5–6. diff --git a/todo.md b/todo.md index 1792a11..c28256a 100644 --- a/todo.md +++ b/todo.md @@ -1,49 +1,38 @@ # Todo -The plan, in order. Nothing here is built yet. - -## Open design questions — settle these first - -None of them have an answer yet, and each one changes what every later milestone implements. - -- [ ] **The flavour's syntax.** Markdown has no syntax for most of ADF. Every node in the inventory - below that is not plain markdown needs one, and the set has to be internally consistent rather - than invented node by node. Prior art worth reading before choosing: CommonMark's generic - directives proposal, MDX, Obsidian's and Pandoc's extensions, and what Atlassian's own - `editor-markdown-transformer` does (it is lossy — read it for the failure modes, not the design). -- [ ] **The unknown-node policy** (AGENTS.md §2). Carried opaquely, refused, or dropped — it decides - the return shape of both functions, so it cannot be retrofitted. -- [ ] **How readable a converted document must stay** to a reader that does not know the flavour. -- [ ] **Whether plain CommonMark is valid input** to `markdownToAdf`. A human typing ordinary - markdown into a comment box is the second consumer's whole write path. -- [ ] **Table fidelity.** ADF tables carry column widths, colspan, rowspan, header rows and cell - background colours; markdown tables carry none of it. -- [ ] **Identity-bearing nodes.** `mention` holds an account id, `media` an attachment id, `emoji` a - shortcode plus an id. The rendered text is not enough to reconstruct them, so the syntax has to - carry the id — and then a document is only portable within the site it came from. +The plan, in order. Nothing is built. Design questions are settled in `AGENTS.md`; remaining spec +detail is settled at its own milestone. ## Milestones -- [ ] **0 — Scaffold.** `package.json` with the §4 contract, `tsconfig.json`, `.npmrc`, LICENSE (MIT, - Larv IT AB), the Docker tooling setup, and `.gitea/workflows/ci.yml` gating branches. Mirror - `plainpages` for the workflow shape: `runs-on: docker-host`, actions pinned to semver tags. -- [ ] **1 — Settle the flavour.** Write the syntax down as this repo's specification before - implementing it, and make the round-trip corpus from it. -- [ ] **2 — `adfToMarkdown`.** The direction the first consumer needs. Ships `0.1.0`. -- [ ] **3 — `markdownToAdf`.** -- [ ] **4 — Round-trip property tests** over a corpus of real Jira documents, both ways. Not a - milestone that follows 2 and 3 so much as the thing that proves them. -- [ ] **5 — Release pipeline.** Tag-triggered publish to public npmjs, `NPM_TOKEN` secret, the repo - made public with the LICENSE in place first (AGENTS.md §4). +- [ ] **0 — Scaffold.** `package.json` per §6, `tsconfig.json`, `.npmrc` (`save-exact=true`), the + Docker tooling, `renovate.json` (§9), and `.gitea/workflows/ci.yml` gating branches: + `runs-on: docker-host`, actions pinned to semver tags. +- [ ] **1 — The flavour spec.** The markdown flavour written as this repo's specification before + any implementation: the directive grammar (attributes, escaping, nesting), each node's + syntax from the inventory below, the opaque-carry spelling, the pipe-vs-directive table + rule, and what CommonMark's raw-HTML constructs become in ADF, which has no raw-HTML node — + likely the §3 element mapping, error otherwise. Start the corpus (§10) from this spec. +- [ ] **2 — `adfToMarkdown`.** +- [ ] **3 — `markdownToAdf`.** The CommonMark parser is the largest single component. +- [ ] **4 — Round-trip property tests** over the corpus, both ways — the thing that proves 2 and 3. +- [ ] **5 — Release pipeline, ship `0.1.0`.** Publish-on-version-change (§9), `NPM_TOKEN` secret, + the repo made public first (§6). `0.1.0` is the markdown round-trip: both markdown + directions, the types, `isAdfDocument`. +- [ ] **6 — The HTML dialect spec.** Element-by-element mapping, the `data-*` fidelity scheme, the + opaque-carry form, and the documented foreign-element set `htmlToAdf` accepts. +- [ ] **7 — HTML, ship `0.2.0`.** `adfToHtml`, `htmlToAdf`, the composed `markdownToHtml` / + `htmlToMarkdown`. CommonMark spec suite runs against `markdownToHtml` from here (§10). +- [ ] **8 — CLI.** A later goal, shaped around the personas once the library exists. ## The ADF inventory to cover From Atlassian's [structure -reference](https://developer.atlassian.com/cloud/jira/platform/apis/document/structure/). **It is -not the whole schema** — real payloads also carry `taskList`/`taskItem`, `decisionList`/`decisionItem`, +reference](https://developer.atlassian.com/cloud/jira/platform/apis/document/structure/) — not the +whole schema: real payloads also carry `taskList`/`taskItem`, `decisionList`/`decisionItem`, `layoutSection`/`layoutColumn`, `blockCard`/`embedCard`, `extension`/`bodiedExtension`/`inlineExtension` -and `placeholder`, none of which are documented there. Treat the documented set as the floor, not the -ceiling, and see the unknown-node policy above. +and `placeholder`, none documented there. The documented set is the floor: the floor gets designed +syntax, the rest rides the opaque carry (§3) until it does too. | | | | --- | --- | @@ -52,6 +41,6 @@ ceiling, and see the unknown-node policy above. | Inline | `date` `emoji` `hardBreak` `inlineCard` `mediaInline` `mention` `status` `text` | | Marks | `border` `code` `em` `link` `strike` `strong` `subsup` `textColor` `underline` | -Plain markdown already covers `blockquote`, `bulletList`, `codeBlock`, `heading`, `orderedList`, +Plain markdown covers `blockquote`, `bulletList`, `codeBlock`, `heading`, `orderedList`, `paragraph`, `rule`, `listItem`, `hardBreak`, `text`, and the `code`, `em`, `link`, `strike` and `strong` marks. Everything else is what the flavour is for.