diff --git a/README.md b/README.md index 02c7f22..45a8b4b 100644 --- a/README.md +++ b/README.md @@ -82,8 +82,8 @@ For structured output a record writes the row for you. A record is a template seen as columns: its fields are the columns, its `format` the whole. `--format json|ndjson|csv|sql` writes the records; the library's -`FakeRecord` (below) hands back the columns. Every column is a string — typed scalars -are on the release checklist, see [`todo.md`](todo.md). Save +`FakeRecord` (below) hands back the columns. A column is a string unless it declares a +[datatype](#datatype), and a [`null`](#null) item draws it as null. Save `mydata/users.json`: ```json @@ -164,7 +164,8 @@ r, err = f.FakeRecordTemplate(`{"format":"{x}","x":["a","b"]}`) // compile + ren | `WithDataFS(fsys)` | layer an `fs.FS`, such as your own `embed.FS` | | `WithoutShippedData()` | load only what you give | -A `*Record` carries its columns via `Columns()`, and serializes them with `JSON()` +A `*Record` carries its columns via `Columns()` — each a `Column` of `Name`, +`DataType`, rendered `Value` and `Null` — and serializes them with `JSON()` (one object), `CSVHeader()`/`CSVLine()`, or `SQLInsert(table)` — the shapes the CLI's `--format` writes. `FakeRecord` and `FakeRecordTemplate` take a record; a path or template that is not one — a bare string, a choice, or a folder — errors. @@ -255,10 +256,53 @@ Renders e.g. `bar foo baz`. Rejected at load: a `separator` without a `repeat`, a `separator` of `""` (the default), and a `repeat` that multiplies to more than 1 048 576 renders along any path of nested repeats. +### Datatype + +A record column may declare `datatype` — `integer`, `number` or `boolean` — so `json` +writes `42` rather than `"42"` and `sql` a bare literal; a column without one is a +string: + +```json +{ "format": "", + "id": { "format": "{seq()}", "datatype": "integer" }, + "paid": { "format": "{p}", "p": ["true", "false"], "datatype": "boolean" }, + "total": { "format": "{calc(net * qty, 2)}", "net": ["19.99", "5.00"], "qty": ["3", "7"], "datatype": "number" } } +``` + +Writes e.g. `{"id":1,"paid":true,"total":59.97}`. A column is a field of the top-level +template, or an item of a choice standing in for one; `datatype` anywhere else is a +load error. So is a column that can render text its datatype rejects — `integer` takes +`-?(0|[1-9][0-9]*)`, `number` a JSON number, `boolean` `true` or `false` — and the +error shows such a render: + +```text +order.id: datatype integer, but it can render "000", which is not an integer +``` + +A `{calc()}` fills an `integer` or `number` column only where it provably prints no +`NaN` or `Inf`: each operand is a plain decimal — a sign, digits, one dot — of at most +300 bytes, or a field holding only such a calc, and no divisor can be zero. An +`integer` column also needs a decimals count of `0`, or integer operands and no `/`. + +### Null + +A `null` item draws a record column as null: `json` writes `null`, `sql` `NULL`, and +`csv` an empty field, with an empty string written `""` so PostgreSQL's `COPY … CSV` +reads both back. `Fake` renders a null as `""`. The other items' weights skew its +odds: + +```json +{ "format": "", "deleted_at": null, "middle": [null, { "format": "{n}", "n": ["Ann", "Eva"], "weight": 3 }] } +``` + +`deleted_at` is null every draw, `middle` a name three draws in four. Rejected at +load: `null` anywhere but a column, naming `""`, and a column whose items declare +different datatypes. + ### Options and fields -`format`, `weight`, `repeat` and `separator` are the only options; **any other -key is a field** (see [Decisions](#decisions)). An object that does nothing a +`format`, `weight`, `repeat`, `separator` and `datatype` are the only options; **any +other key is a field** (see [Decisions](#decisions)). An object that does nothing a string can't — only a `format` — is rejected naming the string, as is a one-item choice naming its item. @@ -437,8 +481,8 @@ tokens add cost in proportion to the output. ## Decisions -- **Options and fields share one namespace.** `format`, `weight`, `repeat` and - `separator` are reserved; every other key is a field. Nesting fields under a +- **Options and fields share one namespace.** `format`, `weight`, `repeat`, + `separator` and `datatype` are reserved; every other key is a field. Nesting fields under a key, or prefixing options, would tax every template to guard against a misspelt option. - **`{a|b}` stays beside nested choices.** `[[…], […]]` picks the same way, but @@ -518,7 +562,8 @@ tokens add cost in proportion to the output. so `a/(b*c)` with `b` fixed at `0` and `c` varying loads and prints `Inf` every draw — catching it needs zero-absorbing algebra for a shape nobody writes. - **In data, a default written out and a constant spelled as a sample are load - errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `int(5,5)`, `float(1,1,2)`, + errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `datatype: "string"`, + `int(5,5)`, `float(1,1,2)`, `+5` and `05` each spell what a shorter form already spells, so each is rejected naming that form. The CLI's numbers follow the shell instead: `--seed 007` and `--repeat +3` are 7 and 3, as every command line reads them. @@ -557,6 +602,9 @@ tokens add cost in proportion to the output. row — would vary per draw. A fixed column set is what the CSV and `INSERT` contracts rest on, so the restriction holds even where a particular choice would happen to agree. +- **Null is a `null` item, not a rate.** A null is one more outcome of a column's + draw, so a choice's weights skew it like any other; a null-rate option would be a + second way to state odds. - **The performance gate asserts allocations, not wall-clock time.** `AllocsPerRun` is deterministic across machines, so a ±10% ceiling does not flake under CI load, while time varies with the machine and its neighbours. A rendering slowdown @@ -612,7 +660,9 @@ hold.go the hold: one draw per expansion for paths and operands, and its reference.go reference sigils, and binding references across the tree graph.go the render graph: edges, cycles, the repeat bound, tree walks builtins.go the {name()} function registry and its implementations -calc.go the {calc()} arithmetic evaluator: parser, eval, validation +calc.go the {calc()} arithmetic evaluator: parser, eval, validation, and the proof a typed column's calc is finite +datatype.go column datatypes: DataType, where datatype and null may sit, and the load check every typed render passes +renderlang.go what text a node can render, as relations over a scalar's grammar data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge cmd/fejkdata/ the fejkdata CLI data/ shipped data (JSON), embedded at build: locale folders + a misc folder diff --git a/builtins.go b/builtins.go index 8b9fec6..c958a54 100644 --- a/builtins.go +++ b/builtins.go @@ -5,6 +5,7 @@ import ( "errors" "fmt" "math" + "slices" "strconv" "strings" "unicode" @@ -24,36 +25,36 @@ const ( // samples read only the rng. A time-based id (uuid v7, ulid) draws its timestamp // from the rng, not the wall clock, so seeded output stays reproducible. var builtins = map[string]builtin{ - "luhn": {arity: 0, prep: derive(func(e string) string { return string(rune('0' + luhnCheck(e))) })}, - "mod11": {arity: 0, prep: derive(mod11Check)}, - "ean": {arity: 0, prep: derive(eanCheck)}, - "uuid": {arity: 0, prep: sample(uuidV7)}, - "ulid": {arity: 0, prep: sample(ulid)}, - "nanoid": {arity: 1, check: posIntArg, prep: chars(nanoidAlphabet)}, - "hex": {arity: 1, check: posIntArg, prep: chars(hexDigits)}, - "digits": {arity: 1, check: posIntArg, prep: chars("0123456789")}, - "upper": {arity: 1, check: posIntArg, prep: chars("ABCDEFGHIJKLMNOPQRSTUVWXYZ")}, - "lower": {arity: 1, check: posIntArg, prep: chars("abcdefghijklmnopqrstuvwxyz")}, + "luhn": {arity: 0, prep: derive(func(e string) string { return string(rune('0' + luhnCheck(e))) }), emits: always(textShape{{{decimalDigits, 1, 1}}})}, + "mod11": {arity: 0, prep: derive(mod11Check), emits: always(textShape{{{decimalDigits + "X", 1, 1}}})}, + "ean": {arity: 0, prep: derive(eanCheck), emits: always(textShape{{{decimalDigits, 1, 1}}})}, + "uuid": {arity: 0, prep: sample(uuidV7), emits: always(uuidShape)}, + "ulid": {arity: 0, prep: sample(ulid), emits: always(textShape{{{crockford[:8], 1, 1}, {crockford, 25, 25}}})}, + "nanoid": sampleOf(nanoidAlphabet), + "hex": sampleOf(hexDigits), + "digits": sampleOf(decimalDigits), + "upper": sampleOf("ABCDEFGHIJKLMNOPQRSTUVWXYZ"), + "lower": sampleOf("abcdefghijklmnopqrstuvwxyz"), "base64": {arity: 1, check: posIntArg, prep: func(a []string) callFn { n := atoi(a[0]) return func(s *session, _ string, _ []string) string { return base64.StdEncoding.EncodeToString(randBytes(s, n)) } - }}, + }, emits: base64Shape}, "int": {arity: 2, check: intRangeArgs, prep: func(a []string) callFn { lo, span := atoi(a[0]), atoi(a[1])-atoi(a[0])+1 return func(s *session, _ string, _ []string) string { return strconv.Itoa(lo + s.IntN(span)) } - }}, + }, emits: intShape}, "float": {arity: 3, check: floatArgs, prep: func(a []string) callFn { lo, hi, dp := atof(a[0]), atof(a[1]), atoi(a[2]) return func(s *session, _ string, _ []string) string { return strconv.FormatFloat(lo+s.Float64()*(hi-lo), 'f', dp, 64) } - }}, + }, emits: func(a []string) textShape { return printedFloat(atof(a[0]), atof(a[1]), atoi(a[2]), false) }}, "iban": {arity: 1, check: ibanArg, prep: func(a []string) callFn { cc := a[0] return func(s *session, _ string, _ []string) string { return iban(s, cc) } - }}, + }, emits: ibanShape}, "calc": {arity: -1, check: checkCalc, prep: calcPrep, operands: calcOperands}, "lowercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToLower), operands: transformOperand}, "uppercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToUpper), operands: transformOperand}, @@ -69,7 +70,7 @@ var builtins = map[string]builtin{ return func(s *session, _ string, _ []string) string { return strconv.FormatUint(s.next(key), 10) } - }}, + }, emits: always(textShape{{{nonZeroDigits, 1, 1}, {decimalDigits, 0, 19}}})}, } // derive and sample are the two argument-free builtin shapes: a derivation reads @@ -94,6 +95,82 @@ func chars(alphabet string) func([]string) callFn { } } +// sampleOf is the builtin that draws n characters from an alphabet. +func sampleOf(alphabet string) builtin { + return builtin{arity: 1, check: posIntArg, prep: chars(alphabet), emits: func(a []string) textShape { + n := atoi(a[0]) + return textShape{{{alphabet, n, n}}} + }} +} + +// always is the emits of a builtin whose args do not change what it can print. +func always(s textShape) func([]string) textShape { + return func([]string) textShape { return s } +} + +var uuidShape = textShape{{{hexDigits, 8, 8}, {"-", 1, 1}, {hexDigits, 4, 4}, {"-", 1, 1}, {"7", 1, 1}, {hexDigits, 3, 3}, {"-", 1, 1}, {"89ab", 1, 1}, {hexDigits, 3, 3}, {"-", 1, 1}, {hexDigits, 12, 12}}} + +const base64Alphabet = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/" + +func base64Shape(a []string) textShape { + n := atoi(a[0]) + pad := (3 - n%3) % 3 + size := 4*((n+2)/3) - pad + return textShape{{{base64Alphabet, size, size}, {"=", pad, pad}}} +} + +// intShape is what int prints: a sign only below zero, and no leading zero. +func intShape(a []string) textShape { + lo, hi := atoi(a[0]), atoi(a[1]) + var s textShape + if lo <= 0 && hi >= 0 { + s = append(s, []charRun{{"0", 1, 1}}) + } + if hi > 0 { + s = append(s, []charRun{{nonZeroDigits, 1, 1}, {decimalDigits, 0, len(a[1]) - 1}}) + } + if lo < 0 { + s = append(s, []charRun{{"-", 1, 1}, {nonZeroDigits, 1, 1}, {decimalDigits, 0, len(a[0]) - 2}}) + } + return s +} + +func ibanShape(a []string) textShape { + cc, digits := a[0], ibanLen[a[0]]-2 + return textShape{{{cc[:1], 1, 1}, {cc[1:], 1, 1}, {decimalDigits, digits, digits}}} +} + +// shortestFraction bounds the fraction FormatFloat's shortest form prints: at most 17 +// significant digits after up to 323 zeros. +const shortestFraction = 340 + +// printedFloat is what strconv.FormatFloat(v, 'f', dp, 64) prints for a v in [lo, hi] +// that is whole when integral. +func printedFloat(lo, hi float64, dp int, integral bool) textShape { + digits := len(strconv.FormatFloat(math.Floor(math.Max(math.Abs(lo), math.Abs(hi))), 'f', 0, 64)) + 1 // one more for a rounding carry + wholes := [][]charRun{{{"0", 1, 1}}, {{nonZeroDigits, 1, 1}, {decimalDigits, 0, digits - 1}}} + fractions := [][]charRun{nil} + switch { + case dp > 0: + fractions = [][]charRun{{{".", 1, 1}, {decimalDigits, dp, dp}}} + case dp < 0 && !integral: + fractions = append(fractions, []charRun{{".", 1, 1}, {decimalDigits, 1, shortestFraction}}) + } + signs := [][]charRun{nil} + if lo < 0 || math.Signbit(lo) { + signs = append(signs, []charRun{{"-", 1, 1}}) + } + var s textShape + for _, sign := range signs { + for _, whole := range wholes { + for _, fraction := range fractions { + s = append(s, slices.Concat(sign, whole, fraction)) + } + } + } + return s +} + const hexDigits = "0123456789abcdef" // transforms are the builtins that rewrite one operand's value; they nest, so @@ -106,20 +183,19 @@ var transforms = map[string]func(string) string{ // unwrapTransform peels nested transform calls off an operand arg, returning the // field it finally names and the transforms to apply, innermost last. -func unwrapTransform(arg string) (leaf string, chain []func(string) string, err error) { +func unwrapTransform(arg string) (leaf string, chain []string, err error) { for { name, args, isCall := funcCall(arg) if !isCall { return arg, chain, nil } - fn, isTransform := transforms[name] - if !isTransform { + if _, isTransform := transforms[name]; !isTransform { return "", nil, fmt.Errorf("%s(%s) is not a transform, so it cannot be an operand", name, strings.Join(args, ",")) } if len(args) != 1 { return "", nil, fmt.Errorf("%s takes 1 arg, got %d", name, len(args)) } - chain = append(chain, fn) + chain = append(chain, name) arg = args[0] } } @@ -150,10 +226,14 @@ func transformPrep(outer func(string) string) func([]string) callFn { if err != nil { panic(fmt.Sprintf("fejkdata: transform arg %q reached prep unvalidated: %v", a[0], err)) } + fns := make([]func(string) string, len(chain)) + for i, name := range chain { + fns[i] = transforms[name] + } return func(_ *session, _ string, operands []string) string { v := operands[0] - for i := len(chain) - 1; i >= 0; i-- { - v = chain[i](v) + for i := len(fns) - 1; i >= 0; i-- { + v = fns[i](v) } return outer(v) } diff --git a/calc.go b/calc.go index cd776b7..7d7f2cd 100644 --- a/calc.go +++ b/calc.go @@ -6,6 +6,7 @@ import ( "strconv" "strings" "unicode" + "unicode/utf8" ) // calcNode is a parsed expression node. It evaluates over the operand values expand @@ -154,10 +155,12 @@ func calcText(n calcNode) string { return "?" } -// neverNumeric reports a node no render of which is a number: fixed text that does -// not parse, or a choice of only such items. text is one such render. +// neverNumeric reports a node no render of which is a number: a null, fixed text that +// does not parse, or a choice of only such items. text is one such render. func neverNumeric(n node) (text string, never bool) { switch n := n.(type) { + case *null: + return "", true case *template: if !n.fixed || n.repeat > 1 { return "", false @@ -191,10 +194,7 @@ func calcPrep(args []string) callFn { at[name] = i } placed := indexVars(expr, at) - dp := -1 - if len(args) == 2 { - dp = atoi(args[1]) - } + dp := calcDecimals(args) return func(_ *session, _ string, operands []string) string { return strconv.FormatFloat(placed.eval(operands), 'f', dp, 64) } @@ -384,3 +384,256 @@ func contains(bs []byte, b byte) bool { } return false } + +// calcDecimals is a calc's decimals count, or -1 for the shortest form. +func calcDecimals(args []string) int { + if len(args) == 2 { + return atoi(args[1]) + } + return -1 +} + +// calcLimit is the largest magnitude a proof accepts as finite, far enough below +// math.MaxFloat64 that rounding in the bounds cannot hide an overflow. +const calcLimit = 1e300 + +// maxOperandLen is the longest operand text a proof bounds by its length, so that +// bound, 10^maxOperandLen, stays within calcLimit. +const maxOperandLen = 300 + +// calcBound is what a proof knows of every value a calc can take: it lies in [lo, hi], +// is at least nonZero from zero unless nonZero is 0, and is whole when integral. +type calcBound struct { + lo, hi, nonZero float64 + integral bool +} + +func magnitude(b calcBound) float64 { return math.Max(math.Abs(b.lo), math.Abs(b.hi)) } + +// doubt is why a proof could not show a calc finite, and the render that shows it. +type doubt struct{ render, why string } + +type bounded struct { + b calcBound + d *doubt +} + +// calcProof bounds a typed column's calcs from their operands' renders, to show each +// prints a number rather than NaN or Inf. +type calcProof struct { + decimal *textLanguage + operands map[node]bounded + lengths map[node]int +} + +func newCalcProof() *calcProof { + p := &calcProof{operands: map[node]bounded{}, lengths: map[node]int{}} + p.decimal = newTextLanguage(decimalGrammar, p) + return p +} + +// call bounds one calc token of t. +func (p *calcProof) call(t *template, args []string) (calcBound, *doubt) { + expr, err := parseCalc(args[0]) + if err != nil { + panic(fmt.Sprintf("fejkdata: calc(%q) reached a proof unparsed: %v", args[0], err)) + } + b, d := p.expr(expr, t.fields) + if d != nil { + return b, &doubt{d.render, fmt.Sprintf("{calc(%s)}: %s", strings.Join(args, ", "), d.why)} + } + return b, nil +} + +func (p *calcProof) expr(n calcNode, fields map[string]node) (calcBound, *doubt) { + switch n := n.(type) { + case calcNum: + v := float64(n) + return calcBound{v, v, v, v == math.Trunc(v)}, nil + case calcVar: + return p.operand(string(n), fields[string(n)]) + case calcNeg: + b, d := p.expr(n.x, fields) + return calcBound{-b.hi, -b.lo, b.nonZero, b.integral}, d + case calcBin: + l, d := p.expr(n.l, fields) + if d != nil { + return l, d + } + r, d := p.expr(n.r, fields) + if d != nil { + return r, d + } + return combine(n, l, r) + } + panic(fmt.Sprintf("fejkdata: calc node %T has no bound", n)) +} + +// combine bounds one operation from the bounds of its sides. +func combine(n calcBin, l, r calcBound) (calcBound, *doubt) { + b := calcBound{integral: l.integral && r.integral} + switch n.op { + case '+': + b.lo, b.hi = l.lo+r.lo, l.hi+r.hi + case '-': + b.lo, b.hi = l.lo-r.hi, l.hi-r.lo + case '*': + b.lo = min(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi) + b.hi = max(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi) + b.nonZero = l.nonZero * r.nonZero + default: + if r.nonZero == 0 { + return b, &doubt{"+Inf", fmt.Sprintf("divides by %s, which can be zero", calcText(n.r))} + } + m := magnitude(l) / r.nonZero + b = calcBound{lo: -m, hi: m, nonZero: l.nonZero / magnitude(r)} + } + if b.lo > 0 || b.hi < 0 { + b.nonZero = math.Max(b.nonZero, math.Min(math.Abs(b.lo), math.Abs(b.hi))) + } + if !(magnitude(b) <= calcLimit) { + return b, &doubt{"+Inf", calcText(n) + " can overflow"} + } + return b, nil +} + +// operand bounds a calc operand, once per node. +func (p *calcProof) operand(name string, n node) (calcBound, *doubt) { + if seen, done := p.operands[n]; done { + return seen.b, seen.d + } + b, d := p.measure(name, n) + p.operands[n] = bounded{b, d} + return b, d +} + +// measure bounds an operand through the calc it renders when that is all it renders, +// and otherwise from its text: a plain decimal of at most maxOperandLen bytes. +func (p *calcProof) measure(name string, n node) (calcBound, *doubt) { + if t, ok := n.(*template); ok { + if args, isCalc := soleCalc(t); isCalc { + b, d := p.call(t, args) + return rounded(b, calcDecimals(args)), d + } + } + text := p.decimal.node(n, nil) + if w, escapes := text.escape(decimalAccept); escapes { + why := fmt.Sprintf("operand %q can render %s, which is not a plain decimal", name, w) + if w.why != "" { + why += ": " + w.why + } + return calcBound{}, &doubt{"NaN", why} + } + size := p.length(n) + if size > maxOperandLen { + return calcBound{}, &doubt{"NaN", fmt.Sprintf("operand %q can render more than %d bytes, too many to bound", name, maxOperandLen)} + } + ends, m := text.to[1], math.Pow(10, float64(size)) + b := calcBound{hi: m, nonZero: 1 / m, integral: ends&decimalFractional == 0} + if ends&decimalNegative != 0 { + b.lo = -m + } + if ends&decimalZero != 0 { + b.nonZero = 0 + } + return b, nil +} + +// soleCalc reports a template that renders one calc and nothing else, with its args. +func soleCalc(t *template) ([]string, bool) { + if t.repeat != 1 || len(t.ops) != 1 || t.ops[0].kind != 'b' { + return nil, false + } + name, args, _ := funcCall(t.format[1 : len(t.format)-1]) + return args, name == "calc" +} + +// rounded is b once printed to dp decimals, which moves a value by up to half a unit. +func rounded(b calcBound, dp int) calcBound { + if dp < 0 { + return b + } + half := math.Pow(10, -float64(dp)) / 2 + return calcBound{b.lo - half, b.hi + half, math.Max(0, b.nonZero-half), b.integral || dp == 0} +} + +// length is the most bytes a render of n can take, anything past maxOperandLen +// reported as maxOperandLen+1. +func (p *calcProof) length(n node) int { + if size, done := p.lengths[n]; done { + return size + } + size := 0 + switch n := n.(type) { + case *choice: + for _, it := range n.items { + size = max(size, p.length(it)) + } + case *template: + size = p.formatLength(n)*n.repeat + len(n.separator)*(n.repeat-1) + } + size = min(size, maxOperandLen+1) + p.lengths[n] = size + return size +} + +func (p *calcProof) formatLength(t *template) int { + size := 0 + _ = eachToken(t.format, func(tok ftoken) error { + if tok.kind == 'l' { + size += utf8.RuneLen(tok.r) + } else { + size += p.tokenLength(t, tok.body) + } + size = min(size, maxOperandLen+1) + return nil + }) + return size +} + +// tokenLength is the most bytes one token can print. A transform never lengthens a +// render that reads as a decimal: it maps each non-ASCII rune, two bytes or more, to at +// most two ASCII letters. +func (p *calcProof) tokenLength(t *template, body string) int { + name, args, isFunc := funcCall(body) + var arms []arm + switch _, isTransform := transforms[name]; { + case !isFunc: + arms = splitArms(body, t.refs) + case isTransform: + leaf, _, _ := unwrapTransform(args[0]) + arms = []arm{splitArm(leaf, t.refs)} + case name == "calc": + b, d := p.call(t, args) + if d != nil { + return len(d.render) + } + return shapeLength(printedFloat(b.lo, b.hi, calcDecimals(args), b.integral)) + default: + return shapeLength(builtins[name].emits(args)) + } + size := 0 + for _, a := range arms { + for _, leaf := range pathLeaves(t.fields[a.key], a.tail) { + size = max(size, p.length(leaf)) + } + } + return size +} + +// shapeLength is the most bytes a shape can emit, anything past maxOperandLen reported +// as maxOperandLen+1. +func shapeLength(s textShape) int { + longest := 0 + for _, alt := range s { + size := 0 + for _, run := range alt { + if run.max < 0 { + return maxOperandLen + 1 + } + size += run.max + } + longest = max(longest, size) + } + return min(longest, maxOperandLen+1) +} diff --git a/datatype.go b/datatype.go new file mode 100644 index 0000000..b92a398 --- /dev/null +++ b/datatype.go @@ -0,0 +1,140 @@ +package fejkdata + +import ( + "errors" + "fmt" +) + +// DataType is what a record column holds, which decides how a record writes its value. +type DataType int + +// The datatypes a column declares with "datatype"; a column without one is a string. +const ( + DataTypeString DataType = iota + DataTypeInteger + DataTypeNumber + DataTypeBoolean +) + +var dataTypeNames = [...]string{"string", "integer", "number", "boolean"} + +// String is the datatype as data spells it. +func (d DataType) String() string { + if d < 0 || int(d) >= len(dataTypeNames) { + return fmt.Sprintf("DataType(%d)", int(d)) + } + return dataTypeNames[d] +} + +// position is where a JSON value sits, which decides whether it may carry a datatype or +// be null. +type position int + +const ( + inFormat position = iota // rendered by a format, so neither + atTop // a category or an inline template, whose fields are the columns + inColumn // a column, or a choice item standing in for one +) + +// datatypeOf reads a template's "datatype" (default DataTypeString). +func datatypeOf(m map[string]any, pos position) (DataType, error) { + v, ok := m["datatype"] + if !ok { + return DataTypeString, nil + } + name, ok := v.(string) + if !ok { + return 0, fmt.Errorf("datatype must be a string, got %T", v) + } + if name == DataTypeString.String() { + return 0, fmt.Errorf("datatype %q is the default, so it has no effect; drop it", name) + } + for d := DataTypeInteger; d <= DataTypeBoolean; d++ { + if name != d.String() { + continue + } + if pos != inColumn { + return 0, errors.New("datatype only types a record column — a field of the top-level template — so it has no effect here") + } + return d, nil + } + return 0, fmt.Errorf(`datatype takes "integer", "number" or "boolean", got %q`, name) +} + +// columnDatatype is the datatype a column's items declare. They must agree, since a +// column holds one; a column only ever null is a string. +func columnDatatype(n node) (DataType, error) { + var declared []DataType + var collect func(node) + collect = func(n node) { + switch n := n.(type) { + case *choice: + for _, it := range n.items { + collect(it) + } + case *template: + declared = append(declared, n.datatype) + } + } + collect(n) + if len(declared) == 0 { + return DataTypeString, nil + } + for _, d := range declared { + if d != declared[0] { + return declared[0], fmt.Errorf("its items declare %s and %s; a column holds one datatype, so give every item the same", declared[0], d) + } + } + return declared[0], nil +} + +// datatypeSpec is what a datatype's text must satisfy: a grammar, the states a render +// may end in, and how an error names the datatype. +type datatypeSpec struct { + grammar *grammar + accept uint32 + noun string +} + +var datatypeSpecs = map[DataType]datatypeSpec{ + DataTypeInteger: {numberGrammar, integerAccept, "an integer"}, + DataTypeNumber: {numberGrammar, numberAccept, "a number"}, + DataTypeBoolean: {booleanGrammar, booleanAccept, "a boolean"}, +} + +// datatypeCheck proves every render of a typed column is text its datatype takes. One +// check covers a scope, so a node several columns reach is read once per grammar. +type datatypeCheck struct { + languages map[*grammar]*textLanguage + proof *calcProof +} + +func (c *datatypeCheck) check(path string, n node) error { + t, ok := n.(*template) + if !ok || t.datatype == DataTypeString { + return nil + } + spec := datatypeSpecs[t.datatype] + w, escapes := c.language(spec.grammar).node(t, nil).escape(spec.accept) + if !escapes { + return nil + } + msg := fmt.Sprintf("%s: datatype %s, but it can render %s, which is not %s", path, t.datatype, w, spec.noun) + if w.why != "" { + msg += ": " + w.why + } + return errors.New(msg) +} + +func (c *datatypeCheck) language(g *grammar) *textLanguage { + if c.proof == nil { + c.proof = newCalcProof() + c.languages = map[*grammar]*textLanguage{} + } + l, made := c.languages[g] + if !made { + l = newTextLanguage(g, c.proof) + c.languages[g] = l + } + return l +} diff --git a/fejkdata.go b/fejkdata.go index c558d9b..35cd73e 100644 --- a/fejkdata.go +++ b/fejkdata.go @@ -162,6 +162,8 @@ func paths(n node) []string { } } return out + case *null: + return []string{""} case *choice: out := []string{""} for p := range n.shared { diff --git a/graph.go b/graph.go index cf5cd5b..f65627e 100644 --- a/graph.go +++ b/graph.go @@ -186,7 +186,10 @@ func checkScope(s nodeScope) error { if err := s(func(path string, n node) error { return repeatCheck(path, n, mem) }); err != nil { return err } - return s(heldCheck) + if err := s(heldCheck); err != nil { + return err + } + return s((&datatypeCheck{}).check) } type reachMemo map[node]int diff --git a/node.go b/node.go index 384a2c6..d5af8e3 100644 --- a/node.go +++ b/node.go @@ -32,6 +32,12 @@ type choice struct { func (*choice) isNode() {} +// null is a record column's missing value, rendered as "". It is not zero-sized, so two +// nulls are two map keys. +type null struct{ _ byte } + +func (*null) isNode() {} + // template renders a format string, substituting {tokens} from fields. A bare // JSON string is a template with no fields. repeat (default 1) renders that format // that many times and joins the results with separator (default ""), each render @@ -41,6 +47,7 @@ type template struct { fields map[string]node repeat int separator string + datatype DataType ops []op // format compiled once (see compileOps); what expand walks grow int // minimum output size, to size the render buffer fixed bool // no op varies, so every render is lit @@ -66,26 +73,37 @@ func (t *template) field(seg string) (node, bool) { return n, ok } -// compile converts parsed JSON into a node tree, validating structure up front. -// Only a choice's items carry a weight, so one here would be inert whatever its type. +// compile converts parsed JSON — a category or an inline template — into a node tree, +// validating structure up front. func compile(v any) (node, error) { + return compileAt(v, atTop) +} + +// compileAt compiles a node that is no choice's item. Only a choice's items carry a +// weight, so one here would be inert whatever its type. +func compileAt(v any, pos position) (node, error) { if m, ok := v.(map[string]any); ok { if _, weighted := m["weight"]; weighted { return nil, fmt.Errorf("weight only skews a choice's items, so it has no effect here; it is an option and can never be a field") } } - return compileItem(v) + return compileItem(v, pos) } // compileItem compiles one node, allowing the weight a choice item may carry. -func compileItem(v any) (node, error) { +func compileItem(v any, pos position) (node, error) { switch v := v.(type) { case string: return compileString(v) case []any: - return compileChoice(v) + return compileChoice(v, pos) case map[string]any: - return compileTemplate(v) + return compileTemplate(v, pos) + case nil: + if pos != inColumn { + return nil, fmt.Errorf(`null is a record column's value; here it only renders "", so write ""`) + } + return &null{}, nil default: return nil, fmt.Errorf("a template value must be a string, a list or an object, not %s", jsonKind(v)) } @@ -99,8 +117,6 @@ func jsonKind(v any) string { return "a number" case bool: return "a boolean" - case nil: - return "null" } return fmt.Sprintf("%T", v) } @@ -136,7 +152,11 @@ func (t *template) compileFormat() error { return checkNoRepeatedRead(t.format, c, t.refs) } -func compileChoice(items []any) (node, error) { +func compileChoice(items []any, pos position) (node, error) { + itemPos := inFormat + if pos == inColumn { + itemPos = inColumn + } if len(items) == 0 { return nil, fmt.Errorf("empty choice") } @@ -160,7 +180,7 @@ func compileChoice(items []any) (node, error) { } total += w cum[i] = total - n, err := compileItem(raw) + n, err := compileItem(raw, itemPos) if err != nil { return nil, err } @@ -203,22 +223,26 @@ func checkNoRepeatedItem(items []any) error { return nil } -func compileTemplate(m map[string]any) (node, error) { - o, err := readOptions(m) +func compileTemplate(m map[string]any, pos position) (node, error) { + o, err := readOptions(m, pos) if err != nil { return nil, err } - fields, err := compileFields(m) + fieldPos := inFormat + if pos == atTop && o.repeat == 1 { + fieldPos = inColumn + } + fields, err := compileFields(m, fieldPos) if err != nil { return nil, err } - if len(fields) == 0 && o.repeat == 1 && !o.weighted { + if len(fields) == 0 && o.repeat == 1 && !o.weighted && o.datatype == DataTypeString { return nil, fmt.Errorf("an object holding only a format is a string; write %q", o.format) } if err := checkTokens(o.format, fields); err != nil { return nil, err } - t := &template{format: o.format, fields: fields, repeat: o.repeat, separator: o.separator} + t := &template{format: o.format, fields: fields, repeat: o.repeat, separator: o.separator, datatype: o.datatype} if err := t.compileFormat(); err != nil { return nil, err } @@ -227,13 +251,14 @@ func compileTemplate(m map[string]any) (node, error) { // templateOptions is what a template object's option keys say. type templateOptions struct { + datatype DataType format string repeat int separator string weighted bool } -func readOptions(m map[string]any) (templateOptions, error) { +func readOptions(m map[string]any, pos position) (templateOptions, error) { var o templateOptions format, ok := m["format"].(string) if !ok { @@ -245,6 +270,9 @@ func readOptions(m map[string]any) (templateOptions, error) { return o, err } o.repeat = repeat + if o.datatype, err = datatypeOf(m, pos); err != nil { + return o, err + } if sv, ok := m["separator"]; ok { if o.separator, ok = sv.(string); !ok { return o, fmt.Errorf("separator must be a string, got %T", sv) @@ -262,7 +290,7 @@ func readOptions(m map[string]any) (templateOptions, error) { // compileFields compiles every non-option key of a template object, in name order // so which of several bad fields is reported does not vary. -func compileFields(m map[string]any) (map[string]node, error) { +func compileFields(m map[string]any, pos position) (map[string]node, error) { fields := make(map[string]node, len(m)) keys := make([]string, 0, len(m)) for k := range m { @@ -276,7 +304,10 @@ func compileFields(m map[string]any) (map[string]node, error) { if err := checkName(k); err != nil { return nil, fmt.Errorf("field %w", err) } - n, err := compile(m[k]) + n, err := compileAt(m[k], pos) + if err == nil && pos == inColumn { + _, err = columnDatatype(n) + } if err != nil { return nil, fmt.Errorf("field %q: %w", k, err) } @@ -360,10 +391,10 @@ func checkName(name string) error { } // isOption reports whether a template key configures the node instead of naming a -// field. These four names can never be fields. +// field. These names can never be fields. func isOption(name string) bool { switch name { - case "format", "repeat", "separator", "weight": + case "datatype", "format", "repeat", "separator", "weight": return true } return false diff --git a/path.go b/path.go index ee9b1c5..687d9b9 100644 --- a/path.go +++ b/path.go @@ -60,7 +60,7 @@ func walkPath(n node, tail []string, w pathWalk) error { } return nil } - return fmt.Errorf("cannot descend into %T at %q", n, tail[0]) + return fmt.Errorf("no field %q", tail[0]) } // carriedByAll is the choice rule a path that must resolve on every call obeys: diff --git a/record.go b/record.go index 803e041..fc53ee0 100644 --- a/record.go +++ b/record.go @@ -9,10 +9,14 @@ import ( "strings" ) -// Column is one rendered column of a record. +// Column is one rendered column of a record. Value is the rendered text, which a +// serializer quotes for DataTypeString and writes bare for any other datatype; a Null +// column has no Value. type Column struct { - Name string - Value string + Name string + DataType DataType + Value string + Null bool } // Record is one record rendered from a template: every direct field is a column, @@ -28,70 +32,93 @@ func (r *Record) Columns() []Column { return append([]Column(nil), r.columns...) } -// JSON renders the record as one JSON object, every column a string. +// JSON renders the record as one JSON object. func (r *Record) JSON() string { - m := make(map[string]string, len(r.columns)) - for _, c := range r.columns { - m[c.Name] = c.Value + var b strings.Builder + b.WriteByte('{') + for i, c := range r.columns { + if i > 0 { + b.WriteByte(',') + } + b.WriteString(jsonString(c.Name)) + b.WriteByte(':') + b.WriteString(literal(c, jsonString, "null")) } - b, _ := json.Marshal(m) + b.WriteByte('}') + return b.String() +} + +func jsonString(s string) string { + b, _ := json.Marshal(s) return string(b) } // CSVHeader renders the column names as one CSV header line. func (r *Record) CSVHeader() string { - return csvLine(r.names()) + fields := make([]string, len(r.columns)) + for i, c := range r.columns { + fields[i] = csvField(c.Name) + } + return strings.Join(fields, ",") } -// CSVLine renders the column values as one CSV row. +// CSVLine renders the column values as one CSV row: a null column an empty field and an +// empty string "", the convention PostgreSQL's COPY reads a null by. func (r *Record) CSVLine() string { - return csvLine(r.values()) -} - -func (r *Record) names() []string { - out := make([]string, len(r.columns)) + fields := make([]string, len(r.columns)) for i, c := range r.columns { - out[i] = c.Name + fields[i] = literal(c, csvField, "") } - return out + if line := strings.Join(fields, ","); line != "" { + return line + } + return `""` // a blank line is a row every CSV reader drops } -func (r *Record) values() []string { - out := make([]string, len(r.columns)) - for i, c := range r.columns { - out[i] = c.Value +func csvField(s string) string { + if s == "" { + return `""` } - return out -} - -func csvLine(cols []string) string { var b strings.Builder w := csv.NewWriter(&b) - _ = w.Write(cols) + _ = w.Write([]string{s}) w.Flush() - line := strings.TrimSuffix(b.String(), "\n") - if line == "" { - return `""` // a blank line is a row every CSV reader drops - } - return line + return strings.TrimSuffix(b.String(), "\n") } -// SQLInsert renders the record as one INSERT statement into table: identifiers in -// ANSI double quotes, every value a single-quoted string literal. +// SQLInsert renders the record as one INSERT statement into table, identifiers in ANSI +// double quotes. func (r *Record) SQLInsert(table string) string { cols := make([]string, len(r.columns)) vals := make([]string, len(r.columns)) for i, c := range r.columns { cols[i] = quoteIdent(c.Name) - vals[i] = "'" + strings.ReplaceAll(c.Value, "'", "''") + "'" + vals[i] = literal(c, sqlString, "NULL") } return fmt.Sprintf("INSERT INTO %s (%s) VALUES (%s);", quoteIdent(table), strings.Join(cols, ", "), strings.Join(vals, ", ")) } +func sqlString(s string) string { + return "'" + strings.ReplaceAll(s, "'", "''") + "'" +} + func quoteIdent(s string) string { return `"` + strings.ReplaceAll(s, `"`, `""`) + `"` } +// literal spells a column the way a serializer writes it: quoted for a string, bare for +// any other datatype, whose every render the load check proved a literal, and nullText +// for a null. +func literal(c Column, quote func(string) string, nullText string) string { + switch { + case c.Null: + return nullText + case c.DataType == DataTypeString: + return quote(c.Value) + } + return c.Value +} + // FakeRecord renders a path as one record: the template it names, with each direct // field drawn as a column. Only a category-level template is a record — a path // that descends into a field, or that names a folder or a choice, is an error. @@ -116,7 +143,7 @@ func (f *Generator) FakeRecord(path string) (*Record, error) { // columns, or why it is not a record. type recordShape struct { t *template - columns []string + columns []Column err error } @@ -140,7 +167,7 @@ func (f *Generator) recordShapeOf(n node) recordShape { type RecordTemplate struct { g *Generator t *template - columns []string + columns []Column } // Fake renders the record with one draw. @@ -175,7 +202,7 @@ func (f *Generator) FakeRecordTemplate(input string) (*Record, error) { // recordOf is the fence both record entry points pass. The columns come back with // the template, fixed for every draw the caller goes on to make. -func recordOf(n node) (*template, []string, error) { +func recordOf(n node) (*template, []Column, error) { t, ok := n.(*template) if !ok { return nil, nil, errors.New("names a choice, not a template; a record is a template whose fields are its columns") @@ -183,13 +210,18 @@ func recordOf(n node) (*template, []string, error) { if t.repeat != 1 { return nil, nil, fmt.Errorf("carries repeat %d, which composes its format into one string; a record projects columns instead — drop the repeat and render the record again for more rows", t.repeat) } - columns := recordColumns(t) - if len(columns) == 0 { + names := recordColumns(t) + if len(names) == 0 { return nil, nil, errors.New("has no fields, so no columns") } - if err := checkColumnRefs(t, columns); err != nil { + if err := checkColumnRefs(t, names); err != nil { return nil, nil, err } + columns := make([]Column, len(names)) + for i, name := range names { + datatype, _ := columnDatatype(t.fields[name]) // compile refused a column whose items disagree + columns[i] = Column{Name: name, DataType: datatype} + } return t, columns, nil } @@ -262,11 +294,16 @@ func columnRefs(t *template, columns []string) ([]columnRef, error) { // renderRecord draws each column once, in the name order recordOf fixed, over one // reference scope shared across them. -func renderRecord(s *session, t *template, columns []string) *Record { +func renderRecord(s *session, t *template, columns []Column) *Record { scope := &draws{variant: map[string]node{}, value: map[string]string{}} - r := &Record{columns: make([]Column, len(columns))} - for i, name := range columns { - r.columns[i] = Column{Name: name, Value: render(s, t.fields[name], scope)} + r := &Record{columns: append([]Column(nil), columns...)} + for i := range r.columns { + n := drawn(s, t.fields[r.columns[i].Name]) + if _, isNull := n.(*null); isNull { + r.columns[i].Null = true + } else { + r.columns[i].Value = render(s, n, scope) + } } return r } diff --git a/render.go b/render.go index 7a7af3a..330ed0f 100644 --- a/render.go +++ b/render.go @@ -56,6 +56,8 @@ func render(s *session, n node, refScope *draws) string { switch n := n.(type) { case *choice: return render(s, pick(s, n), refScope) + case *null: + return "" case *template: if n.repeat == 1 { if n.fixed { diff --git a/renderlang.go b/renderlang.go new file mode 100644 index 0000000..4d56a2d --- /dev/null +++ b/renderlang.go @@ -0,0 +1,403 @@ +package fejkdata + +import ( + "slices" + "strconv" + "strings" + "unicode/utf8" +) + +// grammar is a deterministic automaton over a scalar's text: state 0 is dead, 1 the +// start, and each state lists the runes that leave it and where they lead. +type grammar [][]arc + +type arc struct { + on string + to int +} + +func (g *grammar) run(q int, s string) int { + for _, r := range s { + if q = g.step(q, r); q == 0 { + return 0 + } + } + return q +} + +func (g *grammar) step(q int, r rune) int { + for _, a := range (*g)[q] { + if strings.ContainsRune(a.on, r) { + return a.to + } + } + return 0 +} + +const ( + decimalDigits = "0123456789" + nonZeroDigits = "123456789" +) + +// numberGrammar reads a JSON number. States: 2 "-", 3 "0", 4 more integer digits, 5 ".", +// 6 fraction digits, 7 "e", 8 its sign, 9 exponent digits. +var numberGrammar = &grammar{ + nil, + {{"-", 2}, {"0", 3}, {nonZeroDigits, 4}}, + {{"0", 3}, {nonZeroDigits, 4}}, + {{".", 5}, {"eE", 7}}, + {{decimalDigits, 4}, {".", 5}, {"eE", 7}}, + {{decimalDigits, 6}}, + {{decimalDigits, 6}, {"eE", 7}}, + {{"+-", 8}, {decimalDigits, 9}}, + {{decimalDigits, 9}}, + {{decimalDigits, 9}}, +} + +const ( + integerAccept uint32 = 1<<3 | 1<<4 + numberAccept = integerAccept | 1<<6 | 1<<9 +) + +var booleanGrammar = &grammar{ + nil, + {{"t", 2}, {"f", 6}}, + {{"r", 3}}, {{"u", 4}}, {{"e", 5}}, nil, + {{"a", 7}}, {{"l", 8}}, {{"s", 9}}, {{"e", 10}}, nil, +} + +const booleanAccept uint32 = 1<<5 | 1<<10 + +// decimalGrammar reads what a calc operand must render to be proven finite: a sign, +// digits and at most one dot. Past the sign, states 4–9 are positive and 10–15 their +// negatives: 4 zero digits, 5 a nonzero integer, 6 a leading dot, 7 zero with a dot, +// 8 a nonzero integer with a zero fraction, 9 a nonzero fraction. +var decimalGrammar = &grammar{ + nil, + {{"+", 2}, {"-", 3}, {"0", 4}, {nonZeroDigits, 5}, {".", 6}}, + {{"0", 4}, {nonZeroDigits, 5}, {".", 6}}, + {{"0", 10}, {nonZeroDigits, 11}, {".", 12}}, + {{"0", 4}, {nonZeroDigits, 5}, {".", 7}}, + {{decimalDigits, 5}, {".", 8}}, + {{"0", 7}, {nonZeroDigits, 9}}, + {{"0", 7}, {nonZeroDigits, 9}}, + {{"0", 8}, {nonZeroDigits, 9}}, + {{decimalDigits, 9}}, + {{"0", 10}, {nonZeroDigits, 11}, {".", 13}}, + {{decimalDigits, 11}, {".", 14}}, + {{"0", 13}, {nonZeroDigits, 15}}, + {{"0", 13}, {nonZeroDigits, 15}}, + {{"0", 14}, {nonZeroDigits, 15}}, + {{decimalDigits, 15}}, +} + +const ( + decimalAccept uint32 = 1<<4 | 1<<5 | 1<<7 | 1<<8 | 1<<9 | 1<<10 | 1<<11 | 1<<13 | 1<<14 | 1<<15 + decimalNegative uint32 = 0xfc00 + decimalZero uint32 = 1<<4 | 1<<7 | 1<<10 | 1<<13 + decimalFractional uint32 = 1<<9 | 1<<15 +) + +// relation is what a node's renders do to a grammar: from each state, the states a +// render can end in, and one render reaching each. +type relation struct { + g *grammar + to []uint32 + w []witness // w[from*len(to)+to] +} + +// witness is one render, cut past witnessCap bytes, and why it can occur when the text +// alone does not say. +type witness struct { + text string + cut bool + why string +} + +const witnessCap = 60 + +func (w witness) then(next witness) witness { + if w.why == "" { + w.why = next.why + } + if w.cut { + return w + } + w.text += next.text + w.cut = next.cut + if len(w.text) > witnessCap { + end := witnessCap + for !utf8.RuneStart(w.text[end]) { + end-- + } + w.text, w.cut = w.text[:end], true + } + return w +} + +func (w witness) String() string { + if w.cut { + return strconv.Quote(w.text + "…") + } + return strconv.Quote(w.text) +} + +func newRelation(g *grammar) *relation { + n := len(*g) + return &relation{g: g, to: make([]uint32, n), w: make([]witness, n*n)} +} + +func (r *relation) add(from, to int, w witness) { + if r.to[from]&(1<>= 1; k == 0 { + return out + } + } +} + +// closure is any number of renders of r in a row, where r includes the empty render. +func (r *relation) closure() *relation { + for { + next := r.then(r) + if slices.Equal(next.to, r.to) { + return r + } + r = next + } +} + +// escape finds a render from the start that ends outside accept, preferring one that +// carries a reason. +func (r *relation) escape(accept uint32) (witness, bool) { + var found witness + escapes := false + for to := range r.to { + if (r.to[1]&^accept)&(1< 1 { + r = r.then(l.text(f.apply(n.separator)).then(r).power(n.repeat - 1)) + } + } + l.memo[key] = r + return r +} + +func (l *textLanguage) text(s string) *relation { return textRelation(l.g, s, "") } + +// format reads a template's format the way expand renders it: literal runs and tokens +// in turn. +func (l *textLanguage) format(t *template, f fold) *relation { + r := l.empty + var lit strings.Builder + _ = eachToken(t.format, func(tok ftoken) error { + if tok.kind == 'l' { + lit.WriteRune(tok.r) + return nil + } + r = r.then(l.text(f.apply(lit.String()))).then(l.token(t, tok.body, f)) + lit.Reset() + return nil + }) + return r.then(l.text(f.apply(lit.String()))) +} + +// token reads one {…} token: a field read, a transform over one, a calc, or what a +// builtin emits. +func (l *textLanguage) token(t *template, body string, f fold) *relation { + name, args, isFunc := funcCall(body) + if !isFunc { + var r *relation + for _, a := range splitArms(body, t.refs) { + r = union(r, l.read(t, a, f)) + } + return r + } + if _, isTransform := transforms[name]; isTransform { + leaf, chain, _ := unwrapTransform(args[0]) + inner := slices.Clone(chain) + slices.Reverse(inner) + return l.read(t, splitArm(leaf, t.refs), append(append(inner, name), f...)) + } + if name == "calc" { + return l.calc(t, args, f) + } + return l.shape(builtins[name].emits(args), f) +} + +// read is one arm of a token: every node its path can land on. +func (l *textLanguage) read(t *template, a arm, f fold) *relation { + var r *relation + for _, leaf := range pathLeaves(t.fields[a.key], a.tail) { + r = union(r, l.node(leaf, f)) + } + return r +} + +func (l *textLanguage) calc(t *template, args []string, f fold) *relation { + b, d := l.proof.call(t, args) + if d != nil { + return textRelation(l.g, f.apply(d.render), d.why) + } + return l.shape(printedFloat(b.lo, b.hi, calcDecimals(args), b.integral), f) +} + +func (l *textLanguage) shape(s textShape, f fold) *relation { + var r *relation + for _, alt := range s { + seq := l.empty + for _, run := range alt { + seq = seq.then(l.run(run, f)) + } + r = union(r, seq) + } + return r +} + +// run reads a charRun: min characters, then up to max-min more. +func (l *textLanguage) run(c charRun, f fold) *relation { + one := newRelation(l.g) + for from := range one.to { + for _, ch := range c.chars { + s := f.apply(string(ch)) + one.add(from, l.g.run(from, s), witness{text: s}) + } + } + more := l.empty + switch optional := union(one, l.empty); { + case c.max < 0: + more = optional.closure() + case c.max > c.min: + more = optional.power(c.max - c.min) + } + if c.min == 0 { + return more + } + return one.power(c.min).then(more) +} diff --git a/template.go b/template.go index 7ef9336..a673002 100644 --- a/template.go +++ b/template.go @@ -72,6 +72,9 @@ type builtin struct { // operands names the fields the call reads, which expand renders for it; nil // for a builtin that reads none. operands func(args []string) []string + // emits is the text a call can print, for the datatype check; nil for calc and the + // transforms, whose text the check derives from what they read. + emits func(args []string) textShape } // funcCall splits a "{token}" body shaped name(args) into its parts; ok is false diff --git a/todo.md b/todo.md index 0c2b28f..3e670b2 100644 --- a/todo.md +++ b/todo.md @@ -6,11 +6,6 @@ The record API lands first, so the data update can use it. ### Record API -- Typed columns — a column declares its type, so `json` writes `42` rather than - `"42"` and `sql` an unquoted literal: string, integer, number, boolean, and a - way to write null. A template that can render a value its type rejects is a - load error. The option key is reserved from then on, so a common column name - like `type` is a poor pick. - Struct-filling — fill a Go struct from `fake:"…"` tags holding a path or an inline template, for parity with gofakeit and go-faker. The field's Go type is the column type, through the same conversion and load checks as typed columns,