Typed columns and null #13

Merged
lilleman merged 10 commits from typed-columns into main 2026-09-15 15:19:39 +02:00
13 changed files with 1106 additions and 107 deletions
Showing only changes of commit 94dc562c6c - Show all commits
+59 -9
View File
@@ -82,8 +82,8 @@ For structured output a record writes the row for you.
A record is a template seen as columns: its fields are the columns, its `format`
the whole. `--format json|ndjson|csv|sql` writes the records; the library's
`FakeRecord` (below) hands back the columns. Every column is a string — typed scalars
are on the release checklist, see [`todo.md`](todo.md). Save
`FakeRecord` (below) hands back the columns. A column is a string unless it declares a
[datatype](#datatype), and a [`null`](#null) item draws it as null. Save
`mydata/users.json`:
```json
@@ -164,7 +164,8 @@ r, err = f.FakeRecordTemplate(`{"format":"{x}","x":["a","b"]}`) // compile + ren
| `WithDataFS(fsys)` | layer an `fs.FS`, such as your own `embed.FS` |
| `WithoutShippedData()` | load only what you give |
A `*Record` carries its columns via `Columns()`, and serializes them with `JSON()`
A `*Record` carries its columns via `Columns()` — each a `Column` of `Name`,
`DataType`, rendered `Value` and `Null` — and serializes them with `JSON()`
(one object), `CSVHeader()`/`CSVLine()`, or `SQLInsert(table)` — the shapes the
CLI's `--format` writes. `FakeRecord` and `FakeRecordTemplate` take a record; a
path or template that is not one — a bare string, a choice, or a folder — errors.
@@ -255,10 +256,53 @@ Renders e.g. `bar foo baz`. Rejected at load: a `separator` without a `repeat`,
a `separator` of `""` (the default), and a `repeat` that multiplies to more than
1 048 576 renders along any path of nested repeats.
### Datatype
A record column may declare `datatype` — `integer`, `number` or `boolean` — so `json`
writes `42` rather than `"42"` and `sql` a bare literal; a column without one is a
string:
```json
{ "format": "",
"id": { "format": "{seq()}", "datatype": "integer" },
"paid": { "format": "{p}", "p": ["true", "false"], "datatype": "boolean" },
"total": { "format": "{calc(net * qty, 2)}", "net": ["19.99", "5.00"], "qty": ["3", "7"], "datatype": "number" } }
```
Writes e.g. `{"id":1,"paid":true,"total":59.97}`. A column is a field of the top-level
template, or an item of a choice standing in for one; `datatype` anywhere else is a
load error. So is a column that can render text its datatype rejects — `integer` takes
`-?(0|[1-9][0-9]*)`, `number` a JSON number, `boolean` `true` or `false` — and the
error shows such a render:
```text
order.id: datatype integer, but it can render "000", which is not an integer
```
A `{calc()}` fills an `integer` or `number` column only where it provably prints no
`NaN` or `Inf`: each operand is a plain decimal — a sign, digits, one dot — of at most
300 bytes, or a field holding only such a calc, and no divisor can be zero. An
`integer` column also needs a decimals count of `0`, or integer operands and no `/`.
### Null
A `null` item draws a record column as null: `json` writes `null`, `sql` `NULL`, and
`csv` an empty field, with an empty string written `""` so PostgreSQL's `COPY … CSV`
reads both back. `Fake` renders a null as `""`. The other items' weights skew its
odds:
```json
{ "format": "", "deleted_at": null, "middle": [null, { "format": "{n}", "n": ["Ann", "Eva"], "weight": 3 }] }
```
`deleted_at` is null every draw, `middle` a name three draws in four. Rejected at
load: `null` anywhere but a column, naming `""`, and a column whose items declare
different datatypes.
### Options and fields
`format`, `weight`, `repeat` and `separator` are the only options; **any other
key is a field** (see [Decisions](#decisions)). An object that does nothing a
`format`, `weight`, `repeat`, `separator` and `datatype` are the only options; **any
other key is a field** (see [Decisions](#decisions)). An object that does nothing a
string can't — only a `format` — is rejected naming the string, as is a one-item
choice naming its item.
@@ -437,8 +481,8 @@ tokens add cost in proportion to the output.
## Decisions
- **Options and fields share one namespace.** `format`, `weight`, `repeat` and
`separator` are reserved; every other key is a field. Nesting fields under a
- **Options and fields share one namespace.** `format`, `weight`, `repeat`,
`separator` and `datatype` are reserved; every other key is a field. Nesting fields under a
key, or prefixing options, would tax every template to guard against a
misspelt option.
- **`{a|b}` stays beside nested choices.** `[[…], […]]` picks the same way, but
@@ -518,7 +562,8 @@ tokens add cost in proportion to the output.
so `a/(b*c)` with `b` fixed at `0` and `c` varying loads and prints `Inf` every
draw — catching it needs zero-absorbing algebra for a shape nobody writes.
- **In data, a default written out and a constant spelled as a sample are load
errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `int(5,5)`, `float(1,1,2)`,
errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `datatype: "string"`,
`int(5,5)`, `float(1,1,2)`,
`+5` and `05` each spell what a shorter form already spells, so each is rejected
naming that form. The CLI's numbers follow the shell instead: `--seed 007` and
`--repeat +3` are 7 and 3, as every command line reads them.
@@ -557,6 +602,9 @@ tokens add cost in proportion to the output.
row — would vary per draw. A fixed column set is what the CSV and `INSERT`
contracts rest on, so the restriction holds even where a particular choice would
happen to agree.
- **Null is a `null` item, not a rate.** A null is one more outcome of a column's
draw, so a choice's weights skew it like any other; a null-rate option would be a
second way to state odds.
- **The performance gate asserts allocations, not wall-clock time.** `AllocsPerRun`
is deterministic across machines, so a ±10% ceiling does not flake under CI load,
while time varies with the machine and its neighbours. A rendering slowdown
@@ -612,7 +660,9 @@ hold.go the hold: one draw per expansion for paths and operands, and its
reference.go reference sigils, and binding references across the tree
graph.go the render graph: edges, cycles, the repeat bound, tree walks
builtins.go the {name()} function registry and its implementations
calc.go the {calc()} arithmetic evaluator: parser, eval, validation
calc.go the {calc()} arithmetic evaluator: parser, eval, validation, and the proof a typed column's calc is finite
datatype.go column datatypes: DataType, where datatype and null may sit, and the load check every typed render passes
renderlang.go what text a node can render, as relations over a scalar's grammar
data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge
cmd/fejkdata/ the fejkdata CLI
data/ shipped data (JSON), embedded at build: locale folders + a misc folder
+101 -21
View File
@@ -5,6 +5,7 @@ import (
"errors"
"fmt"
"math"
"slices"
"strconv"
"strings"
"unicode"
@@ -24,36 +25,36 @@ const (
// samples read only the rng. A time-based id (uuid v7, ulid) draws its timestamp
// from the rng, not the wall clock, so seeded output stays reproducible.
var builtins = map[string]builtin{
"luhn": {arity: 0, prep: derive(func(e string) string { return string(rune('0' + luhnCheck(e))) })},
"mod11": {arity: 0, prep: derive(mod11Check)},
"ean": {arity: 0, prep: derive(eanCheck)},
"uuid": {arity: 0, prep: sample(uuidV7)},
"ulid": {arity: 0, prep: sample(ulid)},
"nanoid": {arity: 1, check: posIntArg, prep: chars(nanoidAlphabet)},
"hex": {arity: 1, check: posIntArg, prep: chars(hexDigits)},
"digits": {arity: 1, check: posIntArg, prep: chars("0123456789")},
"upper": {arity: 1, check: posIntArg, prep: chars("ABCDEFGHIJKLMNOPQRSTUVWXYZ")},
"lower": {arity: 1, check: posIntArg, prep: chars("abcdefghijklmnopqrstuvwxyz")},
"luhn": {arity: 0, prep: derive(func(e string) string { return string(rune('0' + luhnCheck(e))) }), emits: always(textShape{{{decimalDigits, 1, 1}}})},
"mod11": {arity: 0, prep: derive(mod11Check), emits: always(textShape{{{decimalDigits + "X", 1, 1}}})},
"ean": {arity: 0, prep: derive(eanCheck), emits: always(textShape{{{decimalDigits, 1, 1}}})},
"uuid": {arity: 0, prep: sample(uuidV7), emits: always(uuidShape)},
"ulid": {arity: 0, prep: sample(ulid), emits: always(textShape{{{crockford[:8], 1, 1}, {crockford, 25, 25}}})},
"nanoid": sampleOf(nanoidAlphabet),
"hex": sampleOf(hexDigits),
"digits": sampleOf(decimalDigits),
"upper": sampleOf("ABCDEFGHIJKLMNOPQRSTUVWXYZ"),
"lower": sampleOf("abcdefghijklmnopqrstuvwxyz"),
"base64": {arity: 1, check: posIntArg, prep: func(a []string) callFn {
n := atoi(a[0])
return func(s *session, _ string, _ []string) string {
return base64.StdEncoding.EncodeToString(randBytes(s, n))
}
}},
}, emits: base64Shape},
"int": {arity: 2, check: intRangeArgs, prep: func(a []string) callFn {
lo, span := atoi(a[0]), atoi(a[1])-atoi(a[0])+1
return func(s *session, _ string, _ []string) string { return strconv.Itoa(lo + s.IntN(span)) }
}},
}, emits: intShape},
"float": {arity: 3, check: floatArgs, prep: func(a []string) callFn {
lo, hi, dp := atof(a[0]), atof(a[1]), atoi(a[2])
return func(s *session, _ string, _ []string) string {
return strconv.FormatFloat(lo+s.Float64()*(hi-lo), 'f', dp, 64)
}
}},
}, emits: func(a []string) textShape { return printedFloat(atof(a[0]), atof(a[1]), atoi(a[2]), false) }},
"iban": {arity: 1, check: ibanArg, prep: func(a []string) callFn {
cc := a[0]
return func(s *session, _ string, _ []string) string { return iban(s, cc) }
}},
}, emits: ibanShape},
"calc": {arity: -1, check: checkCalc, prep: calcPrep, operands: calcOperands},
"lowercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToLower), operands: transformOperand},
"uppercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToUpper), operands: transformOperand},
@@ -69,7 +70,7 @@ var builtins = map[string]builtin{
return func(s *session, _ string, _ []string) string {
return strconv.FormatUint(s.next(key), 10)
}
}},
}, emits: always(textShape{{{nonZeroDigits, 1, 1}, {decimalDigits, 0, 19}}})},
}
// derive and sample are the two argument-free builtin shapes: a derivation reads
@@ -94,6 +95,82 @@ func chars(alphabet string) func([]string) callFn {
}
}
// sampleOf is the builtin that draws n characters from an alphabet.
func sampleOf(alphabet string) builtin {
return builtin{arity: 1, check: posIntArg, prep: chars(alphabet), emits: func(a []string) textShape {
n := atoi(a[0])
return textShape{{{alphabet, n, n}}}
}}
}
// always is the emits of a builtin whose args do not change what it can print.
func always(s textShape) func([]string) textShape {
return func([]string) textShape { return s }
}
var uuidShape = textShape{{{hexDigits, 8, 8}, {"-", 1, 1}, {hexDigits, 4, 4}, {"-", 1, 1}, {"7", 1, 1}, {hexDigits, 3, 3}, {"-", 1, 1}, {"89ab", 1, 1}, {hexDigits, 3, 3}, {"-", 1, 1}, {hexDigits, 12, 12}}}
const base64Alphabet = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/"
func base64Shape(a []string) textShape {
n := atoi(a[0])
pad := (3 - n%3) % 3
size := 4*((n+2)/3) - pad
return textShape{{{base64Alphabet, size, size}, {"=", pad, pad}}}
}
// intShape is what int prints: a sign only below zero, and no leading zero.
func intShape(a []string) textShape {
lo, hi := atoi(a[0]), atoi(a[1])
var s textShape
if lo <= 0 && hi >= 0 {
s = append(s, []charRun{{"0", 1, 1}})
}
if hi > 0 {
s = append(s, []charRun{{nonZeroDigits, 1, 1}, {decimalDigits, 0, len(a[1]) - 1}})
}
if lo < 0 {
s = append(s, []charRun{{"-", 1, 1}, {nonZeroDigits, 1, 1}, {decimalDigits, 0, len(a[0]) - 2}})
}
return s
}
func ibanShape(a []string) textShape {
cc, digits := a[0], ibanLen[a[0]]-2
return textShape{{{cc[:1], 1, 1}, {cc[1:], 1, 1}, {decimalDigits, digits, digits}}}
}
// shortestFraction bounds the fraction FormatFloat's shortest form prints: at most 17
// significant digits after up to 323 zeros.
const shortestFraction = 340
// printedFloat is what strconv.FormatFloat(v, 'f', dp, 64) prints for a v in [lo, hi]
// that is whole when integral.
func printedFloat(lo, hi float64, dp int, integral bool) textShape {
digits := len(strconv.FormatFloat(math.Floor(math.Max(math.Abs(lo), math.Abs(hi))), 'f', 0, 64)) + 1 // one more for a rounding carry
wholes := [][]charRun{{{"0", 1, 1}}, {{nonZeroDigits, 1, 1}, {decimalDigits, 0, digits - 1}}}
fractions := [][]charRun{nil}
switch {
case dp > 0:
fractions = [][]charRun{{{".", 1, 1}, {decimalDigits, dp, dp}}}
case dp < 0 && !integral:
fractions = append(fractions, []charRun{{".", 1, 1}, {decimalDigits, 1, shortestFraction}})
}
signs := [][]charRun{nil}
if lo < 0 || math.Signbit(lo) {
signs = append(signs, []charRun{{"-", 1, 1}})
}
var s textShape
for _, sign := range signs {
for _, whole := range wholes {
for _, fraction := range fractions {
s = append(s, slices.Concat(sign, whole, fraction))
}
}
}
return s
}
const hexDigits = "0123456789abcdef"
// transforms are the builtins that rewrite one operand's value; they nest, so
@@ -106,20 +183,19 @@ var transforms = map[string]func(string) string{
// unwrapTransform peels nested transform calls off an operand arg, returning the
// field it finally names and the transforms to apply, innermost last.
func unwrapTransform(arg string) (leaf string, chain []func(string) string, err error) {
func unwrapTransform(arg string) (leaf string, chain []string, err error) {
for {
name, args, isCall := funcCall(arg)
if !isCall {
return arg, chain, nil
}
fn, isTransform := transforms[name]
if !isTransform {
if _, isTransform := transforms[name]; !isTransform {
return "", nil, fmt.Errorf("%s(%s) is not a transform, so it cannot be an operand", name, strings.Join(args, ","))
}
if len(args) != 1 {
return "", nil, fmt.Errorf("%s takes 1 arg, got %d", name, len(args))
}
chain = append(chain, fn)
chain = append(chain, name)
arg = args[0]
}
}
@@ -150,10 +226,14 @@ func transformPrep(outer func(string) string) func([]string) callFn {
if err != nil {
panic(fmt.Sprintf("fejkdata: transform arg %q reached prep unvalidated: %v", a[0], err))
}
fns := make([]func(string) string, len(chain))
for i, name := range chain {
fns[i] = transforms[name]
}
return func(_ *session, _ string, operands []string) string {
v := operands[0]
for i := len(chain) - 1; i >= 0; i-- {
v = chain[i](v)
for i := len(fns) - 1; i >= 0; i-- {
v = fns[i](v)
}
return outer(v)
}
+259 -6
View File
@@ -6,6 +6,7 @@ import (
"strconv"
"strings"
"unicode"
"unicode/utf8"
)
// calcNode is a parsed expression node. It evaluates over the operand values expand
@@ -154,10 +155,12 @@ func calcText(n calcNode) string {
return "?"
}
// neverNumeric reports a node no render of which is a number: fixed text that does
// not parse, or a choice of only such items. text is one such render.
// neverNumeric reports a node no render of which is a number: a null, fixed text that
// does not parse, or a choice of only such items. text is one such render.
func neverNumeric(n node) (text string, never bool) {
switch n := n.(type) {
case *null:
return "", true
case *template:
if !n.fixed || n.repeat > 1 {
return "", false
@@ -191,10 +194,7 @@ func calcPrep(args []string) callFn {
at[name] = i
}
placed := indexVars(expr, at)
dp := -1
if len(args) == 2 {
dp = atoi(args[1])
}
dp := calcDecimals(args)
return func(_ *session, _ string, operands []string) string {
return strconv.FormatFloat(placed.eval(operands), 'f', dp, 64)
}
@@ -384,3 +384,256 @@ func contains(bs []byte, b byte) bool {
}
return false
}
// calcDecimals is a calc's decimals count, or -1 for the shortest form.
func calcDecimals(args []string) int {
if len(args) == 2 {
return atoi(args[1])
}
return -1
}
// calcLimit is the largest magnitude a proof accepts as finite, far enough below
// math.MaxFloat64 that rounding in the bounds cannot hide an overflow.
const calcLimit = 1e300
// maxOperandLen is the longest operand text a proof bounds by its length, so that
// bound, 10^maxOperandLen, stays within calcLimit.
const maxOperandLen = 300
// calcBound is what a proof knows of every value a calc can take: it lies in [lo, hi],
// is at least nonZero from zero unless nonZero is 0, and is whole when integral.
type calcBound struct {
lo, hi, nonZero float64
integral bool
}
func magnitude(b calcBound) float64 { return math.Max(math.Abs(b.lo), math.Abs(b.hi)) }
// doubt is why a proof could not show a calc finite, and the render that shows it.
type doubt struct{ render, why string }
type bounded struct {
b calcBound
d *doubt
}
// calcProof bounds a typed column's calcs from their operands' renders, to show each
// prints a number rather than NaN or Inf.
type calcProof struct {
decimal *textLanguage
operands map[node]bounded
lengths map[node]int
}
func newCalcProof() *calcProof {
p := &calcProof{operands: map[node]bounded{}, lengths: map[node]int{}}
p.decimal = newTextLanguage(decimalGrammar, p)
return p
}
// call bounds one calc token of t.
func (p *calcProof) call(t *template, args []string) (calcBound, *doubt) {
expr, err := parseCalc(args[0])
if err != nil {
panic(fmt.Sprintf("fejkdata: calc(%q) reached a proof unparsed: %v", args[0], err))
}
b, d := p.expr(expr, t.fields)
if d != nil {
return b, &doubt{d.render, fmt.Sprintf("{calc(%s)}: %s", strings.Join(args, ", "), d.why)}
}
return b, nil
}
func (p *calcProof) expr(n calcNode, fields map[string]node) (calcBound, *doubt) {
switch n := n.(type) {
case calcNum:
v := float64(n)
return calcBound{v, v, v, v == math.Trunc(v)}, nil
case calcVar:
return p.operand(string(n), fields[string(n)])
case calcNeg:
b, d := p.expr(n.x, fields)
return calcBound{-b.hi, -b.lo, b.nonZero, b.integral}, d
case calcBin:
l, d := p.expr(n.l, fields)
if d != nil {
return l, d
}
r, d := p.expr(n.r, fields)
if d != nil {
return r, d
}
return combine(n, l, r)
}
panic(fmt.Sprintf("fejkdata: calc node %T has no bound", n))
}
// combine bounds one operation from the bounds of its sides.
func combine(n calcBin, l, r calcBound) (calcBound, *doubt) {
b := calcBound{integral: l.integral && r.integral}
switch n.op {
case '+':
b.lo, b.hi = l.lo+r.lo, l.hi+r.hi
case '-':
b.lo, b.hi = l.lo-r.hi, l.hi-r.lo
case '*':
b.lo = min(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi)
b.hi = max(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi)
b.nonZero = l.nonZero * r.nonZero
default:
if r.nonZero == 0 {
return b, &doubt{"+Inf", fmt.Sprintf("divides by %s, which can be zero", calcText(n.r))}
}
m := magnitude(l) / r.nonZero
b = calcBound{lo: -m, hi: m, nonZero: l.nonZero / magnitude(r)}
}
if b.lo > 0 || b.hi < 0 {
b.nonZero = math.Max(b.nonZero, math.Min(math.Abs(b.lo), math.Abs(b.hi)))
}
if !(magnitude(b) <= calcLimit) {
return b, &doubt{"+Inf", calcText(n) + " can overflow"}
}
return b, nil
}
// operand bounds a calc operand, once per node.
func (p *calcProof) operand(name string, n node) (calcBound, *doubt) {
if seen, done := p.operands[n]; done {
return seen.b, seen.d
}
b, d := p.measure(name, n)
p.operands[n] = bounded{b, d}
return b, d
}
// measure bounds an operand through the calc it renders when that is all it renders,
// and otherwise from its text: a plain decimal of at most maxOperandLen bytes.
func (p *calcProof) measure(name string, n node) (calcBound, *doubt) {
if t, ok := n.(*template); ok {
if args, isCalc := soleCalc(t); isCalc {
b, d := p.call(t, args)
return rounded(b, calcDecimals(args)), d
}
}
text := p.decimal.node(n, nil)
if w, escapes := text.escape(decimalAccept); escapes {
why := fmt.Sprintf("operand %q can render %s, which is not a plain decimal", name, w)
if w.why != "" {
why += ": " + w.why
}
return calcBound{}, &doubt{"NaN", why}
}
size := p.length(n)
if size > maxOperandLen {
return calcBound{}, &doubt{"NaN", fmt.Sprintf("operand %q can render more than %d bytes, too many to bound", name, maxOperandLen)}
}
ends, m := text.to[1], math.Pow(10, float64(size))
b := calcBound{hi: m, nonZero: 1 / m, integral: ends&decimalFractional == 0}
if ends&decimalNegative != 0 {
b.lo = -m
}
if ends&decimalZero != 0 {
b.nonZero = 0
}
return b, nil
}
// soleCalc reports a template that renders one calc and nothing else, with its args.
func soleCalc(t *template) ([]string, bool) {
if t.repeat != 1 || len(t.ops) != 1 || t.ops[0].kind != 'b' {
return nil, false
}
name, args, _ := funcCall(t.format[1 : len(t.format)-1])
return args, name == "calc"
}
// rounded is b once printed to dp decimals, which moves a value by up to half a unit.
func rounded(b calcBound, dp int) calcBound {
if dp < 0 {
return b
}
half := math.Pow(10, -float64(dp)) / 2
return calcBound{b.lo - half, b.hi + half, math.Max(0, b.nonZero-half), b.integral || dp == 0}
}
// length is the most bytes a render of n can take, anything past maxOperandLen
// reported as maxOperandLen+1.
func (p *calcProof) length(n node) int {
if size, done := p.lengths[n]; done {
return size
}
size := 0
switch n := n.(type) {
case *choice:
for _, it := range n.items {
size = max(size, p.length(it))
}
case *template:
size = p.formatLength(n)*n.repeat + len(n.separator)*(n.repeat-1)
}
size = min(size, maxOperandLen+1)
p.lengths[n] = size
return size
}
func (p *calcProof) formatLength(t *template) int {
size := 0
_ = eachToken(t.format, func(tok ftoken) error {
if tok.kind == 'l' {
size += utf8.RuneLen(tok.r)
} else {
size += p.tokenLength(t, tok.body)
}
size = min(size, maxOperandLen+1)
return nil
})
return size
}
// tokenLength is the most bytes one token can print. A transform never lengthens a
// render that reads as a decimal: it maps each non-ASCII rune, two bytes or more, to at
// most two ASCII letters.
func (p *calcProof) tokenLength(t *template, body string) int {
name, args, isFunc := funcCall(body)
var arms []arm
switch _, isTransform := transforms[name]; {
case !isFunc:
arms = splitArms(body, t.refs)
case isTransform:
leaf, _, _ := unwrapTransform(args[0])
arms = []arm{splitArm(leaf, t.refs)}
case name == "calc":
b, d := p.call(t, args)
if d != nil {
return len(d.render)
}
return shapeLength(printedFloat(b.lo, b.hi, calcDecimals(args), b.integral))
default:
return shapeLength(builtins[name].emits(args))
}
size := 0
for _, a := range arms {
for _, leaf := range pathLeaves(t.fields[a.key], a.tail) {
size = max(size, p.length(leaf))
}
}
return size
}
// shapeLength is the most bytes a shape can emit, anything past maxOperandLen reported
// as maxOperandLen+1.
func shapeLength(s textShape) int {
longest := 0
for _, alt := range s {
size := 0
for _, run := range alt {
if run.max < 0 {
return maxOperandLen + 1
}
size += run.max
}
longest = max(longest, size)
}
return min(longest, maxOperandLen+1)
}
+140
View File
@@ -0,0 +1,140 @@
package fejkdata
import (
"errors"
"fmt"
)
// DataType is what a record column holds, which decides how a record writes its value.
type DataType int
// The datatypes a column declares with "datatype"; a column without one is a string.
const (
DataTypeString DataType = iota
DataTypeInteger
DataTypeNumber
DataTypeBoolean
)
var dataTypeNames = [...]string{"string", "integer", "number", "boolean"}
// String is the datatype as data spells it.
func (d DataType) String() string {
if d < 0 || int(d) >= len(dataTypeNames) {
return fmt.Sprintf("DataType(%d)", int(d))
}
return dataTypeNames[d]
}
// position is where a JSON value sits, which decides whether it may carry a datatype or
// be null.
type position int
const (
inFormat position = iota // rendered by a format, so neither
atTop // a category or an inline template, whose fields are the columns
inColumn // a column, or a choice item standing in for one
)
// datatypeOf reads a template's "datatype" (default DataTypeString).
func datatypeOf(m map[string]any, pos position) (DataType, error) {
v, ok := m["datatype"]
if !ok {
return DataTypeString, nil
}
name, ok := v.(string)
if !ok {
return 0, fmt.Errorf("datatype must be a string, got %T", v)
}
if name == DataTypeString.String() {
return 0, fmt.Errorf("datatype %q is the default, so it has no effect; drop it", name)
}
for d := DataTypeInteger; d <= DataTypeBoolean; d++ {
if name != d.String() {
continue
}
if pos != inColumn {
return 0, errors.New("datatype only types a record column — a field of the top-level template — so it has no effect here")
}
return d, nil
}
return 0, fmt.Errorf(`datatype takes "integer", "number" or "boolean", got %q`, name)
}
// columnDatatype is the datatype a column's items declare. They must agree, since a
// column holds one; a column only ever null is a string.
func columnDatatype(n node) (DataType, error) {
var declared []DataType
var collect func(node)
collect = func(n node) {
switch n := n.(type) {
case *choice:
for _, it := range n.items {
collect(it)
}
case *template:
declared = append(declared, n.datatype)
}
}
collect(n)
if len(declared) == 0 {
return DataTypeString, nil
}
for _, d := range declared {
if d != declared[0] {
return declared[0], fmt.Errorf("its items declare %s and %s; a column holds one datatype, so give every item the same", declared[0], d)
}
}
return declared[0], nil
}
// datatypeSpec is what a datatype's text must satisfy: a grammar, the states a render
// may end in, and how an error names the datatype.
type datatypeSpec struct {
grammar *grammar
accept uint32
noun string
}
var datatypeSpecs = map[DataType]datatypeSpec{
DataTypeInteger: {numberGrammar, integerAccept, "an integer"},
DataTypeNumber: {numberGrammar, numberAccept, "a number"},
DataTypeBoolean: {booleanGrammar, booleanAccept, "a boolean"},
}
// datatypeCheck proves every render of a typed column is text its datatype takes. One
// check covers a scope, so a node several columns reach is read once per grammar.
type datatypeCheck struct {
languages map[*grammar]*textLanguage
proof *calcProof
}
func (c *datatypeCheck) check(path string, n node) error {
t, ok := n.(*template)
if !ok || t.datatype == DataTypeString {
return nil
}
spec := datatypeSpecs[t.datatype]
w, escapes := c.language(spec.grammar).node(t, nil).escape(spec.accept)
if !escapes {
return nil
}
msg := fmt.Sprintf("%s: datatype %s, but it can render %s, which is not %s", path, t.datatype, w, spec.noun)
if w.why != "" {
msg += ": " + w.why
}
return errors.New(msg)
}
func (c *datatypeCheck) language(g *grammar) *textLanguage {
if c.proof == nil {
c.proof = newCalcProof()
c.languages = map[*grammar]*textLanguage{}
}
l, made := c.languages[g]
if !made {
l = newTextLanguage(g, c.proof)
c.languages[g] = l
}
return l
}
+2
View File
@@ -162,6 +162,8 @@ func paths(n node) []string {
}
}
return out
case *null:
return []string{""}
case *choice:
out := []string{""}
for p := range n.shared {
+4 -1
View File
@@ -186,7 +186,10 @@ func checkScope(s nodeScope) error {
if err := s(func(path string, n node) error { return repeatCheck(path, n, mem) }); err != nil {
return err
}
return s(heldCheck)
if err := s(heldCheck); err != nil {
return err
}
return s((&datatypeCheck{}).check)
}
type reachMemo map[node]int
+51 -20
View File
@@ -32,6 +32,12 @@ type choice struct {
func (*choice) isNode() {}
// null is a record column's missing value, rendered as "". It is not zero-sized, so two
// nulls are two map keys.
type null struct{ _ byte }
func (*null) isNode() {}
// template renders a format string, substituting {tokens} from fields. A bare
// JSON string is a template with no fields. repeat (default 1) renders that format
// that many times and joins the results with separator (default ""), each render
@@ -41,6 +47,7 @@ type template struct {
fields map[string]node
repeat int
separator string
datatype DataType
ops []op // format compiled once (see compileOps); what expand walks
grow int // minimum output size, to size the render buffer
fixed bool // no op varies, so every render is lit
@@ -66,26 +73,37 @@ func (t *template) field(seg string) (node, bool) {
return n, ok
}
// compile converts parsed JSON into a node tree, validating structure up front.
// Only a choice's items carry a weight, so one here would be inert whatever its type.
// compile converts parsed JSON — a category or an inline template — into a node tree,
// validating structure up front.
func compile(v any) (node, error) {
return compileAt(v, atTop)
}
// compileAt compiles a node that is no choice's item. Only a choice's items carry a
// weight, so one here would be inert whatever its type.
func compileAt(v any, pos position) (node, error) {
if m, ok := v.(map[string]any); ok {
if _, weighted := m["weight"]; weighted {
return nil, fmt.Errorf("weight only skews a choice's items, so it has no effect here; it is an option and can never be a field")
}
}
return compileItem(v)
return compileItem(v, pos)
}
// compileItem compiles one node, allowing the weight a choice item may carry.
func compileItem(v any) (node, error) {
func compileItem(v any, pos position) (node, error) {
switch v := v.(type) {
case string:
return compileString(v)
case []any:
return compileChoice(v)
return compileChoice(v, pos)
case map[string]any:
return compileTemplate(v)
return compileTemplate(v, pos)
case nil:
if pos != inColumn {
return nil, fmt.Errorf(`null is a record column's value; here it only renders "", so write ""`)
}
return &null{}, nil
default:
return nil, fmt.Errorf("a template value must be a string, a list or an object, not %s", jsonKind(v))
}
@@ -99,8 +117,6 @@ func jsonKind(v any) string {
return "a number"
case bool:
return "a boolean"
case nil:
return "null"
}
return fmt.Sprintf("%T", v)
}
@@ -136,7 +152,11 @@ func (t *template) compileFormat() error {
return checkNoRepeatedRead(t.format, c, t.refs)
}
func compileChoice(items []any) (node, error) {
func compileChoice(items []any, pos position) (node, error) {
itemPos := inFormat
if pos == inColumn {
itemPos = inColumn
}
if len(items) == 0 {
return nil, fmt.Errorf("empty choice")
}
@@ -160,7 +180,7 @@ func compileChoice(items []any) (node, error) {
}
total += w
cum[i] = total
n, err := compileItem(raw)
n, err := compileItem(raw, itemPos)
if err != nil {
return nil, err
}
@@ -203,22 +223,26 @@ func checkNoRepeatedItem(items []any) error {
return nil
}
func compileTemplate(m map[string]any) (node, error) {
o, err := readOptions(m)
func compileTemplate(m map[string]any, pos position) (node, error) {
o, err := readOptions(m, pos)
if err != nil {
return nil, err
}
fields, err := compileFields(m)
fieldPos := inFormat
if pos == atTop && o.repeat == 1 {
fieldPos = inColumn
}
fields, err := compileFields(m, fieldPos)
if err != nil {
return nil, err
}
if len(fields) == 0 && o.repeat == 1 && !o.weighted {
if len(fields) == 0 && o.repeat == 1 && !o.weighted && o.datatype == DataTypeString {
return nil, fmt.Errorf("an object holding only a format is a string; write %q", o.format)
}
if err := checkTokens(o.format, fields); err != nil {
return nil, err
}
t := &template{format: o.format, fields: fields, repeat: o.repeat, separator: o.separator}
t := &template{format: o.format, fields: fields, repeat: o.repeat, separator: o.separator, datatype: o.datatype}
if err := t.compileFormat(); err != nil {
return nil, err
}
@@ -227,13 +251,14 @@ func compileTemplate(m map[string]any) (node, error) {
// templateOptions is what a template object's option keys say.
type templateOptions struct {
datatype DataType
format string
repeat int
separator string
weighted bool
}
func readOptions(m map[string]any) (templateOptions, error) {
func readOptions(m map[string]any, pos position) (templateOptions, error) {
var o templateOptions
format, ok := m["format"].(string)
if !ok {
@@ -245,6 +270,9 @@ func readOptions(m map[string]any) (templateOptions, error) {
return o, err
}
o.repeat = repeat
if o.datatype, err = datatypeOf(m, pos); err != nil {
return o, err
}
if sv, ok := m["separator"]; ok {
if o.separator, ok = sv.(string); !ok {
return o, fmt.Errorf("separator must be a string, got %T", sv)
@@ -262,7 +290,7 @@ func readOptions(m map[string]any) (templateOptions, error) {
// compileFields compiles every non-option key of a template object, in name order
// so which of several bad fields is reported does not vary.
func compileFields(m map[string]any) (map[string]node, error) {
func compileFields(m map[string]any, pos position) (map[string]node, error) {
fields := make(map[string]node, len(m))
keys := make([]string, 0, len(m))
for k := range m {
@@ -276,7 +304,10 @@ func compileFields(m map[string]any) (map[string]node, error) {
if err := checkName(k); err != nil {
return nil, fmt.Errorf("field %w", err)
}
n, err := compile(m[k])
n, err := compileAt(m[k], pos)
if err == nil && pos == inColumn {
_, err = columnDatatype(n)
}
if err != nil {
return nil, fmt.Errorf("field %q: %w", k, err)
}
@@ -360,10 +391,10 @@ func checkName(name string) error {
}
// isOption reports whether a template key configures the node instead of naming a
// field. These four names can never be fields.
// field. These names can never be fields.
func isOption(name string) bool {
switch name {
case "format", "repeat", "separator", "weight":
case "datatype", "format", "repeat", "separator", "weight":
return true
}
return false
+1 -1
View File
@@ -60,7 +60,7 @@ func walkPath(n node, tail []string, w pathWalk) error {
}
return nil
}
return fmt.Errorf("cannot descend into %T at %q", n, tail[0])
return fmt.Errorf("no field %q", tail[0])
}
// carriedByAll is the choice rule a path that must resolve on every call obeys:
+84 -47
View File
@@ -9,10 +9,14 @@ import (
"strings"
)
// Column is one rendered column of a record.
// Column is one rendered column of a record. Value is the rendered text, which a
// serializer quotes for DataTypeString and writes bare for any other datatype; a Null
// column has no Value.
type Column struct {
Name string
DataType DataType
Value string
Null bool
}
// Record is one record rendered from a template: every direct field is a column,
@@ -28,70 +32,93 @@ func (r *Record) Columns() []Column {
return append([]Column(nil), r.columns...)
}
// JSON renders the record as one JSON object, every column a string.
// JSON renders the record as one JSON object.
func (r *Record) JSON() string {
m := make(map[string]string, len(r.columns))
for _, c := range r.columns {
m[c.Name] = c.Value
var b strings.Builder
b.WriteByte('{')
for i, c := range r.columns {
if i > 0 {
b.WriteByte(',')
}
b, _ := json.Marshal(m)
b.WriteString(jsonString(c.Name))
b.WriteByte(':')
b.WriteString(literal(c, jsonString, "null"))
}
b.WriteByte('}')
return b.String()
}
func jsonString(s string) string {
b, _ := json.Marshal(s)
return string(b)
}
// CSVHeader renders the column names as one CSV header line.
func (r *Record) CSVHeader() string {
return csvLine(r.names())
fields := make([]string, len(r.columns))
for i, c := range r.columns {
fields[i] = csvField(c.Name)
}
return strings.Join(fields, ",")
}
// CSVLine renders the column values as one CSV row.
// CSVLine renders the column values as one CSV row: a null column an empty field and an
// empty string "", the convention PostgreSQL's COPY reads a null by.
func (r *Record) CSVLine() string {
return csvLine(r.values())
}
func (r *Record) names() []string {
out := make([]string, len(r.columns))
fields := make([]string, len(r.columns))
for i, c := range r.columns {
out[i] = c.Name
}
return out
}
func (r *Record) values() []string {
out := make([]string, len(r.columns))
for i, c := range r.columns {
out[i] = c.Value
}
return out
}
func csvLine(cols []string) string {
var b strings.Builder
w := csv.NewWriter(&b)
_ = w.Write(cols)
w.Flush()
line := strings.TrimSuffix(b.String(), "\n")
if line == "" {
return `""` // a blank line is a row every CSV reader drops
fields[i] = literal(c, csvField, "")
}
if line := strings.Join(fields, ","); line != "" {
return line
}
return `""` // a blank line is a row every CSV reader drops
}
// SQLInsert renders the record as one INSERT statement into table: identifiers in
// ANSI double quotes, every value a single-quoted string literal.
func csvField(s string) string {
if s == "" {
return `""`
}
var b strings.Builder
w := csv.NewWriter(&b)
_ = w.Write([]string{s})
w.Flush()
return strings.TrimSuffix(b.String(), "\n")
}
// SQLInsert renders the record as one INSERT statement into table, identifiers in ANSI
// double quotes.
func (r *Record) SQLInsert(table string) string {
cols := make([]string, len(r.columns))
vals := make([]string, len(r.columns))
for i, c := range r.columns {
cols[i] = quoteIdent(c.Name)
vals[i] = "'" + strings.ReplaceAll(c.Value, "'", "''") + "'"
vals[i] = literal(c, sqlString, "NULL")
}
return fmt.Sprintf("INSERT INTO %s (%s) VALUES (%s);", quoteIdent(table), strings.Join(cols, ", "), strings.Join(vals, ", "))
}
func sqlString(s string) string {
return "'" + strings.ReplaceAll(s, "'", "''") + "'"
}
func quoteIdent(s string) string {
return `"` + strings.ReplaceAll(s, `"`, `""`) + `"`
}
// literal spells a column the way a serializer writes it: quoted for a string, bare for
// any other datatype, whose every render the load check proved a literal, and nullText
// for a null.
func literal(c Column, quote func(string) string, nullText string) string {
switch {
case c.Null:
return nullText
case c.DataType == DataTypeString:
return quote(c.Value)
}
return c.Value
}
// FakeRecord renders a path as one record: the template it names, with each direct
// field drawn as a column. Only a category-level template is a record — a path
// that descends into a field, or that names a folder or a choice, is an error.
@@ -116,7 +143,7 @@ func (f *Generator) FakeRecord(path string) (*Record, error) {
// columns, or why it is not a record.
type recordShape struct {
t *template
columns []string
columns []Column
err error
}
@@ -140,7 +167,7 @@ func (f *Generator) recordShapeOf(n node) recordShape {
type RecordTemplate struct {
g *Generator
t *template
columns []string
columns []Column
}
// Fake renders the record with one draw.
@@ -175,7 +202,7 @@ func (f *Generator) FakeRecordTemplate(input string) (*Record, error) {
// recordOf is the fence both record entry points pass. The columns come back with
// the template, fixed for every draw the caller goes on to make.
func recordOf(n node) (*template, []string, error) {
func recordOf(n node) (*template, []Column, error) {
t, ok := n.(*template)
if !ok {
return nil, nil, errors.New("names a choice, not a template; a record is a template whose fields are its columns")
@@ -183,13 +210,18 @@ func recordOf(n node) (*template, []string, error) {
if t.repeat != 1 {
return nil, nil, fmt.Errorf("carries repeat %d, which composes its format into one string; a record projects columns instead — drop the repeat and render the record again for more rows", t.repeat)
}
columns := recordColumns(t)
if len(columns) == 0 {
names := recordColumns(t)
if len(names) == 0 {
return nil, nil, errors.New("has no fields, so no columns")
}
if err := checkColumnRefs(t, columns); err != nil {
if err := checkColumnRefs(t, names); err != nil {
return nil, nil, err
}
columns := make([]Column, len(names))
for i, name := range names {
datatype, _ := columnDatatype(t.fields[name]) // compile refused a column whose items disagree
columns[i] = Column{Name: name, DataType: datatype}
}
return t, columns, nil
}
@@ -262,11 +294,16 @@ func columnRefs(t *template, columns []string) ([]columnRef, error) {
// renderRecord draws each column once, in the name order recordOf fixed, over one
// reference scope shared across them.
func renderRecord(s *session, t *template, columns []string) *Record {
func renderRecord(s *session, t *template, columns []Column) *Record {
scope := &draws{variant: map[string]node{}, value: map[string]string{}}
r := &Record{columns: make([]Column, len(columns))}
for i, name := range columns {
r.columns[i] = Column{Name: name, Value: render(s, t.fields[name], scope)}
r := &Record{columns: append([]Column(nil), columns...)}
for i := range r.columns {
n := drawn(s, t.fields[r.columns[i].Name])
if _, isNull := n.(*null); isNull {
r.columns[i].Null = true
} else {
r.columns[i].Value = render(s, n, scope)
}
}
return r
}
+2
View File
@@ -56,6 +56,8 @@ func render(s *session, n node, refScope *draws) string {
switch n := n.(type) {
case *choice:
return render(s, pick(s, n), refScope)
case *null:
return ""
case *template:
if n.repeat == 1 {
if n.fixed {
+403
View File
@@ -0,0 +1,403 @@
package fejkdata
import (
"slices"
"strconv"
"strings"
"unicode/utf8"
)
// grammar is a deterministic automaton over a scalar's text: state 0 is dead, 1 the
// start, and each state lists the runes that leave it and where they lead.
type grammar [][]arc
type arc struct {
on string
to int
}
func (g *grammar) run(q int, s string) int {
for _, r := range s {
if q = g.step(q, r); q == 0 {
return 0
}
}
return q
}
func (g *grammar) step(q int, r rune) int {
for _, a := range (*g)[q] {
if strings.ContainsRune(a.on, r) {
return a.to
}
}
return 0
}
const (
decimalDigits = "0123456789"
nonZeroDigits = "123456789"
)
// numberGrammar reads a JSON number. States: 2 "-", 3 "0", 4 more integer digits, 5 ".",
// 6 fraction digits, 7 "e", 8 its sign, 9 exponent digits.
var numberGrammar = &grammar{
nil,
{{"-", 2}, {"0", 3}, {nonZeroDigits, 4}},
{{"0", 3}, {nonZeroDigits, 4}},
{{".", 5}, {"eE", 7}},
{{decimalDigits, 4}, {".", 5}, {"eE", 7}},
{{decimalDigits, 6}},
{{decimalDigits, 6}, {"eE", 7}},
{{"+-", 8}, {decimalDigits, 9}},
{{decimalDigits, 9}},
{{decimalDigits, 9}},
}
const (
integerAccept uint32 = 1<<3 | 1<<4
numberAccept = integerAccept | 1<<6 | 1<<9
)
var booleanGrammar = &grammar{
nil,
{{"t", 2}, {"f", 6}},
{{"r", 3}}, {{"u", 4}}, {{"e", 5}}, nil,
{{"a", 7}}, {{"l", 8}}, {{"s", 9}}, {{"e", 10}}, nil,
}
const booleanAccept uint32 = 1<<5 | 1<<10
// decimalGrammar reads what a calc operand must render to be proven finite: a sign,
// digits and at most one dot. Past the sign, states 4–9 are positive and 10–15 their
// negatives: 4 zero digits, 5 a nonzero integer, 6 a leading dot, 7 zero with a dot,
// 8 a nonzero integer with a zero fraction, 9 a nonzero fraction.
var decimalGrammar = &grammar{
nil,
{{"+", 2}, {"-", 3}, {"0", 4}, {nonZeroDigits, 5}, {".", 6}},
{{"0", 4}, {nonZeroDigits, 5}, {".", 6}},
{{"0", 10}, {nonZeroDigits, 11}, {".", 12}},
{{"0", 4}, {nonZeroDigits, 5}, {".", 7}},
{{decimalDigits, 5}, {".", 8}},
{{"0", 7}, {nonZeroDigits, 9}},
{{"0", 7}, {nonZeroDigits, 9}},
{{"0", 8}, {nonZeroDigits, 9}},
{{decimalDigits, 9}},
{{"0", 10}, {nonZeroDigits, 11}, {".", 13}},
{{decimalDigits, 11}, {".", 14}},
{{"0", 13}, {nonZeroDigits, 15}},
{{"0", 13}, {nonZeroDigits, 15}},
{{"0", 14}, {nonZeroDigits, 15}},
{{decimalDigits, 15}},
}
const (
decimalAccept uint32 = 1<<4 | 1<<5 | 1<<7 | 1<<8 | 1<<9 | 1<<10 | 1<<11 | 1<<13 | 1<<14 | 1<<15
decimalNegative uint32 = 0xfc00
decimalZero uint32 = 1<<4 | 1<<7 | 1<<10 | 1<<13
decimalFractional uint32 = 1<<9 | 1<<15
)
// relation is what a node's renders do to a grammar: from each state, the states a
// render can end in, and one render reaching each.
type relation struct {
g *grammar
to []uint32
w []witness // w[from*len(to)+to]
}
// witness is one render, cut past witnessCap bytes, and why it can occur when the text
// alone does not say.
type witness struct {
text string
cut bool
why string
}
const witnessCap = 60
func (w witness) then(next witness) witness {
if w.why == "" {
w.why = next.why
}
if w.cut {
return w
}
w.text += next.text
w.cut = next.cut
if len(w.text) > witnessCap {
end := witnessCap
for !utf8.RuneStart(w.text[end]) {
end--
}
w.text, w.cut = w.text[:end], true
}
return w
}
func (w witness) String() string {
if w.cut {
return strconv.Quote(w.text + "…")
}
return strconv.Quote(w.text)
}
func newRelation(g *grammar) *relation {
n := len(*g)
return &relation{g: g, to: make([]uint32, n), w: make([]witness, n*n)}
}
func (r *relation) add(from, to int, w witness) {
if r.to[from]&(1<<to) == 0 {
r.to[from] |= 1 << to
r.w[from*len(r.to)+to] = w
}
}
// textRelation is the relation of a render that is always s.
func textRelation(g *grammar, s, why string) *relation {
r := newRelation(g)
w := witness{why: why}.then(witness{text: s})
for q := range r.to {
r.add(q, g.run(q, s), w)
}
return r
}
// union is the renders of either relation; a nil relation has none.
func union(a, b *relation) *relation {
if a == nil {
return b
}
if b == nil {
return a
}
u := newRelation(a.g)
for _, r := range []*relation{a, b} {
for from, ends := range r.to {
for to := range r.to {
if ends&(1<<to) != 0 {
u.add(from, to, r.w[from*len(r.to)+to])
}
}
}
}
return u
}
// then is a render of r followed by a render of next.
func (r *relation) then(next *relation) *relation {
c := newRelation(r.g)
n := len(r.to)
for from, mids := range r.to {
for mid := 0; mid < n; mid++ {
if mids&(1<<mid) == 0 {
continue
}
for to := 0; to < n; to++ {
if next.to[mid]&(1<<to) != 0 && c.to[from]&(1<<to) == 0 {
c.add(from, to, r.w[from*n+mid].then(next.w[mid*n+to]))
}
}
}
}
return c
}
// power is k renders of r in a row, k at least 1, composed by squaring.
func (r *relation) power(k int) *relation {
var out *relation
for base := r; ; base = base.then(base) {
if k&1 == 1 {
if out == nil {
out = base
} else {
out = out.then(base)
}
}
if k >>= 1; k == 0 {
return out
}
}
}
// closure is any number of renders of r in a row, where r includes the empty render.
func (r *relation) closure() *relation {
for {
next := r.then(r)
if slices.Equal(next.to, r.to) {
return r
}
r = next
}
}
// escape finds a render from the start that ends outside accept, preferring one that
// carries a reason.
func (r *relation) escape(accept uint32) (witness, bool) {
var found witness
escapes := false
for to := range r.to {
if (r.to[1]&^accept)&(1<<to) == 0 {
continue
}
if w := r.w[len(r.to)+to]; !escapes || found.why == "" && w.why != "" {
found, escapes = w, true
}
}
return found, escapes
}
// textShape is the text a builtin can emit: alternatives, each a sequence of runs.
type textShape [][]charRun
// charRun is between min and max characters, each one of chars; max -1 is unbounded.
// chars is ASCII, so a run of k characters is k bytes.
type charRun struct {
chars string
min, max int
}
// textLanguage is what one grammar makes of the renders a check reads, worked out once
// per node and fold.
type textLanguage struct {
g *grammar
proof *calcProof
memo map[languageKey]*relation
empty *relation
}
type languageKey struct {
n node
fold string
}
// fold is the transforms a render passes through before the grammar reads it, innermost
// first. Each rewrites rune by rune, so folding a render is folding each of its pieces.
type fold []string
func (f fold) apply(s string) string {
for _, name := range f {
s = transforms[name](s)
}
return s
}
func newTextLanguage(g *grammar, proof *calcProof) *textLanguage {
return &textLanguage{g: g, proof: proof, memo: map[languageKey]*relation{}, empty: textRelation(g, "", "")}
}
func (l *textLanguage) node(n node, f fold) *relation {
key := languageKey{n, strings.Join(f, ",")}
if r, done := l.memo[key]; done {
return r
}
r := l.empty // a null renders ""
switch n := n.(type) {
case *choice:
r = nil
for _, it := range n.items {
r = union(r, l.node(it, f))
}
case *template:
r = l.format(n, f)
if n.repeat > 1 {
r = r.then(l.text(f.apply(n.separator)).then(r).power(n.repeat - 1))
}
}
l.memo[key] = r
return r
}
func (l *textLanguage) text(s string) *relation { return textRelation(l.g, s, "") }
// format reads a template's format the way expand renders it: literal runs and tokens
// in turn.
func (l *textLanguage) format(t *template, f fold) *relation {
r := l.empty
var lit strings.Builder
_ = eachToken(t.format, func(tok ftoken) error {
if tok.kind == 'l' {
lit.WriteRune(tok.r)
return nil
}
r = r.then(l.text(f.apply(lit.String()))).then(l.token(t, tok.body, f))
lit.Reset()
return nil
})
return r.then(l.text(f.apply(lit.String())))
}
// token reads one {…} token: a field read, a transform over one, a calc, or what a
// builtin emits.
func (l *textLanguage) token(t *template, body string, f fold) *relation {
name, args, isFunc := funcCall(body)
if !isFunc {
var r *relation
for _, a := range splitArms(body, t.refs) {
r = union(r, l.read(t, a, f))
}
return r
}
if _, isTransform := transforms[name]; isTransform {
leaf, chain, _ := unwrapTransform(args[0])
inner := slices.Clone(chain)
slices.Reverse(inner)
return l.read(t, splitArm(leaf, t.refs), append(append(inner, name), f...))
}
if name == "calc" {
return l.calc(t, args, f)
}
return l.shape(builtins[name].emits(args), f)
}
// read is one arm of a token: every node its path can land on.
func (l *textLanguage) read(t *template, a arm, f fold) *relation {
var r *relation
for _, leaf := range pathLeaves(t.fields[a.key], a.tail) {
r = union(r, l.node(leaf, f))
}
return r
}
func (l *textLanguage) calc(t *template, args []string, f fold) *relation {
b, d := l.proof.call(t, args)
if d != nil {
return textRelation(l.g, f.apply(d.render), d.why)
}
return l.shape(printedFloat(b.lo, b.hi, calcDecimals(args), b.integral), f)
}
func (l *textLanguage) shape(s textShape, f fold) *relation {
var r *relation
for _, alt := range s {
seq := l.empty
for _, run := range alt {
seq = seq.then(l.run(run, f))
}
r = union(r, seq)
}
return r
}
// run reads a charRun: min characters, then up to max-min more.
func (l *textLanguage) run(c charRun, f fold) *relation {
one := newRelation(l.g)
for from := range one.to {
for _, ch := range c.chars {
s := f.apply(string(ch))
one.add(from, l.g.run(from, s), witness{text: s})
}
}
more := l.empty
switch optional := union(one, l.empty); {
case c.max < 0:
more = optional.closure()
case c.max > c.min:
more = optional.power(c.max - c.min)
}
if c.min == 0 {
return more
}
return one.power(c.min).then(more)
}
+3
View File
@@ -72,6 +72,9 @@ type builtin struct {
// operands names the fields the call reads, which expand renders for it; nil
// for a builtin that reads none.
operands func(args []string) []string
// emits is the text a call can print, for the datatype check; nil for calc and the
// transforms, whose text the check derives from what they read.
emits func(args []string) textShape
}
// funcCall splits a "{token}" body shaped name(args) into its parts; ok is false
-5
View File
@@ -6,11 +6,6 @@ The record API lands first, so the data update can use it.
### Record API
- Typed columns — a column declares its type, so `json` writes `42` rather than
`"42"` and `sql` an unquoted literal: string, integer, number, boolean, and a
way to write null. A template that can render a value its type rejects is a
load error. The option key is reserved from then on, so a common column name
like `type` is a poor pick.
- Struct-filling — fill a Go struct from `fake:"…"` tags holding a path or an
inline template, for parity with gofakeit and go-faker. The field's Go type is
the column type, through the same conversion and load checks as typed columns,