Typed columns and null: a datatype option, null items, and a load check that every typed render parses
Tests / vet + fmt + tests (pull_request) Successful in 58s

This commit is contained in:
2026-09-15 11:23:57 +02:00
parent 044294dd90
commit 94dc562c6c
13 changed files with 1106 additions and 107 deletions
+59 -9
View File
@@ -82,8 +82,8 @@ For structured output a record writes the row for you.
A record is a template seen as columns: its fields are the columns, its `format` A record is a template seen as columns: its fields are the columns, its `format`
the whole. `--format json|ndjson|csv|sql` writes the records; the library's the whole. `--format json|ndjson|csv|sql` writes the records; the library's
`FakeRecord` (below) hands back the columns. Every column is a string — typed scalars `FakeRecord` (below) hands back the columns. A column is a string unless it declares a
are on the release checklist, see [`todo.md`](todo.md). Save [datatype](#datatype), and a [`null`](#null) item draws it as null. Save
`mydata/users.json`: `mydata/users.json`:
```json ```json
@@ -164,7 +164,8 @@ r, err = f.FakeRecordTemplate(`{"format":"{x}","x":["a","b"]}`) // compile + ren
| `WithDataFS(fsys)` | layer an `fs.FS`, such as your own `embed.FS` | | `WithDataFS(fsys)` | layer an `fs.FS`, such as your own `embed.FS` |
| `WithoutShippedData()` | load only what you give | | `WithoutShippedData()` | load only what you give |
A `*Record` carries its columns via `Columns()`, and serializes them with `JSON()` A `*Record` carries its columns via `Columns()` — each a `Column` of `Name`,
`DataType`, rendered `Value` and `Null` — and serializes them with `JSON()`
(one object), `CSVHeader()`/`CSVLine()`, or `SQLInsert(table)` — the shapes the (one object), `CSVHeader()`/`CSVLine()`, or `SQLInsert(table)` — the shapes the
CLI's `--format` writes. `FakeRecord` and `FakeRecordTemplate` take a record; a CLI's `--format` writes. `FakeRecord` and `FakeRecordTemplate` take a record; a
path or template that is not one — a bare string, a choice, or a folder — errors. path or template that is not one — a bare string, a choice, or a folder — errors.
@@ -255,10 +256,53 @@ Renders e.g. `bar foo baz`. Rejected at load: a `separator` without a `repeat`,
a `separator` of `""` (the default), and a `repeat` that multiplies to more than a `separator` of `""` (the default), and a `repeat` that multiplies to more than
1 048 576 renders along any path of nested repeats. 1 048 576 renders along any path of nested repeats.
### Datatype
A record column may declare `datatype` — `integer`, `number` or `boolean` — so `json`
writes `42` rather than `"42"` and `sql` a bare literal; a column without one is a
string:
```json
{ "format": "",
"id": { "format": "{seq()}", "datatype": "integer" },
"paid": { "format": "{p}", "p": ["true", "false"], "datatype": "boolean" },
"total": { "format": "{calc(net * qty, 2)}", "net": ["19.99", "5.00"], "qty": ["3", "7"], "datatype": "number" } }
```
Writes e.g. `{"id":1,"paid":true,"total":59.97}`. A column is a field of the top-level
template, or an item of a choice standing in for one; `datatype` anywhere else is a
load error. So is a column that can render text its datatype rejects — `integer` takes
`-?(0|[1-9][0-9]*)`, `number` a JSON number, `boolean` `true` or `false` — and the
error shows such a render:
```text
order.id: datatype integer, but it can render "000", which is not an integer
```
A `{calc()}` fills an `integer` or `number` column only where it provably prints no
`NaN` or `Inf`: each operand is a plain decimal — a sign, digits, one dot — of at most
300 bytes, or a field holding only such a calc, and no divisor can be zero. An
`integer` column also needs a decimals count of `0`, or integer operands and no `/`.
### Null
A `null` item draws a record column as null: `json` writes `null`, `sql` `NULL`, and
`csv` an empty field, with an empty string written `""` so PostgreSQL's `COPY … CSV`
reads both back. `Fake` renders a null as `""`. The other items' weights skew its
odds:
```json
{ "format": "", "deleted_at": null, "middle": [null, { "format": "{n}", "n": ["Ann", "Eva"], "weight": 3 }] }
```
`deleted_at` is null every draw, `middle` a name three draws in four. Rejected at
load: `null` anywhere but a column, naming `""`, and a column whose items declare
different datatypes.
### Options and fields ### Options and fields
`format`, `weight`, `repeat` and `separator` are the only options; **any other `format`, `weight`, `repeat`, `separator` and `datatype` are the only options; **any
key is a field** (see [Decisions](#decisions)). An object that does nothing a other key is a field** (see [Decisions](#decisions)). An object that does nothing a
string can't — only a `format` — is rejected naming the string, as is a one-item string can't — only a `format` — is rejected naming the string, as is a one-item
choice naming its item. choice naming its item.
@@ -437,8 +481,8 @@ tokens add cost in proportion to the output.
## Decisions ## Decisions
- **Options and fields share one namespace.** `format`, `weight`, `repeat` and - **Options and fields share one namespace.** `format`, `weight`, `repeat`,
`separator` are reserved; every other key is a field. Nesting fields under a `separator` and `datatype` are reserved; every other key is a field. Nesting fields under a
key, or prefixing options, would tax every template to guard against a key, or prefixing options, would tax every template to guard against a
misspelt option. misspelt option.
- **`{a|b}` stays beside nested choices.** `[[…], […]]` picks the same way, but - **`{a|b}` stays beside nested choices.** `[[…], […]]` picks the same way, but
@@ -518,7 +562,8 @@ tokens add cost in proportion to the output.
so `a/(b*c)` with `b` fixed at `0` and `c` varying loads and prints `Inf` every so `a/(b*c)` with `b` fixed at `0` and `c` varying loads and prints `Inf` every
draw — catching it needs zero-absorbing algebra for a shape nobody writes. draw — catching it needs zero-absorbing algebra for a shape nobody writes.
- **In data, a default written out and a constant spelled as a sample are load - **In data, a default written out and a constant spelled as a sample are load
errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `int(5,5)`, `float(1,1,2)`, errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `datatype: "string"`,
`int(5,5)`, `float(1,1,2)`,
`+5` and `05` each spell what a shorter form already spells, so each is rejected `+5` and `05` each spell what a shorter form already spells, so each is rejected
naming that form. The CLI's numbers follow the shell instead: `--seed 007` and naming that form. The CLI's numbers follow the shell instead: `--seed 007` and
`--repeat +3` are 7 and 3, as every command line reads them. `--repeat +3` are 7 and 3, as every command line reads them.
@@ -557,6 +602,9 @@ tokens add cost in proportion to the output.
row — would vary per draw. A fixed column set is what the CSV and `INSERT` row — would vary per draw. A fixed column set is what the CSV and `INSERT`
contracts rest on, so the restriction holds even where a particular choice would contracts rest on, so the restriction holds even where a particular choice would
happen to agree. happen to agree.
- **Null is a `null` item, not a rate.** A null is one more outcome of a column's
draw, so a choice's weights skew it like any other; a null-rate option would be a
second way to state odds.
- **The performance gate asserts allocations, not wall-clock time.** `AllocsPerRun` - **The performance gate asserts allocations, not wall-clock time.** `AllocsPerRun`
is deterministic across machines, so a ±10% ceiling does not flake under CI load, is deterministic across machines, so a ±10% ceiling does not flake under CI load,
while time varies with the machine and its neighbours. A rendering slowdown while time varies with the machine and its neighbours. A rendering slowdown
@@ -612,7 +660,9 @@ hold.go the hold: one draw per expansion for paths and operands, and its
reference.go reference sigils, and binding references across the tree reference.go reference sigils, and binding references across the tree
graph.go the render graph: edges, cycles, the repeat bound, tree walks graph.go the render graph: edges, cycles, the repeat bound, tree walks
builtins.go the {name()} function registry and its implementations builtins.go the {name()} function registry and its implementations
calc.go the {calc()} arithmetic evaluator: parser, eval, validation calc.go the {calc()} arithmetic evaluator: parser, eval, validation, and the proof a typed column's calc is finite
datatype.go column datatypes: DataType, where datatype and null may sit, and the load check every typed render passes
renderlang.go what text a node can render, as relations over a scalar's grammar
data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge
cmd/fejkdata/ the fejkdata CLI cmd/fejkdata/ the fejkdata CLI
data/ shipped data (JSON), embedded at build: locale folders + a misc folder data/ shipped data (JSON), embedded at build: locale folders + a misc folder
+101 -21
View File
@@ -5,6 +5,7 @@ import (
"errors" "errors"
"fmt" "fmt"
"math" "math"
"slices"
"strconv" "strconv"
"strings" "strings"
"unicode" "unicode"
@@ -24,36 +25,36 @@ const (
// samples read only the rng. A time-based id (uuid v7, ulid) draws its timestamp // samples read only the rng. A time-based id (uuid v7, ulid) draws its timestamp
// from the rng, not the wall clock, so seeded output stays reproducible. // from the rng, not the wall clock, so seeded output stays reproducible.
var builtins = map[string]builtin{ var builtins = map[string]builtin{
"luhn": {arity: 0, prep: derive(func(e string) string { return string(rune('0' + luhnCheck(e))) })}, "luhn": {arity: 0, prep: derive(func(e string) string { return string(rune('0' + luhnCheck(e))) }), emits: always(textShape{{{decimalDigits, 1, 1}}})},
"mod11": {arity: 0, prep: derive(mod11Check)}, "mod11": {arity: 0, prep: derive(mod11Check), emits: always(textShape{{{decimalDigits + "X", 1, 1}}})},
"ean": {arity: 0, prep: derive(eanCheck)}, "ean": {arity: 0, prep: derive(eanCheck), emits: always(textShape{{{decimalDigits, 1, 1}}})},
"uuid": {arity: 0, prep: sample(uuidV7)}, "uuid": {arity: 0, prep: sample(uuidV7), emits: always(uuidShape)},
"ulid": {arity: 0, prep: sample(ulid)}, "ulid": {arity: 0, prep: sample(ulid), emits: always(textShape{{{crockford[:8], 1, 1}, {crockford, 25, 25}}})},
"nanoid": {arity: 1, check: posIntArg, prep: chars(nanoidAlphabet)}, "nanoid": sampleOf(nanoidAlphabet),
"hex": {arity: 1, check: posIntArg, prep: chars(hexDigits)}, "hex": sampleOf(hexDigits),
"digits": {arity: 1, check: posIntArg, prep: chars("0123456789")}, "digits": sampleOf(decimalDigits),
"upper": {arity: 1, check: posIntArg, prep: chars("ABCDEFGHIJKLMNOPQRSTUVWXYZ")}, "upper": sampleOf("ABCDEFGHIJKLMNOPQRSTUVWXYZ"),
"lower": {arity: 1, check: posIntArg, prep: chars("abcdefghijklmnopqrstuvwxyz")}, "lower": sampleOf("abcdefghijklmnopqrstuvwxyz"),
"base64": {arity: 1, check: posIntArg, prep: func(a []string) callFn { "base64": {arity: 1, check: posIntArg, prep: func(a []string) callFn {
n := atoi(a[0]) n := atoi(a[0])
return func(s *session, _ string, _ []string) string { return func(s *session, _ string, _ []string) string {
return base64.StdEncoding.EncodeToString(randBytes(s, n)) return base64.StdEncoding.EncodeToString(randBytes(s, n))
} }
}}, }, emits: base64Shape},
"int": {arity: 2, check: intRangeArgs, prep: func(a []string) callFn { "int": {arity: 2, check: intRangeArgs, prep: func(a []string) callFn {
lo, span := atoi(a[0]), atoi(a[1])-atoi(a[0])+1 lo, span := atoi(a[0]), atoi(a[1])-atoi(a[0])+1
return func(s *session, _ string, _ []string) string { return strconv.Itoa(lo + s.IntN(span)) } return func(s *session, _ string, _ []string) string { return strconv.Itoa(lo + s.IntN(span)) }
}}, }, emits: intShape},
"float": {arity: 3, check: floatArgs, prep: func(a []string) callFn { "float": {arity: 3, check: floatArgs, prep: func(a []string) callFn {
lo, hi, dp := atof(a[0]), atof(a[1]), atoi(a[2]) lo, hi, dp := atof(a[0]), atof(a[1]), atoi(a[2])
return func(s *session, _ string, _ []string) string { return func(s *session, _ string, _ []string) string {
return strconv.FormatFloat(lo+s.Float64()*(hi-lo), 'f', dp, 64) return strconv.FormatFloat(lo+s.Float64()*(hi-lo), 'f', dp, 64)
} }
}}, }, emits: func(a []string) textShape { return printedFloat(atof(a[0]), atof(a[1]), atoi(a[2]), false) }},
"iban": {arity: 1, check: ibanArg, prep: func(a []string) callFn { "iban": {arity: 1, check: ibanArg, prep: func(a []string) callFn {
cc := a[0] cc := a[0]
return func(s *session, _ string, _ []string) string { return iban(s, cc) } return func(s *session, _ string, _ []string) string { return iban(s, cc) }
}}, }, emits: ibanShape},
"calc": {arity: -1, check: checkCalc, prep: calcPrep, operands: calcOperands}, "calc": {arity: -1, check: checkCalc, prep: calcPrep, operands: calcOperands},
"lowercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToLower), operands: transformOperand}, "lowercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToLower), operands: transformOperand},
"uppercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToUpper), operands: transformOperand}, "uppercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToUpper), operands: transformOperand},
@@ -69,7 +70,7 @@ var builtins = map[string]builtin{
return func(s *session, _ string, _ []string) string { return func(s *session, _ string, _ []string) string {
return strconv.FormatUint(s.next(key), 10) return strconv.FormatUint(s.next(key), 10)
} }
}}, }, emits: always(textShape{{{nonZeroDigits, 1, 1}, {decimalDigits, 0, 19}}})},
} }
// derive and sample are the two argument-free builtin shapes: a derivation reads // derive and sample are the two argument-free builtin shapes: a derivation reads
@@ -94,6 +95,82 @@ func chars(alphabet string) func([]string) callFn {
} }
} }
// sampleOf is the builtin that draws n characters from an alphabet.
func sampleOf(alphabet string) builtin {
return builtin{arity: 1, check: posIntArg, prep: chars(alphabet), emits: func(a []string) textShape {
n := atoi(a[0])
return textShape{{{alphabet, n, n}}}
}}
}
// always is the emits of a builtin whose args do not change what it can print.
func always(s textShape) func([]string) textShape {
return func([]string) textShape { return s }
}
var uuidShape = textShape{{{hexDigits, 8, 8}, {"-", 1, 1}, {hexDigits, 4, 4}, {"-", 1, 1}, {"7", 1, 1}, {hexDigits, 3, 3}, {"-", 1, 1}, {"89ab", 1, 1}, {hexDigits, 3, 3}, {"-", 1, 1}, {hexDigits, 12, 12}}}
const base64Alphabet = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/"
func base64Shape(a []string) textShape {
n := atoi(a[0])
pad := (3 - n%3) % 3
size := 4*((n+2)/3) - pad
return textShape{{{base64Alphabet, size, size}, {"=", pad, pad}}}
}
// intShape is what int prints: a sign only below zero, and no leading zero.
func intShape(a []string) textShape {
lo, hi := atoi(a[0]), atoi(a[1])
var s textShape
if lo <= 0 && hi >= 0 {
s = append(s, []charRun{{"0", 1, 1}})
}
if hi > 0 {
s = append(s, []charRun{{nonZeroDigits, 1, 1}, {decimalDigits, 0, len(a[1]) - 1}})
}
if lo < 0 {
s = append(s, []charRun{{"-", 1, 1}, {nonZeroDigits, 1, 1}, {decimalDigits, 0, len(a[0]) - 2}})
}
return s
}
func ibanShape(a []string) textShape {
cc, digits := a[0], ibanLen[a[0]]-2
return textShape{{{cc[:1], 1, 1}, {cc[1:], 1, 1}, {decimalDigits, digits, digits}}}
}
// shortestFraction bounds the fraction FormatFloat's shortest form prints: at most 17
// significant digits after up to 323 zeros.
const shortestFraction = 340
// printedFloat is what strconv.FormatFloat(v, 'f', dp, 64) prints for a v in [lo, hi]
// that is whole when integral.
func printedFloat(lo, hi float64, dp int, integral bool) textShape {
digits := len(strconv.FormatFloat(math.Floor(math.Max(math.Abs(lo), math.Abs(hi))), 'f', 0, 64)) + 1 // one more for a rounding carry
wholes := [][]charRun{{{"0", 1, 1}}, {{nonZeroDigits, 1, 1}, {decimalDigits, 0, digits - 1}}}
fractions := [][]charRun{nil}
switch {
case dp > 0:
fractions = [][]charRun{{{".", 1, 1}, {decimalDigits, dp, dp}}}
case dp < 0 && !integral:
fractions = append(fractions, []charRun{{".", 1, 1}, {decimalDigits, 1, shortestFraction}})
}
signs := [][]charRun{nil}
if lo < 0 || math.Signbit(lo) {
signs = append(signs, []charRun{{"-", 1, 1}})
}
var s textShape
for _, sign := range signs {
for _, whole := range wholes {
for _, fraction := range fractions {
s = append(s, slices.Concat(sign, whole, fraction))
}
}
}
return s
}
const hexDigits = "0123456789abcdef" const hexDigits = "0123456789abcdef"
// transforms are the builtins that rewrite one operand's value; they nest, so // transforms are the builtins that rewrite one operand's value; they nest, so
@@ -106,20 +183,19 @@ var transforms = map[string]func(string) string{
// unwrapTransform peels nested transform calls off an operand arg, returning the // unwrapTransform peels nested transform calls off an operand arg, returning the
// field it finally names and the transforms to apply, innermost last. // field it finally names and the transforms to apply, innermost last.
func unwrapTransform(arg string) (leaf string, chain []func(string) string, err error) { func unwrapTransform(arg string) (leaf string, chain []string, err error) {
for { for {
name, args, isCall := funcCall(arg) name, args, isCall := funcCall(arg)
if !isCall { if !isCall {
return arg, chain, nil return arg, chain, nil
} }
fn, isTransform := transforms[name] if _, isTransform := transforms[name]; !isTransform {
if !isTransform {
return "", nil, fmt.Errorf("%s(%s) is not a transform, so it cannot be an operand", name, strings.Join(args, ",")) return "", nil, fmt.Errorf("%s(%s) is not a transform, so it cannot be an operand", name, strings.Join(args, ","))
} }
if len(args) != 1 { if len(args) != 1 {
return "", nil, fmt.Errorf("%s takes 1 arg, got %d", name, len(args)) return "", nil, fmt.Errorf("%s takes 1 arg, got %d", name, len(args))
} }
chain = append(chain, fn) chain = append(chain, name)
arg = args[0] arg = args[0]
} }
} }
@@ -150,10 +226,14 @@ func transformPrep(outer func(string) string) func([]string) callFn {
if err != nil { if err != nil {
panic(fmt.Sprintf("fejkdata: transform arg %q reached prep unvalidated: %v", a[0], err)) panic(fmt.Sprintf("fejkdata: transform arg %q reached prep unvalidated: %v", a[0], err))
} }
fns := make([]func(string) string, len(chain))
for i, name := range chain {
fns[i] = transforms[name]
}
return func(_ *session, _ string, operands []string) string { return func(_ *session, _ string, operands []string) string {
v := operands[0] v := operands[0]
for i := len(chain) - 1; i >= 0; i-- { for i := len(fns) - 1; i >= 0; i-- {
v = chain[i](v) v = fns[i](v)
} }
return outer(v) return outer(v)
} }
+259 -6
View File
@@ -6,6 +6,7 @@ import (
"strconv" "strconv"
"strings" "strings"
"unicode" "unicode"
"unicode/utf8"
) )
// calcNode is a parsed expression node. It evaluates over the operand values expand // calcNode is a parsed expression node. It evaluates over the operand values expand
@@ -154,10 +155,12 @@ func calcText(n calcNode) string {
return "?" return "?"
} }
// neverNumeric reports a node no render of which is a number: fixed text that does // neverNumeric reports a node no render of which is a number: a null, fixed text that
// not parse, or a choice of only such items. text is one such render. // does not parse, or a choice of only such items. text is one such render.
func neverNumeric(n node) (text string, never bool) { func neverNumeric(n node) (text string, never bool) {
switch n := n.(type) { switch n := n.(type) {
case *null:
return "", true
case *template: case *template:
if !n.fixed || n.repeat > 1 { if !n.fixed || n.repeat > 1 {
return "", false return "", false
@@ -191,10 +194,7 @@ func calcPrep(args []string) callFn {
at[name] = i at[name] = i
} }
placed := indexVars(expr, at) placed := indexVars(expr, at)
dp := -1 dp := calcDecimals(args)
if len(args) == 2 {
dp = atoi(args[1])
}
return func(_ *session, _ string, operands []string) string { return func(_ *session, _ string, operands []string) string {
return strconv.FormatFloat(placed.eval(operands), 'f', dp, 64) return strconv.FormatFloat(placed.eval(operands), 'f', dp, 64)
} }
@@ -384,3 +384,256 @@ func contains(bs []byte, b byte) bool {
} }
return false return false
} }
// calcDecimals is a calc's decimals count, or -1 for the shortest form.
func calcDecimals(args []string) int {
if len(args) == 2 {
return atoi(args[1])
}
return -1
}
// calcLimit is the largest magnitude a proof accepts as finite, far enough below
// math.MaxFloat64 that rounding in the bounds cannot hide an overflow.
const calcLimit = 1e300
// maxOperandLen is the longest operand text a proof bounds by its length, so that
// bound, 10^maxOperandLen, stays within calcLimit.
const maxOperandLen = 300
// calcBound is what a proof knows of every value a calc can take: it lies in [lo, hi],
// is at least nonZero from zero unless nonZero is 0, and is whole when integral.
type calcBound struct {
lo, hi, nonZero float64
integral bool
}
func magnitude(b calcBound) float64 { return math.Max(math.Abs(b.lo), math.Abs(b.hi)) }
// doubt is why a proof could not show a calc finite, and the render that shows it.
type doubt struct{ render, why string }
type bounded struct {
b calcBound
d *doubt
}
// calcProof bounds a typed column's calcs from their operands' renders, to show each
// prints a number rather than NaN or Inf.
type calcProof struct {
decimal *textLanguage
operands map[node]bounded
lengths map[node]int
}
func newCalcProof() *calcProof {
p := &calcProof{operands: map[node]bounded{}, lengths: map[node]int{}}
p.decimal = newTextLanguage(decimalGrammar, p)
return p
}
// call bounds one calc token of t.
func (p *calcProof) call(t *template, args []string) (calcBound, *doubt) {
expr, err := parseCalc(args[0])
if err != nil {
panic(fmt.Sprintf("fejkdata: calc(%q) reached a proof unparsed: %v", args[0], err))
}
b, d := p.expr(expr, t.fields)
if d != nil {
return b, &doubt{d.render, fmt.Sprintf("{calc(%s)}: %s", strings.Join(args, ", "), d.why)}
}
return b, nil
}
func (p *calcProof) expr(n calcNode, fields map[string]node) (calcBound, *doubt) {
switch n := n.(type) {
case calcNum:
v := float64(n)
return calcBound{v, v, v, v == math.Trunc(v)}, nil
case calcVar:
return p.operand(string(n), fields[string(n)])
case calcNeg:
b, d := p.expr(n.x, fields)
return calcBound{-b.hi, -b.lo, b.nonZero, b.integral}, d
case calcBin:
l, d := p.expr(n.l, fields)
if d != nil {
return l, d
}
r, d := p.expr(n.r, fields)
if d != nil {
return r, d
}
return combine(n, l, r)
}
panic(fmt.Sprintf("fejkdata: calc node %T has no bound", n))
}
// combine bounds one operation from the bounds of its sides.
func combine(n calcBin, l, r calcBound) (calcBound, *doubt) {
b := calcBound{integral: l.integral && r.integral}
switch n.op {
case '+':
b.lo, b.hi = l.lo+r.lo, l.hi+r.hi
case '-':
b.lo, b.hi = l.lo-r.hi, l.hi-r.lo
case '*':
b.lo = min(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi)
b.hi = max(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi)
b.nonZero = l.nonZero * r.nonZero
default:
if r.nonZero == 0 {
return b, &doubt{"+Inf", fmt.Sprintf("divides by %s, which can be zero", calcText(n.r))}
}
m := magnitude(l) / r.nonZero
b = calcBound{lo: -m, hi: m, nonZero: l.nonZero / magnitude(r)}
}
if b.lo > 0 || b.hi < 0 {
b.nonZero = math.Max(b.nonZero, math.Min(math.Abs(b.lo), math.Abs(b.hi)))
}
if !(magnitude(b) <= calcLimit) {
return b, &doubt{"+Inf", calcText(n) + " can overflow"}
}
return b, nil
}
// operand bounds a calc operand, once per node.
func (p *calcProof) operand(name string, n node) (calcBound, *doubt) {
if seen, done := p.operands[n]; done {
return seen.b, seen.d
}
b, d := p.measure(name, n)
p.operands[n] = bounded{b, d}
return b, d
}
// measure bounds an operand through the calc it renders when that is all it renders,
// and otherwise from its text: a plain decimal of at most maxOperandLen bytes.
func (p *calcProof) measure(name string, n node) (calcBound, *doubt) {
if t, ok := n.(*template); ok {
if args, isCalc := soleCalc(t); isCalc {
b, d := p.call(t, args)
return rounded(b, calcDecimals(args)), d
}
}
text := p.decimal.node(n, nil)
if w, escapes := text.escape(decimalAccept); escapes {
why := fmt.Sprintf("operand %q can render %s, which is not a plain decimal", name, w)
if w.why != "" {
why += ": " + w.why
}
return calcBound{}, &doubt{"NaN", why}
}
size := p.length(n)
if size > maxOperandLen {
return calcBound{}, &doubt{"NaN", fmt.Sprintf("operand %q can render more than %d bytes, too many to bound", name, maxOperandLen)}
}
ends, m := text.to[1], math.Pow(10, float64(size))
b := calcBound{hi: m, nonZero: 1 / m, integral: ends&decimalFractional == 0}
if ends&decimalNegative != 0 {
b.lo = -m
}
if ends&decimalZero != 0 {
b.nonZero = 0
}
return b, nil
}
// soleCalc reports a template that renders one calc and nothing else, with its args.
func soleCalc(t *template) ([]string, bool) {
if t.repeat != 1 || len(t.ops) != 1 || t.ops[0].kind != 'b' {
return nil, false
}
name, args, _ := funcCall(t.format[1 : len(t.format)-1])
return args, name == "calc"
}
// rounded is b once printed to dp decimals, which moves a value by up to half a unit.
func rounded(b calcBound, dp int) calcBound {
if dp < 0 {
return b
}
half := math.Pow(10, -float64(dp)) / 2
return calcBound{b.lo - half, b.hi + half, math.Max(0, b.nonZero-half), b.integral || dp == 0}
}
// length is the most bytes a render of n can take, anything past maxOperandLen
// reported as maxOperandLen+1.
func (p *calcProof) length(n node) int {
if size, done := p.lengths[n]; done {
return size
}
size := 0
switch n := n.(type) {
case *choice:
for _, it := range n.items {
size = max(size, p.length(it))
}
case *template:
size = p.formatLength(n)*n.repeat + len(n.separator)*(n.repeat-1)
}
size = min(size, maxOperandLen+1)
p.lengths[n] = size
return size
}
func (p *calcProof) formatLength(t *template) int {
size := 0
_ = eachToken(t.format, func(tok ftoken) error {
if tok.kind == 'l' {
size += utf8.RuneLen(tok.r)
} else {
size += p.tokenLength(t, tok.body)
}
size = min(size, maxOperandLen+1)
return nil
})
return size
}
// tokenLength is the most bytes one token can print. A transform never lengthens a
// render that reads as a decimal: it maps each non-ASCII rune, two bytes or more, to at
// most two ASCII letters.
func (p *calcProof) tokenLength(t *template, body string) int {
name, args, isFunc := funcCall(body)
var arms []arm
switch _, isTransform := transforms[name]; {
case !isFunc:
arms = splitArms(body, t.refs)
case isTransform:
leaf, _, _ := unwrapTransform(args[0])
arms = []arm{splitArm(leaf, t.refs)}
case name == "calc":
b, d := p.call(t, args)
if d != nil {
return len(d.render)
}
return shapeLength(printedFloat(b.lo, b.hi, calcDecimals(args), b.integral))
default:
return shapeLength(builtins[name].emits(args))
}
size := 0
for _, a := range arms {
for _, leaf := range pathLeaves(t.fields[a.key], a.tail) {
size = max(size, p.length(leaf))
}
}
return size
}
// shapeLength is the most bytes a shape can emit, anything past maxOperandLen reported
// as maxOperandLen+1.
func shapeLength(s textShape) int {
longest := 0
for _, alt := range s {
size := 0
for _, run := range alt {
if run.max < 0 {
return maxOperandLen + 1
}
size += run.max
}
longest = max(longest, size)
}
return min(longest, maxOperandLen+1)
}
+140
View File
@@ -0,0 +1,140 @@
package fejkdata
import (
"errors"
"fmt"
)
// DataType is what a record column holds, which decides how a record writes its value.
type DataType int
// The datatypes a column declares with "datatype"; a column without one is a string.
const (
DataTypeString DataType = iota
DataTypeInteger
DataTypeNumber
DataTypeBoolean
)
var dataTypeNames = [...]string{"string", "integer", "number", "boolean"}
// String is the datatype as data spells it.
func (d DataType) String() string {
if d < 0 || int(d) >= len(dataTypeNames) {
return fmt.Sprintf("DataType(%d)", int(d))
}
return dataTypeNames[d]
}
// position is where a JSON value sits, which decides whether it may carry a datatype or
// be null.
type position int
const (
inFormat position = iota // rendered by a format, so neither
atTop // a category or an inline template, whose fields are the columns
inColumn // a column, or a choice item standing in for one
)
// datatypeOf reads a template's "datatype" (default DataTypeString).
func datatypeOf(m map[string]any, pos position) (DataType, error) {
v, ok := m["datatype"]
if !ok {
return DataTypeString, nil
}
name, ok := v.(string)
if !ok {
return 0, fmt.Errorf("datatype must be a string, got %T", v)
}
if name == DataTypeString.String() {
return 0, fmt.Errorf("datatype %q is the default, so it has no effect; drop it", name)
}
for d := DataTypeInteger; d <= DataTypeBoolean; d++ {
if name != d.String() {
continue
}
if pos != inColumn {
return 0, errors.New("datatype only types a record column — a field of the top-level template — so it has no effect here")
}
return d, nil
}
return 0, fmt.Errorf(`datatype takes "integer", "number" or "boolean", got %q`, name)
}
// columnDatatype is the datatype a column's items declare. They must agree, since a
// column holds one; a column only ever null is a string.
func columnDatatype(n node) (DataType, error) {
var declared []DataType
var collect func(node)
collect = func(n node) {
switch n := n.(type) {
case *choice:
for _, it := range n.items {
collect(it)
}
case *template:
declared = append(declared, n.datatype)
}
}
collect(n)
if len(declared) == 0 {
return DataTypeString, nil
}
for _, d := range declared {
if d != declared[0] {
return declared[0], fmt.Errorf("its items declare %s and %s; a column holds one datatype, so give every item the same", declared[0], d)
}
}
return declared[0], nil
}
// datatypeSpec is what a datatype's text must satisfy: a grammar, the states a render
// may end in, and how an error names the datatype.
type datatypeSpec struct {
grammar *grammar
accept uint32
noun string
}
var datatypeSpecs = map[DataType]datatypeSpec{
DataTypeInteger: {numberGrammar, integerAccept, "an integer"},
DataTypeNumber: {numberGrammar, numberAccept, "a number"},
DataTypeBoolean: {booleanGrammar, booleanAccept, "a boolean"},
}
// datatypeCheck proves every render of a typed column is text its datatype takes. One
// check covers a scope, so a node several columns reach is read once per grammar.
type datatypeCheck struct {
languages map[*grammar]*textLanguage
proof *calcProof
}
func (c *datatypeCheck) check(path string, n node) error {
t, ok := n.(*template)
if !ok || t.datatype == DataTypeString {
return nil
}
spec := datatypeSpecs[t.datatype]
w, escapes := c.language(spec.grammar).node(t, nil).escape(spec.accept)
if !escapes {
return nil
}
msg := fmt.Sprintf("%s: datatype %s, but it can render %s, which is not %s", path, t.datatype, w, spec.noun)
if w.why != "" {
msg += ": " + w.why
}
return errors.New(msg)
}
func (c *datatypeCheck) language(g *grammar) *textLanguage {
if c.proof == nil {
c.proof = newCalcProof()
c.languages = map[*grammar]*textLanguage{}
}
l, made := c.languages[g]
if !made {
l = newTextLanguage(g, c.proof)
c.languages[g] = l
}
return l
}
+2
View File
@@ -162,6 +162,8 @@ func paths(n node) []string {
} }
} }
return out return out
case *null:
return []string{""}
case *choice: case *choice:
out := []string{""} out := []string{""}
for p := range n.shared { for p := range n.shared {
+4 -1
View File
@@ -186,7 +186,10 @@ func checkScope(s nodeScope) error {
if err := s(func(path string, n node) error { return repeatCheck(path, n, mem) }); err != nil { if err := s(func(path string, n node) error { return repeatCheck(path, n, mem) }); err != nil {
return err return err
} }
return s(heldCheck) if err := s(heldCheck); err != nil {
return err
}
return s((&datatypeCheck{}).check)
} }
type reachMemo map[node]int type reachMemo map[node]int
+51 -20
View File
@@ -32,6 +32,12 @@ type choice struct {
func (*choice) isNode() {} func (*choice) isNode() {}
// null is a record column's missing value, rendered as "". It is not zero-sized, so two
// nulls are two map keys.
type null struct{ _ byte }
func (*null) isNode() {}
// template renders a format string, substituting {tokens} from fields. A bare // template renders a format string, substituting {tokens} from fields. A bare
// JSON string is a template with no fields. repeat (default 1) renders that format // JSON string is a template with no fields. repeat (default 1) renders that format
// that many times and joins the results with separator (default ""), each render // that many times and joins the results with separator (default ""), each render
@@ -41,6 +47,7 @@ type template struct {
fields map[string]node fields map[string]node
repeat int repeat int
separator string separator string
datatype DataType
ops []op // format compiled once (see compileOps); what expand walks ops []op // format compiled once (see compileOps); what expand walks
grow int // minimum output size, to size the render buffer grow int // minimum output size, to size the render buffer
fixed bool // no op varies, so every render is lit fixed bool // no op varies, so every render is lit
@@ -66,26 +73,37 @@ func (t *template) field(seg string) (node, bool) {
return n, ok return n, ok
} }
// compile converts parsed JSON into a node tree, validating structure up front. // compile converts parsed JSON — a category or an inline template — into a node tree,
// Only a choice's items carry a weight, so one here would be inert whatever its type. // validating structure up front.
func compile(v any) (node, error) { func compile(v any) (node, error) {
return compileAt(v, atTop)
}
// compileAt compiles a node that is no choice's item. Only a choice's items carry a
// weight, so one here would be inert whatever its type.
func compileAt(v any, pos position) (node, error) {
if m, ok := v.(map[string]any); ok { if m, ok := v.(map[string]any); ok {
if _, weighted := m["weight"]; weighted { if _, weighted := m["weight"]; weighted {
return nil, fmt.Errorf("weight only skews a choice's items, so it has no effect here; it is an option and can never be a field") return nil, fmt.Errorf("weight only skews a choice's items, so it has no effect here; it is an option and can never be a field")
} }
} }
return compileItem(v) return compileItem(v, pos)
} }
// compileItem compiles one node, allowing the weight a choice item may carry. // compileItem compiles one node, allowing the weight a choice item may carry.
func compileItem(v any) (node, error) { func compileItem(v any, pos position) (node, error) {
switch v := v.(type) { switch v := v.(type) {
case string: case string:
return compileString(v) return compileString(v)
case []any: case []any:
return compileChoice(v) return compileChoice(v, pos)
case map[string]any: case map[string]any:
return compileTemplate(v) return compileTemplate(v, pos)
case nil:
if pos != inColumn {
return nil, fmt.Errorf(`null is a record column's value; here it only renders "", so write ""`)
}
return &null{}, nil
default: default:
return nil, fmt.Errorf("a template value must be a string, a list or an object, not %s", jsonKind(v)) return nil, fmt.Errorf("a template value must be a string, a list or an object, not %s", jsonKind(v))
} }
@@ -99,8 +117,6 @@ func jsonKind(v any) string {
return "a number" return "a number"
case bool: case bool:
return "a boolean" return "a boolean"
case nil:
return "null"
} }
return fmt.Sprintf("%T", v) return fmt.Sprintf("%T", v)
} }
@@ -136,7 +152,11 @@ func (t *template) compileFormat() error {
return checkNoRepeatedRead(t.format, c, t.refs) return checkNoRepeatedRead(t.format, c, t.refs)
} }
func compileChoice(items []any) (node, error) { func compileChoice(items []any, pos position) (node, error) {
itemPos := inFormat
if pos == inColumn {
itemPos = inColumn
}
if len(items) == 0 { if len(items) == 0 {
return nil, fmt.Errorf("empty choice") return nil, fmt.Errorf("empty choice")
} }
@@ -160,7 +180,7 @@ func compileChoice(items []any) (node, error) {
} }
total += w total += w
cum[i] = total cum[i] = total
n, err := compileItem(raw) n, err := compileItem(raw, itemPos)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -203,22 +223,26 @@ func checkNoRepeatedItem(items []any) error {
return nil return nil
} }
func compileTemplate(m map[string]any) (node, error) { func compileTemplate(m map[string]any, pos position) (node, error) {
o, err := readOptions(m) o, err := readOptions(m, pos)
if err != nil { if err != nil {
return nil, err return nil, err
} }
fields, err := compileFields(m) fieldPos := inFormat
if pos == atTop && o.repeat == 1 {
fieldPos = inColumn
}
fields, err := compileFields(m, fieldPos)
if err != nil { if err != nil {
return nil, err return nil, err
} }
if len(fields) == 0 && o.repeat == 1 && !o.weighted { if len(fields) == 0 && o.repeat == 1 && !o.weighted && o.datatype == DataTypeString {
return nil, fmt.Errorf("an object holding only a format is a string; write %q", o.format) return nil, fmt.Errorf("an object holding only a format is a string; write %q", o.format)
} }
if err := checkTokens(o.format, fields); err != nil { if err := checkTokens(o.format, fields); err != nil {
return nil, err return nil, err
} }
t := &template{format: o.format, fields: fields, repeat: o.repeat, separator: o.separator} t := &template{format: o.format, fields: fields, repeat: o.repeat, separator: o.separator, datatype: o.datatype}
if err := t.compileFormat(); err != nil { if err := t.compileFormat(); err != nil {
return nil, err return nil, err
} }
@@ -227,13 +251,14 @@ func compileTemplate(m map[string]any) (node, error) {
// templateOptions is what a template object's option keys say. // templateOptions is what a template object's option keys say.
type templateOptions struct { type templateOptions struct {
datatype DataType
format string format string
repeat int repeat int
separator string separator string
weighted bool weighted bool
} }
func readOptions(m map[string]any) (templateOptions, error) { func readOptions(m map[string]any, pos position) (templateOptions, error) {
var o templateOptions var o templateOptions
format, ok := m["format"].(string) format, ok := m["format"].(string)
if !ok { if !ok {
@@ -245,6 +270,9 @@ func readOptions(m map[string]any) (templateOptions, error) {
return o, err return o, err
} }
o.repeat = repeat o.repeat = repeat
if o.datatype, err = datatypeOf(m, pos); err != nil {
return o, err
}
if sv, ok := m["separator"]; ok { if sv, ok := m["separator"]; ok {
if o.separator, ok = sv.(string); !ok { if o.separator, ok = sv.(string); !ok {
return o, fmt.Errorf("separator must be a string, got %T", sv) return o, fmt.Errorf("separator must be a string, got %T", sv)
@@ -262,7 +290,7 @@ func readOptions(m map[string]any) (templateOptions, error) {
// compileFields compiles every non-option key of a template object, in name order // compileFields compiles every non-option key of a template object, in name order
// so which of several bad fields is reported does not vary. // so which of several bad fields is reported does not vary.
func compileFields(m map[string]any) (map[string]node, error) { func compileFields(m map[string]any, pos position) (map[string]node, error) {
fields := make(map[string]node, len(m)) fields := make(map[string]node, len(m))
keys := make([]string, 0, len(m)) keys := make([]string, 0, len(m))
for k := range m { for k := range m {
@@ -276,7 +304,10 @@ func compileFields(m map[string]any) (map[string]node, error) {
if err := checkName(k); err != nil { if err := checkName(k); err != nil {
return nil, fmt.Errorf("field %w", err) return nil, fmt.Errorf("field %w", err)
} }
n, err := compile(m[k]) n, err := compileAt(m[k], pos)
if err == nil && pos == inColumn {
_, err = columnDatatype(n)
}
if err != nil { if err != nil {
return nil, fmt.Errorf("field %q: %w", k, err) return nil, fmt.Errorf("field %q: %w", k, err)
} }
@@ -360,10 +391,10 @@ func checkName(name string) error {
} }
// isOption reports whether a template key configures the node instead of naming a // isOption reports whether a template key configures the node instead of naming a
// field. These four names can never be fields. // field. These names can never be fields.
func isOption(name string) bool { func isOption(name string) bool {
switch name { switch name {
case "format", "repeat", "separator", "weight": case "datatype", "format", "repeat", "separator", "weight":
return true return true
} }
return false return false
+1 -1
View File
@@ -60,7 +60,7 @@ func walkPath(n node, tail []string, w pathWalk) error {
} }
return nil return nil
} }
return fmt.Errorf("cannot descend into %T at %q", n, tail[0]) return fmt.Errorf("no field %q", tail[0])
} }
// carriedByAll is the choice rule a path that must resolve on every call obeys: // carriedByAll is the choice rule a path that must resolve on every call obeys:
+79 -42
View File
@@ -9,10 +9,14 @@ import (
"strings" "strings"
) )
// Column is one rendered column of a record. // Column is one rendered column of a record. Value is the rendered text, which a
// serializer quotes for DataTypeString and writes bare for any other datatype; a Null
// column has no Value.
type Column struct { type Column struct {
Name string Name string
DataType DataType
Value string Value string
Null bool
} }
// Record is one record rendered from a template: every direct field is a column, // Record is one record rendered from a template: every direct field is a column,
@@ -28,70 +32,93 @@ func (r *Record) Columns() []Column {
return append([]Column(nil), r.columns...) return append([]Column(nil), r.columns...)
} }
// JSON renders the record as one JSON object, every column a string. // JSON renders the record as one JSON object.
func (r *Record) JSON() string { func (r *Record) JSON() string {
m := make(map[string]string, len(r.columns)) var b strings.Builder
for _, c := range r.columns { b.WriteByte('{')
m[c.Name] = c.Value for i, c := range r.columns {
if i > 0 {
b.WriteByte(',')
} }
b, _ := json.Marshal(m) b.WriteString(jsonString(c.Name))
b.WriteByte(':')
b.WriteString(literal(c, jsonString, "null"))
}
b.WriteByte('}')
return b.String()
}
func jsonString(s string) string {
b, _ := json.Marshal(s)
return string(b) return string(b)
} }
// CSVHeader renders the column names as one CSV header line. // CSVHeader renders the column names as one CSV header line.
func (r *Record) CSVHeader() string { func (r *Record) CSVHeader() string {
return csvLine(r.names()) fields := make([]string, len(r.columns))
for i, c := range r.columns {
fields[i] = csvField(c.Name)
}
return strings.Join(fields, ",")
} }
// CSVLine renders the column values as one CSV row. // CSVLine renders the column values as one CSV row: a null column an empty field and an
// empty string "", the convention PostgreSQL's COPY reads a null by.
func (r *Record) CSVLine() string { func (r *Record) CSVLine() string {
return csvLine(r.values()) fields := make([]string, len(r.columns))
}
func (r *Record) names() []string {
out := make([]string, len(r.columns))
for i, c := range r.columns { for i, c := range r.columns {
out[i] = c.Name fields[i] = literal(c, csvField, "")
} }
return out if line := strings.Join(fields, ","); line != "" {
return line
}
return `""` // a blank line is a row every CSV reader drops
} }
func (r *Record) values() []string { func csvField(s string) string {
out := make([]string, len(r.columns)) if s == "" {
for i, c := range r.columns { return `""`
out[i] = c.Value
} }
return out
}
func csvLine(cols []string) string {
var b strings.Builder var b strings.Builder
w := csv.NewWriter(&b) w := csv.NewWriter(&b)
_ = w.Write(cols) _ = w.Write([]string{s})
w.Flush() w.Flush()
line := strings.TrimSuffix(b.String(), "\n") return strings.TrimSuffix(b.String(), "\n")
if line == "" {
return `""` // a blank line is a row every CSV reader drops
}
return line
} }
// SQLInsert renders the record as one INSERT statement into table: identifiers in // SQLInsert renders the record as one INSERT statement into table, identifiers in ANSI
// ANSI double quotes, every value a single-quoted string literal. // double quotes.
func (r *Record) SQLInsert(table string) string { func (r *Record) SQLInsert(table string) string {
cols := make([]string, len(r.columns)) cols := make([]string, len(r.columns))
vals := make([]string, len(r.columns)) vals := make([]string, len(r.columns))
for i, c := range r.columns { for i, c := range r.columns {
cols[i] = quoteIdent(c.Name) cols[i] = quoteIdent(c.Name)
vals[i] = "'" + strings.ReplaceAll(c.Value, "'", "''") + "'" vals[i] = literal(c, sqlString, "NULL")
} }
return fmt.Sprintf("INSERT INTO %s (%s) VALUES (%s);", quoteIdent(table), strings.Join(cols, ", "), strings.Join(vals, ", ")) return fmt.Sprintf("INSERT INTO %s (%s) VALUES (%s);", quoteIdent(table), strings.Join(cols, ", "), strings.Join(vals, ", "))
} }
func sqlString(s string) string {
return "'" + strings.ReplaceAll(s, "'", "''") + "'"
}
func quoteIdent(s string) string { func quoteIdent(s string) string {
return `"` + strings.ReplaceAll(s, `"`, `""`) + `"` return `"` + strings.ReplaceAll(s, `"`, `""`) + `"`
} }
// literal spells a column the way a serializer writes it: quoted for a string, bare for
// any other datatype, whose every render the load check proved a literal, and nullText
// for a null.
func literal(c Column, quote func(string) string, nullText string) string {
switch {
case c.Null:
return nullText
case c.DataType == DataTypeString:
return quote(c.Value)
}
return c.Value
}
// FakeRecord renders a path as one record: the template it names, with each direct // FakeRecord renders a path as one record: the template it names, with each direct
// field drawn as a column. Only a category-level template is a record — a path // field drawn as a column. Only a category-level template is a record — a path
// that descends into a field, or that names a folder or a choice, is an error. // that descends into a field, or that names a folder or a choice, is an error.
@@ -116,7 +143,7 @@ func (f *Generator) FakeRecord(path string) (*Record, error) {
// columns, or why it is not a record. // columns, or why it is not a record.
type recordShape struct { type recordShape struct {
t *template t *template
columns []string columns []Column
err error err error
} }
@@ -140,7 +167,7 @@ func (f *Generator) recordShapeOf(n node) recordShape {
type RecordTemplate struct { type RecordTemplate struct {
g *Generator g *Generator
t *template t *template
columns []string columns []Column
} }
// Fake renders the record with one draw. // Fake renders the record with one draw.
@@ -175,7 +202,7 @@ func (f *Generator) FakeRecordTemplate(input string) (*Record, error) {
// recordOf is the fence both record entry points pass. The columns come back with // recordOf is the fence both record entry points pass. The columns come back with
// the template, fixed for every draw the caller goes on to make. // the template, fixed for every draw the caller goes on to make.
func recordOf(n node) (*template, []string, error) { func recordOf(n node) (*template, []Column, error) {
t, ok := n.(*template) t, ok := n.(*template)
if !ok { if !ok {
return nil, nil, errors.New("names a choice, not a template; a record is a template whose fields are its columns") return nil, nil, errors.New("names a choice, not a template; a record is a template whose fields are its columns")
@@ -183,13 +210,18 @@ func recordOf(n node) (*template, []string, error) {
if t.repeat != 1 { if t.repeat != 1 {
return nil, nil, fmt.Errorf("carries repeat %d, which composes its format into one string; a record projects columns instead — drop the repeat and render the record again for more rows", t.repeat) return nil, nil, fmt.Errorf("carries repeat %d, which composes its format into one string; a record projects columns instead — drop the repeat and render the record again for more rows", t.repeat)
} }
columns := recordColumns(t) names := recordColumns(t)
if len(columns) == 0 { if len(names) == 0 {
return nil, nil, errors.New("has no fields, so no columns") return nil, nil, errors.New("has no fields, so no columns")
} }
if err := checkColumnRefs(t, columns); err != nil { if err := checkColumnRefs(t, names); err != nil {
return nil, nil, err return nil, nil, err
} }
columns := make([]Column, len(names))
for i, name := range names {
datatype, _ := columnDatatype(t.fields[name]) // compile refused a column whose items disagree
columns[i] = Column{Name: name, DataType: datatype}
}
return t, columns, nil return t, columns, nil
} }
@@ -262,11 +294,16 @@ func columnRefs(t *template, columns []string) ([]columnRef, error) {
// renderRecord draws each column once, in the name order recordOf fixed, over one // renderRecord draws each column once, in the name order recordOf fixed, over one
// reference scope shared across them. // reference scope shared across them.
func renderRecord(s *session, t *template, columns []string) *Record { func renderRecord(s *session, t *template, columns []Column) *Record {
scope := &draws{variant: map[string]node{}, value: map[string]string{}} scope := &draws{variant: map[string]node{}, value: map[string]string{}}
r := &Record{columns: make([]Column, len(columns))} r := &Record{columns: append([]Column(nil), columns...)}
for i, name := range columns { for i := range r.columns {
r.columns[i] = Column{Name: name, Value: render(s, t.fields[name], scope)} n := drawn(s, t.fields[r.columns[i].Name])
if _, isNull := n.(*null); isNull {
r.columns[i].Null = true
} else {
r.columns[i].Value = render(s, n, scope)
}
} }
return r return r
} }
+2
View File
@@ -56,6 +56,8 @@ func render(s *session, n node, refScope *draws) string {
switch n := n.(type) { switch n := n.(type) {
case *choice: case *choice:
return render(s, pick(s, n), refScope) return render(s, pick(s, n), refScope)
case *null:
return ""
case *template: case *template:
if n.repeat == 1 { if n.repeat == 1 {
if n.fixed { if n.fixed {
+403
View File
@@ -0,0 +1,403 @@
package fejkdata
import (
"slices"
"strconv"
"strings"
"unicode/utf8"
)
// grammar is a deterministic automaton over a scalar's text: state 0 is dead, 1 the
// start, and each state lists the runes that leave it and where they lead.
type grammar [][]arc
type arc struct {
on string
to int
}
func (g *grammar) run(q int, s string) int {
for _, r := range s {
if q = g.step(q, r); q == 0 {
return 0
}
}
return q
}
func (g *grammar) step(q int, r rune) int {
for _, a := range (*g)[q] {
if strings.ContainsRune(a.on, r) {
return a.to
}
}
return 0
}
const (
decimalDigits = "0123456789"
nonZeroDigits = "123456789"
)
// numberGrammar reads a JSON number. States: 2 "-", 3 "0", 4 more integer digits, 5 ".",
// 6 fraction digits, 7 "e", 8 its sign, 9 exponent digits.
var numberGrammar = &grammar{
nil,
{{"-", 2}, {"0", 3}, {nonZeroDigits, 4}},
{{"0", 3}, {nonZeroDigits, 4}},
{{".", 5}, {"eE", 7}},
{{decimalDigits, 4}, {".", 5}, {"eE", 7}},
{{decimalDigits, 6}},
{{decimalDigits, 6}, {"eE", 7}},
{{"+-", 8}, {decimalDigits, 9}},
{{decimalDigits, 9}},
{{decimalDigits, 9}},
}
const (
integerAccept uint32 = 1<<3 | 1<<4
numberAccept = integerAccept | 1<<6 | 1<<9
)
var booleanGrammar = &grammar{
nil,
{{"t", 2}, {"f", 6}},
{{"r", 3}}, {{"u", 4}}, {{"e", 5}}, nil,
{{"a", 7}}, {{"l", 8}}, {{"s", 9}}, {{"e", 10}}, nil,
}
const booleanAccept uint32 = 1<<5 | 1<<10
// decimalGrammar reads what a calc operand must render to be proven finite: a sign,
// digits and at most one dot. Past the sign, states 4–9 are positive and 10–15 their
// negatives: 4 zero digits, 5 a nonzero integer, 6 a leading dot, 7 zero with a dot,
// 8 a nonzero integer with a zero fraction, 9 a nonzero fraction.
var decimalGrammar = &grammar{
nil,
{{"+", 2}, {"-", 3}, {"0", 4}, {nonZeroDigits, 5}, {".", 6}},
{{"0", 4}, {nonZeroDigits, 5}, {".", 6}},
{{"0", 10}, {nonZeroDigits, 11}, {".", 12}},
{{"0", 4}, {nonZeroDigits, 5}, {".", 7}},
{{decimalDigits, 5}, {".", 8}},
{{"0", 7}, {nonZeroDigits, 9}},
{{"0", 7}, {nonZeroDigits, 9}},
{{"0", 8}, {nonZeroDigits, 9}},
{{decimalDigits, 9}},
{{"0", 10}, {nonZeroDigits, 11}, {".", 13}},
{{decimalDigits, 11}, {".", 14}},
{{"0", 13}, {nonZeroDigits, 15}},
{{"0", 13}, {nonZeroDigits, 15}},
{{"0", 14}, {nonZeroDigits, 15}},
{{decimalDigits, 15}},
}
const (
decimalAccept uint32 = 1<<4 | 1<<5 | 1<<7 | 1<<8 | 1<<9 | 1<<10 | 1<<11 | 1<<13 | 1<<14 | 1<<15
decimalNegative uint32 = 0xfc00
decimalZero uint32 = 1<<4 | 1<<7 | 1<<10 | 1<<13
decimalFractional uint32 = 1<<9 | 1<<15
)
// relation is what a node's renders do to a grammar: from each state, the states a
// render can end in, and one render reaching each.
type relation struct {
g *grammar
to []uint32
w []witness // w[from*len(to)+to]
}
// witness is one render, cut past witnessCap bytes, and why it can occur when the text
// alone does not say.
type witness struct {
text string
cut bool
why string
}
const witnessCap = 60
func (w witness) then(next witness) witness {
if w.why == "" {
w.why = next.why
}
if w.cut {
return w
}
w.text += next.text
w.cut = next.cut
if len(w.text) > witnessCap {
end := witnessCap
for !utf8.RuneStart(w.text[end]) {
end--
}
w.text, w.cut = w.text[:end], true
}
return w
}
func (w witness) String() string {
if w.cut {
return strconv.Quote(w.text + "…")
}
return strconv.Quote(w.text)
}
func newRelation(g *grammar) *relation {
n := len(*g)
return &relation{g: g, to: make([]uint32, n), w: make([]witness, n*n)}
}
func (r *relation) add(from, to int, w witness) {
if r.to[from]&(1<<to) == 0 {
r.to[from] |= 1 << to
r.w[from*len(r.to)+to] = w
}
}
// textRelation is the relation of a render that is always s.
func textRelation(g *grammar, s, why string) *relation {
r := newRelation(g)
w := witness{why: why}.then(witness{text: s})
for q := range r.to {
r.add(q, g.run(q, s), w)
}
return r
}
// union is the renders of either relation; a nil relation has none.
func union(a, b *relation) *relation {
if a == nil {
return b
}
if b == nil {
return a
}
u := newRelation(a.g)
for _, r := range []*relation{a, b} {
for from, ends := range r.to {
for to := range r.to {
if ends&(1<<to) != 0 {
u.add(from, to, r.w[from*len(r.to)+to])
}
}
}
}
return u
}
// then is a render of r followed by a render of next.
func (r *relation) then(next *relation) *relation {
c := newRelation(r.g)
n := len(r.to)
for from, mids := range r.to {
for mid := 0; mid < n; mid++ {
if mids&(1<<mid) == 0 {
continue
}
for to := 0; to < n; to++ {
if next.to[mid]&(1<<to) != 0 && c.to[from]&(1<<to) == 0 {
c.add(from, to, r.w[from*n+mid].then(next.w[mid*n+to]))
}
}
}
}
return c
}
// power is k renders of r in a row, k at least 1, composed by squaring.
func (r *relation) power(k int) *relation {
var out *relation
for base := r; ; base = base.then(base) {
if k&1 == 1 {
if out == nil {
out = base
} else {
out = out.then(base)
}
}
if k >>= 1; k == 0 {
return out
}
}
}
// closure is any number of renders of r in a row, where r includes the empty render.
func (r *relation) closure() *relation {
for {
next := r.then(r)
if slices.Equal(next.to, r.to) {
return r
}
r = next
}
}
// escape finds a render from the start that ends outside accept, preferring one that
// carries a reason.
func (r *relation) escape(accept uint32) (witness, bool) {
var found witness
escapes := false
for to := range r.to {
if (r.to[1]&^accept)&(1<<to) == 0 {
continue
}
if w := r.w[len(r.to)+to]; !escapes || found.why == "" && w.why != "" {
found, escapes = w, true
}
}
return found, escapes
}
// textShape is the text a builtin can emit: alternatives, each a sequence of runs.
type textShape [][]charRun
// charRun is between min and max characters, each one of chars; max -1 is unbounded.
// chars is ASCII, so a run of k characters is k bytes.
type charRun struct {
chars string
min, max int
}
// textLanguage is what one grammar makes of the renders a check reads, worked out once
// per node and fold.
type textLanguage struct {
g *grammar
proof *calcProof
memo map[languageKey]*relation
empty *relation
}
type languageKey struct {
n node
fold string
}
// fold is the transforms a render passes through before the grammar reads it, innermost
// first. Each rewrites rune by rune, so folding a render is folding each of its pieces.
type fold []string
func (f fold) apply(s string) string {
for _, name := range f {
s = transforms[name](s)
}
return s
}
func newTextLanguage(g *grammar, proof *calcProof) *textLanguage {
return &textLanguage{g: g, proof: proof, memo: map[languageKey]*relation{}, empty: textRelation(g, "", "")}
}
func (l *textLanguage) node(n node, f fold) *relation {
key := languageKey{n, strings.Join(f, ",")}
if r, done := l.memo[key]; done {
return r
}
r := l.empty // a null renders ""
switch n := n.(type) {
case *choice:
r = nil
for _, it := range n.items {
r = union(r, l.node(it, f))
}
case *template:
r = l.format(n, f)
if n.repeat > 1 {
r = r.then(l.text(f.apply(n.separator)).then(r).power(n.repeat - 1))
}
}
l.memo[key] = r
return r
}
func (l *textLanguage) text(s string) *relation { return textRelation(l.g, s, "") }
// format reads a template's format the way expand renders it: literal runs and tokens
// in turn.
func (l *textLanguage) format(t *template, f fold) *relation {
r := l.empty
var lit strings.Builder
_ = eachToken(t.format, func(tok ftoken) error {
if tok.kind == 'l' {
lit.WriteRune(tok.r)
return nil
}
r = r.then(l.text(f.apply(lit.String()))).then(l.token(t, tok.body, f))
lit.Reset()
return nil
})
return r.then(l.text(f.apply(lit.String())))
}
// token reads one {…} token: a field read, a transform over one, a calc, or what a
// builtin emits.
func (l *textLanguage) token(t *template, body string, f fold) *relation {
name, args, isFunc := funcCall(body)
if !isFunc {
var r *relation
for _, a := range splitArms(body, t.refs) {
r = union(r, l.read(t, a, f))
}
return r
}
if _, isTransform := transforms[name]; isTransform {
leaf, chain, _ := unwrapTransform(args[0])
inner := slices.Clone(chain)
slices.Reverse(inner)
return l.read(t, splitArm(leaf, t.refs), append(append(inner, name), f...))
}
if name == "calc" {
return l.calc(t, args, f)
}
return l.shape(builtins[name].emits(args), f)
}
// read is one arm of a token: every node its path can land on.
func (l *textLanguage) read(t *template, a arm, f fold) *relation {
var r *relation
for _, leaf := range pathLeaves(t.fields[a.key], a.tail) {
r = union(r, l.node(leaf, f))
}
return r
}
func (l *textLanguage) calc(t *template, args []string, f fold) *relation {
b, d := l.proof.call(t, args)
if d != nil {
return textRelation(l.g, f.apply(d.render), d.why)
}
return l.shape(printedFloat(b.lo, b.hi, calcDecimals(args), b.integral), f)
}
func (l *textLanguage) shape(s textShape, f fold) *relation {
var r *relation
for _, alt := range s {
seq := l.empty
for _, run := range alt {
seq = seq.then(l.run(run, f))
}
r = union(r, seq)
}
return r
}
// run reads a charRun: min characters, then up to max-min more.
func (l *textLanguage) run(c charRun, f fold) *relation {
one := newRelation(l.g)
for from := range one.to {
for _, ch := range c.chars {
s := f.apply(string(ch))
one.add(from, l.g.run(from, s), witness{text: s})
}
}
more := l.empty
switch optional := union(one, l.empty); {
case c.max < 0:
more = optional.closure()
case c.max > c.min:
more = optional.power(c.max - c.min)
}
if c.min == 0 {
return more
}
return one.power(c.min).then(more)
}
+3
View File
@@ -72,6 +72,9 @@ type builtin struct {
// operands names the fields the call reads, which expand renders for it; nil // operands names the fields the call reads, which expand renders for it; nil
// for a builtin that reads none. // for a builtin that reads none.
operands func(args []string) []string operands func(args []string) []string
// emits is the text a call can print, for the datatype check; nil for calc and the
// transforms, whose text the check derives from what they read.
emits func(args []string) textShape
} }
// funcCall splits a "{token}" body shaped name(args) into its parts; ok is false // funcCall splits a "{token}" body shaped name(args) into its parts; ok is false
-5
View File
@@ -6,11 +6,6 @@ The record API lands first, so the data update can use it.
### Record API ### Record API
- Typed columns — a column declares its type, so `json` writes `42` rather than
`"42"` and `sql` an unquoted literal: string, integer, number, boolean, and a
way to write null. A template that can render a value its type rejects is a
load error. The option key is reserved from then on, so a common column name
like `type` is a poor pick.
- Struct-filling — fill a Go struct from `fake:"…"` tags holding a path or an - Struct-filling — fill a Go struct from `fake:"…"` tags holding a path or an
inline template, for parity with gofakeit and go-faker. The field's Go type is inline template, for parity with gofakeit and go-faker. The field's Go type is
the column type, through the same conversion and load checks as typed columns, the column type, through the same conversion and load checks as typed columns,