Typed columns and null #13
@@ -82,8 +82,8 @@ For structured output a record writes the row for you.
|
|||||||
|
|
||||||
A record is a template seen as columns: its fields are the columns, its `format`
|
A record is a template seen as columns: its fields are the columns, its `format`
|
||||||
the whole. `--format json|ndjson|csv|sql` writes the records; the library's
|
the whole. `--format json|ndjson|csv|sql` writes the records; the library's
|
||||||
`FakeRecord` (below) hands back the columns. Every column is a string — typed scalars
|
`FakeRecord` (below) hands back the columns. A column is a string unless it declares a
|
||||||
are on the release checklist, see [`todo.md`](todo.md). Save
|
[datatype](#datatype), and a [`null`](#null) item draws it as null. Save
|
||||||
`mydata/users.json`:
|
`mydata/users.json`:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
@@ -164,7 +164,8 @@ r, err = f.FakeRecordTemplate(`{"format":"{x}","x":["a","b"]}`) // compile + ren
|
|||||||
| `WithDataFS(fsys)` | layer an `fs.FS`, such as your own `embed.FS` |
|
| `WithDataFS(fsys)` | layer an `fs.FS`, such as your own `embed.FS` |
|
||||||
| `WithoutShippedData()` | load only what you give |
|
| `WithoutShippedData()` | load only what you give |
|
||||||
|
|
||||||
A `*Record` carries its columns via `Columns()`, and serializes them with `JSON()`
|
A `*Record` carries its columns via `Columns()` — each a `Column` of `Name`,
|
||||||
|
`DataType`, rendered `Value` and `Null` — and serializes them with `JSON()`
|
||||||
(one object), `CSVHeader()`/`CSVLine()`, or `SQLInsert(table)` — the shapes the
|
(one object), `CSVHeader()`/`CSVLine()`, or `SQLInsert(table)` — the shapes the
|
||||||
CLI's `--format` writes. `FakeRecord` and `FakeRecordTemplate` take a record; a
|
CLI's `--format` writes. `FakeRecord` and `FakeRecordTemplate` take a record; a
|
||||||
path or template that is not one — a bare string, a choice, or a folder — errors.
|
path or template that is not one — a bare string, a choice, or a folder — errors.
|
||||||
@@ -255,10 +256,53 @@ Renders e.g. `bar foo baz`. Rejected at load: a `separator` without a `repeat`,
|
|||||||
a `separator` of `""` (the default), and a `repeat` that multiplies to more than
|
a `separator` of `""` (the default), and a `repeat` that multiplies to more than
|
||||||
1 048 576 renders along any path of nested repeats.
|
1 048 576 renders along any path of nested repeats.
|
||||||
|
|
||||||
|
### Datatype
|
||||||
|
|
||||||
|
A record column may declare `datatype` — `integer`, `number` or `boolean` — so `json`
|
||||||
|
writes `42` rather than `"42"` and `sql` a bare literal; a column without one is a
|
||||||
|
string:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{ "format": "",
|
||||||
|
"id": { "format": "{seq()}", "datatype": "integer" },
|
||||||
|
"paid": { "format": "{p}", "p": ["true", "false"], "datatype": "boolean" },
|
||||||
|
"total": { "format": "{calc(net * qty, 2)}", "net": ["19.99", "5.00"], "qty": ["3", "7"], "datatype": "number" } }
|
||||||
|
```
|
||||||
|
|
||||||
|
Writes e.g. `{"id":1,"paid":true,"total":59.97}`. A column is a field of the top-level
|
||||||
|
template, or an item of a choice standing in for one; `datatype` anywhere else is a
|
||||||
|
load error. So is a column that can render text its datatype rejects — `integer` takes
|
||||||
|
`-?(0|[1-9][0-9]*)`, `number` a JSON number, `boolean` `true` or `false` — and the
|
||||||
|
error shows such a render:
|
||||||
|
|
||||||
|
```text
|
||||||
|
order.id: datatype integer, but it can render "000", which is not an integer
|
||||||
|
```
|
||||||
|
|
||||||
|
A `{calc()}` fills an `integer` or `number` column only where it provably prints no
|
||||||
|
`NaN` or `Inf`: each operand is a plain decimal — a sign, digits, one dot — of at most
|
||||||
|
300 bytes, or a field holding only such a calc, and no divisor can be zero. An
|
||||||
|
`integer` column also needs a decimals count of `0`, or integer operands and no `/`.
|
||||||
|
|
||||||
|
### Null
|
||||||
|
|
||||||
|
A `null` item draws a record column as null: `json` writes `null`, `sql` `NULL`, and
|
||||||
|
`csv` an empty field, with an empty string written `""` so PostgreSQL's `COPY … CSV`
|
||||||
|
reads both back. `Fake` renders a null as `""`. The other items' weights skew its
|
||||||
|
odds:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{ "format": "", "deleted_at": null, "middle": [null, { "format": "{n}", "n": ["Ann", "Eva"], "weight": 3 }] }
|
||||||
|
```
|
||||||
|
|
||||||
|
`deleted_at` is null every draw, `middle` a name three draws in four. Rejected at
|
||||||
|
load: `null` anywhere but a column, naming `""`, and a column whose items declare
|
||||||
|
different datatypes.
|
||||||
|
|
||||||
### Options and fields
|
### Options and fields
|
||||||
|
|
||||||
`format`, `weight`, `repeat` and `separator` are the only options; **any other
|
`format`, `weight`, `repeat`, `separator` and `datatype` are the only options; **any
|
||||||
key is a field** (see [Decisions](#decisions)). An object that does nothing a
|
other key is a field** (see [Decisions](#decisions)). An object that does nothing a
|
||||||
string can't — only a `format` — is rejected naming the string, as is a one-item
|
string can't — only a `format` — is rejected naming the string, as is a one-item
|
||||||
choice naming its item.
|
choice naming its item.
|
||||||
|
|
||||||
@@ -437,8 +481,8 @@ tokens add cost in proportion to the output.
|
|||||||
|
|
||||||
## Decisions
|
## Decisions
|
||||||
|
|
||||||
- **Options and fields share one namespace.** `format`, `weight`, `repeat` and
|
- **Options and fields share one namespace.** `format`, `weight`, `repeat`,
|
||||||
`separator` are reserved; every other key is a field. Nesting fields under a
|
`separator` and `datatype` are reserved; every other key is a field. Nesting fields under a
|
||||||
key, or prefixing options, would tax every template to guard against a
|
key, or prefixing options, would tax every template to guard against a
|
||||||
misspelt option.
|
misspelt option.
|
||||||
- **`{a|b}` stays beside nested choices.** `[[…], […]]` picks the same way, but
|
- **`{a|b}` stays beside nested choices.** `[[…], […]]` picks the same way, but
|
||||||
@@ -518,7 +562,8 @@ tokens add cost in proportion to the output.
|
|||||||
so `a/(b*c)` with `b` fixed at `0` and `c` varying loads and prints `Inf` every
|
so `a/(b*c)` with `b` fixed at `0` and `c` varying loads and prints `Inf` every
|
||||||
draw — catching it needs zero-absorbing algebra for a shape nobody writes.
|
draw — catching it needs zero-absorbing algebra for a shape nobody writes.
|
||||||
- **In data, a default written out and a constant spelled as a sample are load
|
- **In data, a default written out and a constant spelled as a sample are load
|
||||||
errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `int(5,5)`, `float(1,1,2)`,
|
errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `datatype: "string"`,
|
||||||
|
`int(5,5)`, `float(1,1,2)`,
|
||||||
`+5` and `05` each spell what a shorter form already spells, so each is rejected
|
`+5` and `05` each spell what a shorter form already spells, so each is rejected
|
||||||
naming that form. The CLI's numbers follow the shell instead: `--seed 007` and
|
naming that form. The CLI's numbers follow the shell instead: `--seed 007` and
|
||||||
`--repeat +3` are 7 and 3, as every command line reads them.
|
`--repeat +3` are 7 and 3, as every command line reads them.
|
||||||
@@ -557,6 +602,9 @@ tokens add cost in proportion to the output.
|
|||||||
row — would vary per draw. A fixed column set is what the CSV and `INSERT`
|
row — would vary per draw. A fixed column set is what the CSV and `INSERT`
|
||||||
contracts rest on, so the restriction holds even where a particular choice would
|
contracts rest on, so the restriction holds even where a particular choice would
|
||||||
happen to agree.
|
happen to agree.
|
||||||
|
- **Null is a `null` item, not a rate.** A null is one more outcome of a column's
|
||||||
|
draw, so a choice's weights skew it like any other; a null-rate option would be a
|
||||||
|
second way to state odds.
|
||||||
- **The performance gate asserts allocations, not wall-clock time.** `AllocsPerRun`
|
- **The performance gate asserts allocations, not wall-clock time.** `AllocsPerRun`
|
||||||
is deterministic across machines, so a ±10% ceiling does not flake under CI load,
|
is deterministic across machines, so a ±10% ceiling does not flake under CI load,
|
||||||
while time varies with the machine and its neighbours. A rendering slowdown
|
while time varies with the machine and its neighbours. A rendering slowdown
|
||||||
@@ -612,7 +660,9 @@ hold.go the hold: one draw per expansion for paths and operands, and its
|
|||||||
reference.go reference sigils, and binding references across the tree
|
reference.go reference sigils, and binding references across the tree
|
||||||
graph.go the render graph: edges, cycles, the repeat bound, tree walks
|
graph.go the render graph: edges, cycles, the repeat bound, tree walks
|
||||||
builtins.go the {name()} function registry and its implementations
|
builtins.go the {name()} function registry and its implementations
|
||||||
calc.go the {calc()} arithmetic evaluator: parser, eval, validation
|
calc.go the {calc()} arithmetic evaluator: parser, eval, validation, and the proof a typed column's calc is finite
|
||||||
|
datatype.go column datatypes: DataType, where datatype and null may sit, and the load check every typed render passes
|
||||||
|
renderlang.go what text a node can render, as relations over a scalar's grammar
|
||||||
data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge
|
data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge
|
||||||
cmd/fejkdata/ the fejkdata CLI
|
cmd/fejkdata/ the fejkdata CLI
|
||||||
data/ shipped data (JSON), embedded at build: locale folders + a misc folder
|
data/ shipped data (JSON), embedded at build: locale folders + a misc folder
|
||||||
|
|||||||
+101
-21
@@ -5,6 +5,7 @@ import (
|
|||||||
"errors"
|
"errors"
|
||||||
"fmt"
|
"fmt"
|
||||||
"math"
|
"math"
|
||||||
|
"slices"
|
||||||
"strconv"
|
"strconv"
|
||||||
"strings"
|
"strings"
|
||||||
"unicode"
|
"unicode"
|
||||||
@@ -24,36 +25,36 @@ const (
|
|||||||
// samples read only the rng. A time-based id (uuid v7, ulid) draws its timestamp
|
// samples read only the rng. A time-based id (uuid v7, ulid) draws its timestamp
|
||||||
// from the rng, not the wall clock, so seeded output stays reproducible.
|
// from the rng, not the wall clock, so seeded output stays reproducible.
|
||||||
var builtins = map[string]builtin{
|
var builtins = map[string]builtin{
|
||||||
"luhn": {arity: 0, prep: derive(func(e string) string { return string(rune('0' + luhnCheck(e))) })},
|
"luhn": {arity: 0, prep: derive(func(e string) string { return string(rune('0' + luhnCheck(e))) }), emits: always(textShape{{{decimalDigits, 1, 1}}})},
|
||||||
"mod11": {arity: 0, prep: derive(mod11Check)},
|
"mod11": {arity: 0, prep: derive(mod11Check), emits: always(textShape{{{decimalDigits + "X", 1, 1}}})},
|
||||||
"ean": {arity: 0, prep: derive(eanCheck)},
|
"ean": {arity: 0, prep: derive(eanCheck), emits: always(textShape{{{decimalDigits, 1, 1}}})},
|
||||||
"uuid": {arity: 0, prep: sample(uuidV7)},
|
"uuid": {arity: 0, prep: sample(uuidV7), emits: always(uuidShape)},
|
||||||
"ulid": {arity: 0, prep: sample(ulid)},
|
"ulid": {arity: 0, prep: sample(ulid), emits: always(textShape{{{crockford[:8], 1, 1}, {crockford, 25, 25}}})},
|
||||||
"nanoid": {arity: 1, check: posIntArg, prep: chars(nanoidAlphabet)},
|
"nanoid": sampleOf(nanoidAlphabet),
|
||||||
"hex": {arity: 1, check: posIntArg, prep: chars(hexDigits)},
|
"hex": sampleOf(hexDigits),
|
||||||
"digits": {arity: 1, check: posIntArg, prep: chars("0123456789")},
|
"digits": sampleOf(decimalDigits),
|
||||||
"upper": {arity: 1, check: posIntArg, prep: chars("ABCDEFGHIJKLMNOPQRSTUVWXYZ")},
|
"upper": sampleOf("ABCDEFGHIJKLMNOPQRSTUVWXYZ"),
|
||||||
"lower": {arity: 1, check: posIntArg, prep: chars("abcdefghijklmnopqrstuvwxyz")},
|
"lower": sampleOf("abcdefghijklmnopqrstuvwxyz"),
|
||||||
"base64": {arity: 1, check: posIntArg, prep: func(a []string) callFn {
|
"base64": {arity: 1, check: posIntArg, prep: func(a []string) callFn {
|
||||||
n := atoi(a[0])
|
n := atoi(a[0])
|
||||||
return func(s *session, _ string, _ []string) string {
|
return func(s *session, _ string, _ []string) string {
|
||||||
return base64.StdEncoding.EncodeToString(randBytes(s, n))
|
return base64.StdEncoding.EncodeToString(randBytes(s, n))
|
||||||
}
|
}
|
||||||
}},
|
}, emits: base64Shape},
|
||||||
"int": {arity: 2, check: intRangeArgs, prep: func(a []string) callFn {
|
"int": {arity: 2, check: intRangeArgs, prep: func(a []string) callFn {
|
||||||
lo, span := atoi(a[0]), atoi(a[1])-atoi(a[0])+1
|
lo, span := atoi(a[0]), atoi(a[1])-atoi(a[0])+1
|
||||||
return func(s *session, _ string, _ []string) string { return strconv.Itoa(lo + s.IntN(span)) }
|
return func(s *session, _ string, _ []string) string { return strconv.Itoa(lo + s.IntN(span)) }
|
||||||
}},
|
}, emits: intShape},
|
||||||
"float": {arity: 3, check: floatArgs, prep: func(a []string) callFn {
|
"float": {arity: 3, check: floatArgs, prep: func(a []string) callFn {
|
||||||
lo, hi, dp := atof(a[0]), atof(a[1]), atoi(a[2])
|
lo, hi, dp := atof(a[0]), atof(a[1]), atoi(a[2])
|
||||||
return func(s *session, _ string, _ []string) string {
|
return func(s *session, _ string, _ []string) string {
|
||||||
return strconv.FormatFloat(lo+s.Float64()*(hi-lo), 'f', dp, 64)
|
return strconv.FormatFloat(lo+s.Float64()*(hi-lo), 'f', dp, 64)
|
||||||
}
|
}
|
||||||
}},
|
}, emits: func(a []string) textShape { return printedFloat(atof(a[0]), atof(a[1]), atoi(a[2]), false) }},
|
||||||
"iban": {arity: 1, check: ibanArg, prep: func(a []string) callFn {
|
"iban": {arity: 1, check: ibanArg, prep: func(a []string) callFn {
|
||||||
cc := a[0]
|
cc := a[0]
|
||||||
return func(s *session, _ string, _ []string) string { return iban(s, cc) }
|
return func(s *session, _ string, _ []string) string { return iban(s, cc) }
|
||||||
}},
|
}, emits: ibanShape},
|
||||||
"calc": {arity: -1, check: checkCalc, prep: calcPrep, operands: calcOperands},
|
"calc": {arity: -1, check: checkCalc, prep: calcPrep, operands: calcOperands},
|
||||||
"lowercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToLower), operands: transformOperand},
|
"lowercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToLower), operands: transformOperand},
|
||||||
"uppercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToUpper), operands: transformOperand},
|
"uppercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToUpper), operands: transformOperand},
|
||||||
@@ -69,7 +70,7 @@ var builtins = map[string]builtin{
|
|||||||
return func(s *session, _ string, _ []string) string {
|
return func(s *session, _ string, _ []string) string {
|
||||||
return strconv.FormatUint(s.next(key), 10)
|
return strconv.FormatUint(s.next(key), 10)
|
||||||
}
|
}
|
||||||
}},
|
}, emits: always(textShape{{{nonZeroDigits, 1, 1}, {decimalDigits, 0, 19}}})},
|
||||||
}
|
}
|
||||||
|
|
||||||
// derive and sample are the two argument-free builtin shapes: a derivation reads
|
// derive and sample are the two argument-free builtin shapes: a derivation reads
|
||||||
@@ -94,6 +95,82 @@ func chars(alphabet string) func([]string) callFn {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// sampleOf is the builtin that draws n characters from an alphabet.
|
||||||
|
func sampleOf(alphabet string) builtin {
|
||||||
|
return builtin{arity: 1, check: posIntArg, prep: chars(alphabet), emits: func(a []string) textShape {
|
||||||
|
n := atoi(a[0])
|
||||||
|
return textShape{{{alphabet, n, n}}}
|
||||||
|
}}
|
||||||
|
}
|
||||||
|
|
||||||
|
// always is the emits of a builtin whose args do not change what it can print.
|
||||||
|
func always(s textShape) func([]string) textShape {
|
||||||
|
return func([]string) textShape { return s }
|
||||||
|
}
|
||||||
|
|
||||||
|
var uuidShape = textShape{{{hexDigits, 8, 8}, {"-", 1, 1}, {hexDigits, 4, 4}, {"-", 1, 1}, {"7", 1, 1}, {hexDigits, 3, 3}, {"-", 1, 1}, {"89ab", 1, 1}, {hexDigits, 3, 3}, {"-", 1, 1}, {hexDigits, 12, 12}}}
|
||||||
|
|
||||||
|
const base64Alphabet = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/"
|
||||||
|
|
||||||
|
func base64Shape(a []string) textShape {
|
||||||
|
n := atoi(a[0])
|
||||||
|
pad := (3 - n%3) % 3
|
||||||
|
size := 4*((n+2)/3) - pad
|
||||||
|
return textShape{{{base64Alphabet, size, size}, {"=", pad, pad}}}
|
||||||
|
}
|
||||||
|
|
||||||
|
// intShape is what int prints: a sign only below zero, and no leading zero.
|
||||||
|
func intShape(a []string) textShape {
|
||||||
|
lo, hi := atoi(a[0]), atoi(a[1])
|
||||||
|
var s textShape
|
||||||
|
if lo <= 0 && hi >= 0 {
|
||||||
|
s = append(s, []charRun{{"0", 1, 1}})
|
||||||
|
}
|
||||||
|
if hi > 0 {
|
||||||
|
s = append(s, []charRun{{nonZeroDigits, 1, 1}, {decimalDigits, 0, len(a[1]) - 1}})
|
||||||
|
}
|
||||||
|
if lo < 0 {
|
||||||
|
s = append(s, []charRun{{"-", 1, 1}, {nonZeroDigits, 1, 1}, {decimalDigits, 0, len(a[0]) - 2}})
|
||||||
|
}
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
|
||||||
|
func ibanShape(a []string) textShape {
|
||||||
|
cc, digits := a[0], ibanLen[a[0]]-2
|
||||||
|
return textShape{{{cc[:1], 1, 1}, {cc[1:], 1, 1}, {decimalDigits, digits, digits}}}
|
||||||
|
}
|
||||||
|
|
||||||
|
// shortestFraction bounds the fraction FormatFloat's shortest form prints: at most 17
|
||||||
|
// significant digits after up to 323 zeros.
|
||||||
|
const shortestFraction = 340
|
||||||
|
|
||||||
|
// printedFloat is what strconv.FormatFloat(v, 'f', dp, 64) prints for a v in [lo, hi]
|
||||||
|
// that is whole when integral.
|
||||||
|
func printedFloat(lo, hi float64, dp int, integral bool) textShape {
|
||||||
|
digits := len(strconv.FormatFloat(math.Floor(math.Max(math.Abs(lo), math.Abs(hi))), 'f', 0, 64)) + 1 // one more for a rounding carry
|
||||||
|
wholes := [][]charRun{{{"0", 1, 1}}, {{nonZeroDigits, 1, 1}, {decimalDigits, 0, digits - 1}}}
|
||||||
|
fractions := [][]charRun{nil}
|
||||||
|
switch {
|
||||||
|
case dp > 0:
|
||||||
|
fractions = [][]charRun{{{".", 1, 1}, {decimalDigits, dp, dp}}}
|
||||||
|
case dp < 0 && !integral:
|
||||||
|
fractions = append(fractions, []charRun{{".", 1, 1}, {decimalDigits, 1, shortestFraction}})
|
||||||
|
}
|
||||||
|
signs := [][]charRun{nil}
|
||||||
|
if lo < 0 || math.Signbit(lo) {
|
||||||
|
signs = append(signs, []charRun{{"-", 1, 1}})
|
||||||
|
}
|
||||||
|
var s textShape
|
||||||
|
for _, sign := range signs {
|
||||||
|
for _, whole := range wholes {
|
||||||
|
for _, fraction := range fractions {
|
||||||
|
s = append(s, slices.Concat(sign, whole, fraction))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
|
||||||
const hexDigits = "0123456789abcdef"
|
const hexDigits = "0123456789abcdef"
|
||||||
|
|
||||||
// transforms are the builtins that rewrite one operand's value; they nest, so
|
// transforms are the builtins that rewrite one operand's value; they nest, so
|
||||||
@@ -106,20 +183,19 @@ var transforms = map[string]func(string) string{
|
|||||||
|
|
||||||
// unwrapTransform peels nested transform calls off an operand arg, returning the
|
// unwrapTransform peels nested transform calls off an operand arg, returning the
|
||||||
// field it finally names and the transforms to apply, innermost last.
|
// field it finally names and the transforms to apply, innermost last.
|
||||||
func unwrapTransform(arg string) (leaf string, chain []func(string) string, err error) {
|
func unwrapTransform(arg string) (leaf string, chain []string, err error) {
|
||||||
for {
|
for {
|
||||||
name, args, isCall := funcCall(arg)
|
name, args, isCall := funcCall(arg)
|
||||||
if !isCall {
|
if !isCall {
|
||||||
return arg, chain, nil
|
return arg, chain, nil
|
||||||
}
|
}
|
||||||
fn, isTransform := transforms[name]
|
if _, isTransform := transforms[name]; !isTransform {
|
||||||
if !isTransform {
|
|
||||||
return "", nil, fmt.Errorf("%s(%s) is not a transform, so it cannot be an operand", name, strings.Join(args, ","))
|
return "", nil, fmt.Errorf("%s(%s) is not a transform, so it cannot be an operand", name, strings.Join(args, ","))
|
||||||
}
|
}
|
||||||
if len(args) != 1 {
|
if len(args) != 1 {
|
||||||
return "", nil, fmt.Errorf("%s takes 1 arg, got %d", name, len(args))
|
return "", nil, fmt.Errorf("%s takes 1 arg, got %d", name, len(args))
|
||||||
}
|
}
|
||||||
chain = append(chain, fn)
|
chain = append(chain, name)
|
||||||
arg = args[0]
|
arg = args[0]
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -150,10 +226,14 @@ func transformPrep(outer func(string) string) func([]string) callFn {
|
|||||||
if err != nil {
|
if err != nil {
|
||||||
panic(fmt.Sprintf("fejkdata: transform arg %q reached prep unvalidated: %v", a[0], err))
|
panic(fmt.Sprintf("fejkdata: transform arg %q reached prep unvalidated: %v", a[0], err))
|
||||||
}
|
}
|
||||||
|
fns := make([]func(string) string, len(chain))
|
||||||
|
for i, name := range chain {
|
||||||
|
fns[i] = transforms[name]
|
||||||
|
}
|
||||||
return func(_ *session, _ string, operands []string) string {
|
return func(_ *session, _ string, operands []string) string {
|
||||||
v := operands[0]
|
v := operands[0]
|
||||||
for i := len(chain) - 1; i >= 0; i-- {
|
for i := len(fns) - 1; i >= 0; i-- {
|
||||||
v = chain[i](v)
|
v = fns[i](v)
|
||||||
}
|
}
|
||||||
return outer(v)
|
return outer(v)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -6,6 +6,7 @@ import (
|
|||||||
"strconv"
|
"strconv"
|
||||||
"strings"
|
"strings"
|
||||||
"unicode"
|
"unicode"
|
||||||
|
"unicode/utf8"
|
||||||
)
|
)
|
||||||
|
|
||||||
// calcNode is a parsed expression node. It evaluates over the operand values expand
|
// calcNode is a parsed expression node. It evaluates over the operand values expand
|
||||||
@@ -154,10 +155,12 @@ func calcText(n calcNode) string {
|
|||||||
return "?"
|
return "?"
|
||||||
}
|
}
|
||||||
|
|
||||||
// neverNumeric reports a node no render of which is a number: fixed text that does
|
// neverNumeric reports a node no render of which is a number: a null, fixed text that
|
||||||
// not parse, or a choice of only such items. text is one such render.
|
// does not parse, or a choice of only such items. text is one such render.
|
||||||
func neverNumeric(n node) (text string, never bool) {
|
func neverNumeric(n node) (text string, never bool) {
|
||||||
switch n := n.(type) {
|
switch n := n.(type) {
|
||||||
|
case *null:
|
||||||
|
return "", true
|
||||||
case *template:
|
case *template:
|
||||||
if !n.fixed || n.repeat > 1 {
|
if !n.fixed || n.repeat > 1 {
|
||||||
return "", false
|
return "", false
|
||||||
@@ -191,10 +194,7 @@ func calcPrep(args []string) callFn {
|
|||||||
at[name] = i
|
at[name] = i
|
||||||
}
|
}
|
||||||
placed := indexVars(expr, at)
|
placed := indexVars(expr, at)
|
||||||
dp := -1
|
dp := calcDecimals(args)
|
||||||
if len(args) == 2 {
|
|
||||||
dp = atoi(args[1])
|
|
||||||
}
|
|
||||||
return func(_ *session, _ string, operands []string) string {
|
return func(_ *session, _ string, operands []string) string {
|
||||||
return strconv.FormatFloat(placed.eval(operands), 'f', dp, 64)
|
return strconv.FormatFloat(placed.eval(operands), 'f', dp, 64)
|
||||||
}
|
}
|
||||||
@@ -384,3 +384,256 @@ func contains(bs []byte, b byte) bool {
|
|||||||
}
|
}
|
||||||
return false
|
return false
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// calcDecimals is a calc's decimals count, or -1 for the shortest form.
|
||||||
|
func calcDecimals(args []string) int {
|
||||||
|
if len(args) == 2 {
|
||||||
|
return atoi(args[1])
|
||||||
|
}
|
||||||
|
return -1
|
||||||
|
}
|
||||||
|
|
||||||
|
// calcLimit is the largest magnitude a proof accepts as finite, far enough below
|
||||||
|
// math.MaxFloat64 that rounding in the bounds cannot hide an overflow.
|
||||||
|
const calcLimit = 1e300
|
||||||
|
|
||||||
|
// maxOperandLen is the longest operand text a proof bounds by its length, so that
|
||||||
|
// bound, 10^maxOperandLen, stays within calcLimit.
|
||||||
|
const maxOperandLen = 300
|
||||||
|
|
||||||
|
// calcBound is what a proof knows of every value a calc can take: it lies in [lo, hi],
|
||||||
|
// is at least nonZero from zero unless nonZero is 0, and is whole when integral.
|
||||||
|
type calcBound struct {
|
||||||
|
lo, hi, nonZero float64
|
||||||
|
integral bool
|
||||||
|
}
|
||||||
|
|
||||||
|
func magnitude(b calcBound) float64 { return math.Max(math.Abs(b.lo), math.Abs(b.hi)) }
|
||||||
|
|
||||||
|
// doubt is why a proof could not show a calc finite, and the render that shows it.
|
||||||
|
type doubt struct{ render, why string }
|
||||||
|
|
||||||
|
type bounded struct {
|
||||||
|
b calcBound
|
||||||
|
d *doubt
|
||||||
|
}
|
||||||
|
|
||||||
|
// calcProof bounds a typed column's calcs from their operands' renders, to show each
|
||||||
|
// prints a number rather than NaN or Inf.
|
||||||
|
type calcProof struct {
|
||||||
|
decimal *textLanguage
|
||||||
|
operands map[node]bounded
|
||||||
|
lengths map[node]int
|
||||||
|
}
|
||||||
|
|
||||||
|
func newCalcProof() *calcProof {
|
||||||
|
p := &calcProof{operands: map[node]bounded{}, lengths: map[node]int{}}
|
||||||
|
p.decimal = newTextLanguage(decimalGrammar, p)
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
// call bounds one calc token of t.
|
||||||
|
func (p *calcProof) call(t *template, args []string) (calcBound, *doubt) {
|
||||||
|
expr, err := parseCalc(args[0])
|
||||||
|
if err != nil {
|
||||||
|
panic(fmt.Sprintf("fejkdata: calc(%q) reached a proof unparsed: %v", args[0], err))
|
||||||
|
}
|
||||||
|
b, d := p.expr(expr, t.fields)
|
||||||
|
if d != nil {
|
||||||
|
return b, &doubt{d.render, fmt.Sprintf("{calc(%s)}: %s", strings.Join(args, ", "), d.why)}
|
||||||
|
}
|
||||||
|
return b, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (p *calcProof) expr(n calcNode, fields map[string]node) (calcBound, *doubt) {
|
||||||
|
switch n := n.(type) {
|
||||||
|
case calcNum:
|
||||||
|
v := float64(n)
|
||||||
|
return calcBound{v, v, v, v == math.Trunc(v)}, nil
|
||||||
|
case calcVar:
|
||||||
|
return p.operand(string(n), fields[string(n)])
|
||||||
|
case calcNeg:
|
||||||
|
b, d := p.expr(n.x, fields)
|
||||||
|
return calcBound{-b.hi, -b.lo, b.nonZero, b.integral}, d
|
||||||
|
case calcBin:
|
||||||
|
l, d := p.expr(n.l, fields)
|
||||||
|
if d != nil {
|
||||||
|
return l, d
|
||||||
|
}
|
||||||
|
r, d := p.expr(n.r, fields)
|
||||||
|
if d != nil {
|
||||||
|
return r, d
|
||||||
|
}
|
||||||
|
return combine(n, l, r)
|
||||||
|
}
|
||||||
|
panic(fmt.Sprintf("fejkdata: calc node %T has no bound", n))
|
||||||
|
}
|
||||||
|
|
||||||
|
// combine bounds one operation from the bounds of its sides.
|
||||||
|
func combine(n calcBin, l, r calcBound) (calcBound, *doubt) {
|
||||||
|
b := calcBound{integral: l.integral && r.integral}
|
||||||
|
switch n.op {
|
||||||
|
case '+':
|
||||||
|
b.lo, b.hi = l.lo+r.lo, l.hi+r.hi
|
||||||
|
case '-':
|
||||||
|
b.lo, b.hi = l.lo-r.hi, l.hi-r.lo
|
||||||
|
case '*':
|
||||||
|
b.lo = min(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi)
|
||||||
|
b.hi = max(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi)
|
||||||
|
b.nonZero = l.nonZero * r.nonZero
|
||||||
|
default:
|
||||||
|
if r.nonZero == 0 {
|
||||||
|
return b, &doubt{"+Inf", fmt.Sprintf("divides by %s, which can be zero", calcText(n.r))}
|
||||||
|
}
|
||||||
|
m := magnitude(l) / r.nonZero
|
||||||
|
b = calcBound{lo: -m, hi: m, nonZero: l.nonZero / magnitude(r)}
|
||||||
|
}
|
||||||
|
if b.lo > 0 || b.hi < 0 {
|
||||||
|
b.nonZero = math.Max(b.nonZero, math.Min(math.Abs(b.lo), math.Abs(b.hi)))
|
||||||
|
}
|
||||||
|
if !(magnitude(b) <= calcLimit) {
|
||||||
|
return b, &doubt{"+Inf", calcText(n) + " can overflow"}
|
||||||
|
}
|
||||||
|
return b, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// operand bounds a calc operand, once per node.
|
||||||
|
func (p *calcProof) operand(name string, n node) (calcBound, *doubt) {
|
||||||
|
if seen, done := p.operands[n]; done {
|
||||||
|
return seen.b, seen.d
|
||||||
|
}
|
||||||
|
b, d := p.measure(name, n)
|
||||||
|
p.operands[n] = bounded{b, d}
|
||||||
|
return b, d
|
||||||
|
}
|
||||||
|
|
||||||
|
// measure bounds an operand through the calc it renders when that is all it renders,
|
||||||
|
// and otherwise from its text: a plain decimal of at most maxOperandLen bytes.
|
||||||
|
func (p *calcProof) measure(name string, n node) (calcBound, *doubt) {
|
||||||
|
if t, ok := n.(*template); ok {
|
||||||
|
if args, isCalc := soleCalc(t); isCalc {
|
||||||
|
b, d := p.call(t, args)
|
||||||
|
return rounded(b, calcDecimals(args)), d
|
||||||
|
}
|
||||||
|
}
|
||||||
|
text := p.decimal.node(n, nil)
|
||||||
|
if w, escapes := text.escape(decimalAccept); escapes {
|
||||||
|
why := fmt.Sprintf("operand %q can render %s, which is not a plain decimal", name, w)
|
||||||
|
if w.why != "" {
|
||||||
|
why += ": " + w.why
|
||||||
|
}
|
||||||
|
return calcBound{}, &doubt{"NaN", why}
|
||||||
|
}
|
||||||
|
size := p.length(n)
|
||||||
|
if size > maxOperandLen {
|
||||||
|
return calcBound{}, &doubt{"NaN", fmt.Sprintf("operand %q can render more than %d bytes, too many to bound", name, maxOperandLen)}
|
||||||
|
}
|
||||||
|
ends, m := text.to[1], math.Pow(10, float64(size))
|
||||||
|
b := calcBound{hi: m, nonZero: 1 / m, integral: ends&decimalFractional == 0}
|
||||||
|
if ends&decimalNegative != 0 {
|
||||||
|
b.lo = -m
|
||||||
|
}
|
||||||
|
if ends&decimalZero != 0 {
|
||||||
|
b.nonZero = 0
|
||||||
|
}
|
||||||
|
return b, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// soleCalc reports a template that renders one calc and nothing else, with its args.
|
||||||
|
func soleCalc(t *template) ([]string, bool) {
|
||||||
|
if t.repeat != 1 || len(t.ops) != 1 || t.ops[0].kind != 'b' {
|
||||||
|
return nil, false
|
||||||
|
}
|
||||||
|
name, args, _ := funcCall(t.format[1 : len(t.format)-1])
|
||||||
|
return args, name == "calc"
|
||||||
|
}
|
||||||
|
|
||||||
|
// rounded is b once printed to dp decimals, which moves a value by up to half a unit.
|
||||||
|
func rounded(b calcBound, dp int) calcBound {
|
||||||
|
if dp < 0 {
|
||||||
|
return b
|
||||||
|
}
|
||||||
|
half := math.Pow(10, -float64(dp)) / 2
|
||||||
|
return calcBound{b.lo - half, b.hi + half, math.Max(0, b.nonZero-half), b.integral || dp == 0}
|
||||||
|
}
|
||||||
|
|
||||||
|
// length is the most bytes a render of n can take, anything past maxOperandLen
|
||||||
|
// reported as maxOperandLen+1.
|
||||||
|
func (p *calcProof) length(n node) int {
|
||||||
|
if size, done := p.lengths[n]; done {
|
||||||
|
return size
|
||||||
|
}
|
||||||
|
size := 0
|
||||||
|
switch n := n.(type) {
|
||||||
|
case *choice:
|
||||||
|
for _, it := range n.items {
|
||||||
|
size = max(size, p.length(it))
|
||||||
|
}
|
||||||
|
case *template:
|
||||||
|
size = p.formatLength(n)*n.repeat + len(n.separator)*(n.repeat-1)
|
||||||
|
}
|
||||||
|
size = min(size, maxOperandLen+1)
|
||||||
|
p.lengths[n] = size
|
||||||
|
return size
|
||||||
|
}
|
||||||
|
|
||||||
|
func (p *calcProof) formatLength(t *template) int {
|
||||||
|
size := 0
|
||||||
|
_ = eachToken(t.format, func(tok ftoken) error {
|
||||||
|
if tok.kind == 'l' {
|
||||||
|
size += utf8.RuneLen(tok.r)
|
||||||
|
} else {
|
||||||
|
size += p.tokenLength(t, tok.body)
|
||||||
|
}
|
||||||
|
size = min(size, maxOperandLen+1)
|
||||||
|
return nil
|
||||||
|
})
|
||||||
|
return size
|
||||||
|
}
|
||||||
|
|
||||||
|
// tokenLength is the most bytes one token can print. A transform never lengthens a
|
||||||
|
// render that reads as a decimal: it maps each non-ASCII rune, two bytes or more, to at
|
||||||
|
// most two ASCII letters.
|
||||||
|
func (p *calcProof) tokenLength(t *template, body string) int {
|
||||||
|
name, args, isFunc := funcCall(body)
|
||||||
|
var arms []arm
|
||||||
|
switch _, isTransform := transforms[name]; {
|
||||||
|
case !isFunc:
|
||||||
|
arms = splitArms(body, t.refs)
|
||||||
|
case isTransform:
|
||||||
|
leaf, _, _ := unwrapTransform(args[0])
|
||||||
|
arms = []arm{splitArm(leaf, t.refs)}
|
||||||
|
case name == "calc":
|
||||||
|
b, d := p.call(t, args)
|
||||||
|
if d != nil {
|
||||||
|
return len(d.render)
|
||||||
|
}
|
||||||
|
return shapeLength(printedFloat(b.lo, b.hi, calcDecimals(args), b.integral))
|
||||||
|
default:
|
||||||
|
return shapeLength(builtins[name].emits(args))
|
||||||
|
}
|
||||||
|
size := 0
|
||||||
|
for _, a := range arms {
|
||||||
|
for _, leaf := range pathLeaves(t.fields[a.key], a.tail) {
|
||||||
|
size = max(size, p.length(leaf))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return size
|
||||||
|
}
|
||||||
|
|
||||||
|
// shapeLength is the most bytes a shape can emit, anything past maxOperandLen reported
|
||||||
|
// as maxOperandLen+1.
|
||||||
|
func shapeLength(s textShape) int {
|
||||||
|
longest := 0
|
||||||
|
for _, alt := range s {
|
||||||
|
size := 0
|
||||||
|
for _, run := range alt {
|
||||||
|
if run.max < 0 {
|
||||||
|
return maxOperandLen + 1
|
||||||
|
}
|
||||||
|
size += run.max
|
||||||
|
}
|
||||||
|
longest = max(longest, size)
|
||||||
|
}
|
||||||
|
return min(longest, maxOperandLen+1)
|
||||||
|
}
|
||||||
|
|||||||
+140
@@ -0,0 +1,140 @@
|
|||||||
|
package fejkdata
|
||||||
|
|
||||||
|
import (
|
||||||
|
"errors"
|
||||||
|
"fmt"
|
||||||
|
)
|
||||||
|
|
||||||
|
// DataType is what a record column holds, which decides how a record writes its value.
|
||||||
|
type DataType int
|
||||||
|
|
||||||
|
// The datatypes a column declares with "datatype"; a column without one is a string.
|
||||||
|
const (
|
||||||
|
DataTypeString DataType = iota
|
||||||
|
DataTypeInteger
|
||||||
|
DataTypeNumber
|
||||||
|
DataTypeBoolean
|
||||||
|
)
|
||||||
|
|
||||||
|
var dataTypeNames = [...]string{"string", "integer", "number", "boolean"}
|
||||||
|
|
||||||
|
// String is the datatype as data spells it.
|
||||||
|
func (d DataType) String() string {
|
||||||
|
if d < 0 || int(d) >= len(dataTypeNames) {
|
||||||
|
return fmt.Sprintf("DataType(%d)", int(d))
|
||||||
|
}
|
||||||
|
return dataTypeNames[d]
|
||||||
|
}
|
||||||
|
|
||||||
|
// position is where a JSON value sits, which decides whether it may carry a datatype or
|
||||||
|
// be null.
|
||||||
|
type position int
|
||||||
|
|
||||||
|
const (
|
||||||
|
inFormat position = iota // rendered by a format, so neither
|
||||||
|
atTop // a category or an inline template, whose fields are the columns
|
||||||
|
inColumn // a column, or a choice item standing in for one
|
||||||
|
)
|
||||||
|
|
||||||
|
// datatypeOf reads a template's "datatype" (default DataTypeString).
|
||||||
|
func datatypeOf(m map[string]any, pos position) (DataType, error) {
|
||||||
|
v, ok := m["datatype"]
|
||||||
|
if !ok {
|
||||||
|
return DataTypeString, nil
|
||||||
|
}
|
||||||
|
name, ok := v.(string)
|
||||||
|
if !ok {
|
||||||
|
return 0, fmt.Errorf("datatype must be a string, got %T", v)
|
||||||
|
}
|
||||||
|
if name == DataTypeString.String() {
|
||||||
|
return 0, fmt.Errorf("datatype %q is the default, so it has no effect; drop it", name)
|
||||||
|
}
|
||||||
|
for d := DataTypeInteger; d <= DataTypeBoolean; d++ {
|
||||||
|
if name != d.String() {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if pos != inColumn {
|
||||||
|
return 0, errors.New("datatype only types a record column — a field of the top-level template — so it has no effect here")
|
||||||
|
}
|
||||||
|
return d, nil
|
||||||
|
}
|
||||||
|
return 0, fmt.Errorf(`datatype takes "integer", "number" or "boolean", got %q`, name)
|
||||||
|
}
|
||||||
|
|
||||||
|
// columnDatatype is the datatype a column's items declare. They must agree, since a
|
||||||
|
// column holds one; a column only ever null is a string.
|
||||||
|
func columnDatatype(n node) (DataType, error) {
|
||||||
|
var declared []DataType
|
||||||
|
var collect func(node)
|
||||||
|
collect = func(n node) {
|
||||||
|
switch n := n.(type) {
|
||||||
|
case *choice:
|
||||||
|
for _, it := range n.items {
|
||||||
|
collect(it)
|
||||||
|
}
|
||||||
|
case *template:
|
||||||
|
declared = append(declared, n.datatype)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
collect(n)
|
||||||
|
if len(declared) == 0 {
|
||||||
|
return DataTypeString, nil
|
||||||
|
}
|
||||||
|
for _, d := range declared {
|
||||||
|
if d != declared[0] {
|
||||||
|
return declared[0], fmt.Errorf("its items declare %s and %s; a column holds one datatype, so give every item the same", declared[0], d)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return declared[0], nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// datatypeSpec is what a datatype's text must satisfy: a grammar, the states a render
|
||||||
|
// may end in, and how an error names the datatype.
|
||||||
|
type datatypeSpec struct {
|
||||||
|
grammar *grammar
|
||||||
|
accept uint32
|
||||||
|
noun string
|
||||||
|
}
|
||||||
|
|
||||||
|
var datatypeSpecs = map[DataType]datatypeSpec{
|
||||||
|
DataTypeInteger: {numberGrammar, integerAccept, "an integer"},
|
||||||
|
DataTypeNumber: {numberGrammar, numberAccept, "a number"},
|
||||||
|
DataTypeBoolean: {booleanGrammar, booleanAccept, "a boolean"},
|
||||||
|
}
|
||||||
|
|
||||||
|
// datatypeCheck proves every render of a typed column is text its datatype takes. One
|
||||||
|
// check covers a scope, so a node several columns reach is read once per grammar.
|
||||||
|
type datatypeCheck struct {
|
||||||
|
languages map[*grammar]*textLanguage
|
||||||
|
proof *calcProof
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *datatypeCheck) check(path string, n node) error {
|
||||||
|
t, ok := n.(*template)
|
||||||
|
if !ok || t.datatype == DataTypeString {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
spec := datatypeSpecs[t.datatype]
|
||||||
|
w, escapes := c.language(spec.grammar).node(t, nil).escape(spec.accept)
|
||||||
|
if !escapes {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
msg := fmt.Sprintf("%s: datatype %s, but it can render %s, which is not %s", path, t.datatype, w, spec.noun)
|
||||||
|
if w.why != "" {
|
||||||
|
msg += ": " + w.why
|
||||||
|
}
|
||||||
|
return errors.New(msg)
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *datatypeCheck) language(g *grammar) *textLanguage {
|
||||||
|
if c.proof == nil {
|
||||||
|
c.proof = newCalcProof()
|
||||||
|
c.languages = map[*grammar]*textLanguage{}
|
||||||
|
}
|
||||||
|
l, made := c.languages[g]
|
||||||
|
if !made {
|
||||||
|
l = newTextLanguage(g, c.proof)
|
||||||
|
c.languages[g] = l
|
||||||
|
}
|
||||||
|
return l
|
||||||
|
}
|
||||||
@@ -162,6 +162,8 @@ func paths(n node) []string {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
return out
|
return out
|
||||||
|
case *null:
|
||||||
|
return []string{""}
|
||||||
case *choice:
|
case *choice:
|
||||||
out := []string{""}
|
out := []string{""}
|
||||||
for p := range n.shared {
|
for p := range n.shared {
|
||||||
|
|||||||
@@ -186,7 +186,10 @@ func checkScope(s nodeScope) error {
|
|||||||
if err := s(func(path string, n node) error { return repeatCheck(path, n, mem) }); err != nil {
|
if err := s(func(path string, n node) error { return repeatCheck(path, n, mem) }); err != nil {
|
||||||
return err
|
return err
|
||||||
}
|
}
|
||||||
return s(heldCheck)
|
if err := s(heldCheck); err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
return s((&datatypeCheck{}).check)
|
||||||
}
|
}
|
||||||
|
|
||||||
type reachMemo map[node]int
|
type reachMemo map[node]int
|
||||||
|
|||||||
@@ -32,6 +32,12 @@ type choice struct {
|
|||||||
|
|
||||||
func (*choice) isNode() {}
|
func (*choice) isNode() {}
|
||||||
|
|
||||||
|
// null is a record column's missing value, rendered as "". It is not zero-sized, so two
|
||||||
|
// nulls are two map keys.
|
||||||
|
type null struct{ _ byte }
|
||||||
|
|
||||||
|
func (*null) isNode() {}
|
||||||
|
|
||||||
// template renders a format string, substituting {tokens} from fields. A bare
|
// template renders a format string, substituting {tokens} from fields. A bare
|
||||||
// JSON string is a template with no fields. repeat (default 1) renders that format
|
// JSON string is a template with no fields. repeat (default 1) renders that format
|
||||||
// that many times and joins the results with separator (default ""), each render
|
// that many times and joins the results with separator (default ""), each render
|
||||||
@@ -41,6 +47,7 @@ type template struct {
|
|||||||
fields map[string]node
|
fields map[string]node
|
||||||
repeat int
|
repeat int
|
||||||
separator string
|
separator string
|
||||||
|
datatype DataType
|
||||||
ops []op // format compiled once (see compileOps); what expand walks
|
ops []op // format compiled once (see compileOps); what expand walks
|
||||||
grow int // minimum output size, to size the render buffer
|
grow int // minimum output size, to size the render buffer
|
||||||
fixed bool // no op varies, so every render is lit
|
fixed bool // no op varies, so every render is lit
|
||||||
@@ -66,26 +73,37 @@ func (t *template) field(seg string) (node, bool) {
|
|||||||
return n, ok
|
return n, ok
|
||||||
}
|
}
|
||||||
|
|
||||||
// compile converts parsed JSON into a node tree, validating structure up front.
|
// compile converts parsed JSON — a category or an inline template — into a node tree,
|
||||||
// Only a choice's items carry a weight, so one here would be inert whatever its type.
|
// validating structure up front.
|
||||||
func compile(v any) (node, error) {
|
func compile(v any) (node, error) {
|
||||||
|
return compileAt(v, atTop)
|
||||||
|
}
|
||||||
|
|
||||||
|
// compileAt compiles a node that is no choice's item. Only a choice's items carry a
|
||||||
|
// weight, so one here would be inert whatever its type.
|
||||||
|
func compileAt(v any, pos position) (node, error) {
|
||||||
if m, ok := v.(map[string]any); ok {
|
if m, ok := v.(map[string]any); ok {
|
||||||
if _, weighted := m["weight"]; weighted {
|
if _, weighted := m["weight"]; weighted {
|
||||||
return nil, fmt.Errorf("weight only skews a choice's items, so it has no effect here; it is an option and can never be a field")
|
return nil, fmt.Errorf("weight only skews a choice's items, so it has no effect here; it is an option and can never be a field")
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
return compileItem(v)
|
return compileItem(v, pos)
|
||||||
}
|
}
|
||||||
|
|
||||||
// compileItem compiles one node, allowing the weight a choice item may carry.
|
// compileItem compiles one node, allowing the weight a choice item may carry.
|
||||||
func compileItem(v any) (node, error) {
|
func compileItem(v any, pos position) (node, error) {
|
||||||
switch v := v.(type) {
|
switch v := v.(type) {
|
||||||
case string:
|
case string:
|
||||||
return compileString(v)
|
return compileString(v)
|
||||||
case []any:
|
case []any:
|
||||||
return compileChoice(v)
|
return compileChoice(v, pos)
|
||||||
case map[string]any:
|
case map[string]any:
|
||||||
return compileTemplate(v)
|
return compileTemplate(v, pos)
|
||||||
|
case nil:
|
||||||
|
if pos != inColumn {
|
||||||
|
return nil, fmt.Errorf(`null is a record column's value; here it only renders "", so write ""`)
|
||||||
|
}
|
||||||
|
return &null{}, nil
|
||||||
default:
|
default:
|
||||||
return nil, fmt.Errorf("a template value must be a string, a list or an object, not %s", jsonKind(v))
|
return nil, fmt.Errorf("a template value must be a string, a list or an object, not %s", jsonKind(v))
|
||||||
}
|
}
|
||||||
@@ -99,8 +117,6 @@ func jsonKind(v any) string {
|
|||||||
return "a number"
|
return "a number"
|
||||||
case bool:
|
case bool:
|
||||||
return "a boolean"
|
return "a boolean"
|
||||||
case nil:
|
|
||||||
return "null"
|
|
||||||
}
|
}
|
||||||
return fmt.Sprintf("%T", v)
|
return fmt.Sprintf("%T", v)
|
||||||
}
|
}
|
||||||
@@ -136,7 +152,11 @@ func (t *template) compileFormat() error {
|
|||||||
return checkNoRepeatedRead(t.format, c, t.refs)
|
return checkNoRepeatedRead(t.format, c, t.refs)
|
||||||
}
|
}
|
||||||
|
|
||||||
func compileChoice(items []any) (node, error) {
|
func compileChoice(items []any, pos position) (node, error) {
|
||||||
|
itemPos := inFormat
|
||||||
|
if pos == inColumn {
|
||||||
|
itemPos = inColumn
|
||||||
|
}
|
||||||
if len(items) == 0 {
|
if len(items) == 0 {
|
||||||
return nil, fmt.Errorf("empty choice")
|
return nil, fmt.Errorf("empty choice")
|
||||||
}
|
}
|
||||||
@@ -160,7 +180,7 @@ func compileChoice(items []any) (node, error) {
|
|||||||
}
|
}
|
||||||
total += w
|
total += w
|
||||||
cum[i] = total
|
cum[i] = total
|
||||||
n, err := compileItem(raw)
|
n, err := compileItem(raw, itemPos)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, err
|
return nil, err
|
||||||
}
|
}
|
||||||
@@ -203,22 +223,26 @@ func checkNoRepeatedItem(items []any) error {
|
|||||||
return nil
|
return nil
|
||||||
}
|
}
|
||||||
|
|
||||||
func compileTemplate(m map[string]any) (node, error) {
|
func compileTemplate(m map[string]any, pos position) (node, error) {
|
||||||
o, err := readOptions(m)
|
o, err := readOptions(m, pos)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, err
|
return nil, err
|
||||||
}
|
}
|
||||||
fields, err := compileFields(m)
|
fieldPos := inFormat
|
||||||
|
if pos == atTop && o.repeat == 1 {
|
||||||
|
fieldPos = inColumn
|
||||||
|
}
|
||||||
|
fields, err := compileFields(m, fieldPos)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, err
|
return nil, err
|
||||||
}
|
}
|
||||||
if len(fields) == 0 && o.repeat == 1 && !o.weighted {
|
if len(fields) == 0 && o.repeat == 1 && !o.weighted && o.datatype == DataTypeString {
|
||||||
return nil, fmt.Errorf("an object holding only a format is a string; write %q", o.format)
|
return nil, fmt.Errorf("an object holding only a format is a string; write %q", o.format)
|
||||||
}
|
}
|
||||||
if err := checkTokens(o.format, fields); err != nil {
|
if err := checkTokens(o.format, fields); err != nil {
|
||||||
return nil, err
|
return nil, err
|
||||||
}
|
}
|
||||||
t := &template{format: o.format, fields: fields, repeat: o.repeat, separator: o.separator}
|
t := &template{format: o.format, fields: fields, repeat: o.repeat, separator: o.separator, datatype: o.datatype}
|
||||||
if err := t.compileFormat(); err != nil {
|
if err := t.compileFormat(); err != nil {
|
||||||
return nil, err
|
return nil, err
|
||||||
}
|
}
|
||||||
@@ -227,13 +251,14 @@ func compileTemplate(m map[string]any) (node, error) {
|
|||||||
|
|
||||||
// templateOptions is what a template object's option keys say.
|
// templateOptions is what a template object's option keys say.
|
||||||
type templateOptions struct {
|
type templateOptions struct {
|
||||||
|
datatype DataType
|
||||||
format string
|
format string
|
||||||
repeat int
|
repeat int
|
||||||
separator string
|
separator string
|
||||||
weighted bool
|
weighted bool
|
||||||
}
|
}
|
||||||
|
|
||||||
func readOptions(m map[string]any) (templateOptions, error) {
|
func readOptions(m map[string]any, pos position) (templateOptions, error) {
|
||||||
var o templateOptions
|
var o templateOptions
|
||||||
format, ok := m["format"].(string)
|
format, ok := m["format"].(string)
|
||||||
if !ok {
|
if !ok {
|
||||||
@@ -245,6 +270,9 @@ func readOptions(m map[string]any) (templateOptions, error) {
|
|||||||
return o, err
|
return o, err
|
||||||
}
|
}
|
||||||
o.repeat = repeat
|
o.repeat = repeat
|
||||||
|
if o.datatype, err = datatypeOf(m, pos); err != nil {
|
||||||
|
return o, err
|
||||||
|
}
|
||||||
if sv, ok := m["separator"]; ok {
|
if sv, ok := m["separator"]; ok {
|
||||||
if o.separator, ok = sv.(string); !ok {
|
if o.separator, ok = sv.(string); !ok {
|
||||||
return o, fmt.Errorf("separator must be a string, got %T", sv)
|
return o, fmt.Errorf("separator must be a string, got %T", sv)
|
||||||
@@ -262,7 +290,7 @@ func readOptions(m map[string]any) (templateOptions, error) {
|
|||||||
|
|
||||||
// compileFields compiles every non-option key of a template object, in name order
|
// compileFields compiles every non-option key of a template object, in name order
|
||||||
// so which of several bad fields is reported does not vary.
|
// so which of several bad fields is reported does not vary.
|
||||||
func compileFields(m map[string]any) (map[string]node, error) {
|
func compileFields(m map[string]any, pos position) (map[string]node, error) {
|
||||||
fields := make(map[string]node, len(m))
|
fields := make(map[string]node, len(m))
|
||||||
keys := make([]string, 0, len(m))
|
keys := make([]string, 0, len(m))
|
||||||
for k := range m {
|
for k := range m {
|
||||||
@@ -276,7 +304,10 @@ func compileFields(m map[string]any) (map[string]node, error) {
|
|||||||
if err := checkName(k); err != nil {
|
if err := checkName(k); err != nil {
|
||||||
return nil, fmt.Errorf("field %w", err)
|
return nil, fmt.Errorf("field %w", err)
|
||||||
}
|
}
|
||||||
n, err := compile(m[k])
|
n, err := compileAt(m[k], pos)
|
||||||
|
if err == nil && pos == inColumn {
|
||||||
|
_, err = columnDatatype(n)
|
||||||
|
}
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, fmt.Errorf("field %q: %w", k, err)
|
return nil, fmt.Errorf("field %q: %w", k, err)
|
||||||
}
|
}
|
||||||
@@ -360,10 +391,10 @@ func checkName(name string) error {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// isOption reports whether a template key configures the node instead of naming a
|
// isOption reports whether a template key configures the node instead of naming a
|
||||||
// field. These four names can never be fields.
|
// field. These names can never be fields.
|
||||||
func isOption(name string) bool {
|
func isOption(name string) bool {
|
||||||
switch name {
|
switch name {
|
||||||
case "format", "repeat", "separator", "weight":
|
case "datatype", "format", "repeat", "separator", "weight":
|
||||||
return true
|
return true
|
||||||
}
|
}
|
||||||
return false
|
return false
|
||||||
|
|||||||
@@ -60,7 +60,7 @@ func walkPath(n node, tail []string, w pathWalk) error {
|
|||||||
}
|
}
|
||||||
return nil
|
return nil
|
||||||
}
|
}
|
||||||
return fmt.Errorf("cannot descend into %T at %q", n, tail[0])
|
return fmt.Errorf("no field %q", tail[0])
|
||||||
}
|
}
|
||||||
|
|
||||||
// carriedByAll is the choice rule a path that must resolve on every call obeys:
|
// carriedByAll is the choice rule a path that must resolve on every call obeys:
|
||||||
|
|||||||
@@ -9,10 +9,14 @@ import (
|
|||||||
"strings"
|
"strings"
|
||||||
)
|
)
|
||||||
|
|
||||||
// Column is one rendered column of a record.
|
// Column is one rendered column of a record. Value is the rendered text, which a
|
||||||
|
// serializer quotes for DataTypeString and writes bare for any other datatype; a Null
|
||||||
|
// column has no Value.
|
||||||
type Column struct {
|
type Column struct {
|
||||||
Name string
|
Name string
|
||||||
|
DataType DataType
|
||||||
Value string
|
Value string
|
||||||
|
Null bool
|
||||||
}
|
}
|
||||||
|
|
||||||
// Record is one record rendered from a template: every direct field is a column,
|
// Record is one record rendered from a template: every direct field is a column,
|
||||||
@@ -28,70 +32,93 @@ func (r *Record) Columns() []Column {
|
|||||||
return append([]Column(nil), r.columns...)
|
return append([]Column(nil), r.columns...)
|
||||||
}
|
}
|
||||||
|
|
||||||
// JSON renders the record as one JSON object, every column a string.
|
// JSON renders the record as one JSON object.
|
||||||
func (r *Record) JSON() string {
|
func (r *Record) JSON() string {
|
||||||
m := make(map[string]string, len(r.columns))
|
var b strings.Builder
|
||||||
for _, c := range r.columns {
|
b.WriteByte('{')
|
||||||
m[c.Name] = c.Value
|
for i, c := range r.columns {
|
||||||
|
if i > 0 {
|
||||||
|
b.WriteByte(',')
|
||||||
}
|
}
|
||||||
b, _ := json.Marshal(m)
|
b.WriteString(jsonString(c.Name))
|
||||||
|
b.WriteByte(':')
|
||||||
|
b.WriteString(literal(c, jsonString, "null"))
|
||||||
|
}
|
||||||
|
b.WriteByte('}')
|
||||||
|
return b.String()
|
||||||
|
}
|
||||||
|
|
||||||
|
func jsonString(s string) string {
|
||||||
|
b, _ := json.Marshal(s)
|
||||||
return string(b)
|
return string(b)
|
||||||
}
|
}
|
||||||
|
|
||||||
// CSVHeader renders the column names as one CSV header line.
|
// CSVHeader renders the column names as one CSV header line.
|
||||||
func (r *Record) CSVHeader() string {
|
func (r *Record) CSVHeader() string {
|
||||||
return csvLine(r.names())
|
fields := make([]string, len(r.columns))
|
||||||
|
for i, c := range r.columns {
|
||||||
|
fields[i] = csvField(c.Name)
|
||||||
|
}
|
||||||
|
return strings.Join(fields, ",")
|
||||||
}
|
}
|
||||||
|
|
||||||
// CSVLine renders the column values as one CSV row.
|
// CSVLine renders the column values as one CSV row: a null column an empty field and an
|
||||||
|
// empty string "", the convention PostgreSQL's COPY reads a null by.
|
||||||
func (r *Record) CSVLine() string {
|
func (r *Record) CSVLine() string {
|
||||||
return csvLine(r.values())
|
fields := make([]string, len(r.columns))
|
||||||
}
|
|
||||||
|
|
||||||
func (r *Record) names() []string {
|
|
||||||
out := make([]string, len(r.columns))
|
|
||||||
for i, c := range r.columns {
|
for i, c := range r.columns {
|
||||||
out[i] = c.Name
|
fields[i] = literal(c, csvField, "")
|
||||||
}
|
}
|
||||||
return out
|
if line := strings.Join(fields, ","); line != "" {
|
||||||
|
return line
|
||||||
|
}
|
||||||
|
return `""` // a blank line is a row every CSV reader drops
|
||||||
}
|
}
|
||||||
|
|
||||||
func (r *Record) values() []string {
|
func csvField(s string) string {
|
||||||
out := make([]string, len(r.columns))
|
if s == "" {
|
||||||
for i, c := range r.columns {
|
return `""`
|
||||||
out[i] = c.Value
|
|
||||||
}
|
}
|
||||||
return out
|
|
||||||
}
|
|
||||||
|
|
||||||
func csvLine(cols []string) string {
|
|
||||||
var b strings.Builder
|
var b strings.Builder
|
||||||
w := csv.NewWriter(&b)
|
w := csv.NewWriter(&b)
|
||||||
_ = w.Write(cols)
|
_ = w.Write([]string{s})
|
||||||
w.Flush()
|
w.Flush()
|
||||||
line := strings.TrimSuffix(b.String(), "\n")
|
return strings.TrimSuffix(b.String(), "\n")
|
||||||
if line == "" {
|
|
||||||
return `""` // a blank line is a row every CSV reader drops
|
|
||||||
}
|
|
||||||
return line
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// SQLInsert renders the record as one INSERT statement into table: identifiers in
|
// SQLInsert renders the record as one INSERT statement into table, identifiers in ANSI
|
||||||
// ANSI double quotes, every value a single-quoted string literal.
|
// double quotes.
|
||||||
func (r *Record) SQLInsert(table string) string {
|
func (r *Record) SQLInsert(table string) string {
|
||||||
cols := make([]string, len(r.columns))
|
cols := make([]string, len(r.columns))
|
||||||
vals := make([]string, len(r.columns))
|
vals := make([]string, len(r.columns))
|
||||||
for i, c := range r.columns {
|
for i, c := range r.columns {
|
||||||
cols[i] = quoteIdent(c.Name)
|
cols[i] = quoteIdent(c.Name)
|
||||||
vals[i] = "'" + strings.ReplaceAll(c.Value, "'", "''") + "'"
|
vals[i] = literal(c, sqlString, "NULL")
|
||||||
}
|
}
|
||||||
return fmt.Sprintf("INSERT INTO %s (%s) VALUES (%s);", quoteIdent(table), strings.Join(cols, ", "), strings.Join(vals, ", "))
|
return fmt.Sprintf("INSERT INTO %s (%s) VALUES (%s);", quoteIdent(table), strings.Join(cols, ", "), strings.Join(vals, ", "))
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func sqlString(s string) string {
|
||||||
|
return "'" + strings.ReplaceAll(s, "'", "''") + "'"
|
||||||
|
}
|
||||||
|
|
||||||
func quoteIdent(s string) string {
|
func quoteIdent(s string) string {
|
||||||
return `"` + strings.ReplaceAll(s, `"`, `""`) + `"`
|
return `"` + strings.ReplaceAll(s, `"`, `""`) + `"`
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// literal spells a column the way a serializer writes it: quoted for a string, bare for
|
||||||
|
// any other datatype, whose every render the load check proved a literal, and nullText
|
||||||
|
// for a null.
|
||||||
|
func literal(c Column, quote func(string) string, nullText string) string {
|
||||||
|
switch {
|
||||||
|
case c.Null:
|
||||||
|
return nullText
|
||||||
|
case c.DataType == DataTypeString:
|
||||||
|
return quote(c.Value)
|
||||||
|
}
|
||||||
|
return c.Value
|
||||||
|
}
|
||||||
|
|
||||||
// FakeRecord renders a path as one record: the template it names, with each direct
|
// FakeRecord renders a path as one record: the template it names, with each direct
|
||||||
// field drawn as a column. Only a category-level template is a record — a path
|
// field drawn as a column. Only a category-level template is a record — a path
|
||||||
// that descends into a field, or that names a folder or a choice, is an error.
|
// that descends into a field, or that names a folder or a choice, is an error.
|
||||||
@@ -116,7 +143,7 @@ func (f *Generator) FakeRecord(path string) (*Record, error) {
|
|||||||
// columns, or why it is not a record.
|
// columns, or why it is not a record.
|
||||||
type recordShape struct {
|
type recordShape struct {
|
||||||
t *template
|
t *template
|
||||||
columns []string
|
columns []Column
|
||||||
err error
|
err error
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -140,7 +167,7 @@ func (f *Generator) recordShapeOf(n node) recordShape {
|
|||||||
type RecordTemplate struct {
|
type RecordTemplate struct {
|
||||||
g *Generator
|
g *Generator
|
||||||
t *template
|
t *template
|
||||||
columns []string
|
columns []Column
|
||||||
}
|
}
|
||||||
|
|
||||||
// Fake renders the record with one draw.
|
// Fake renders the record with one draw.
|
||||||
@@ -175,7 +202,7 @@ func (f *Generator) FakeRecordTemplate(input string) (*Record, error) {
|
|||||||
|
|
||||||
// recordOf is the fence both record entry points pass. The columns come back with
|
// recordOf is the fence both record entry points pass. The columns come back with
|
||||||
// the template, fixed for every draw the caller goes on to make.
|
// the template, fixed for every draw the caller goes on to make.
|
||||||
func recordOf(n node) (*template, []string, error) {
|
func recordOf(n node) (*template, []Column, error) {
|
||||||
t, ok := n.(*template)
|
t, ok := n.(*template)
|
||||||
if !ok {
|
if !ok {
|
||||||
return nil, nil, errors.New("names a choice, not a template; a record is a template whose fields are its columns")
|
return nil, nil, errors.New("names a choice, not a template; a record is a template whose fields are its columns")
|
||||||
@@ -183,13 +210,18 @@ func recordOf(n node) (*template, []string, error) {
|
|||||||
if t.repeat != 1 {
|
if t.repeat != 1 {
|
||||||
return nil, nil, fmt.Errorf("carries repeat %d, which composes its format into one string; a record projects columns instead — drop the repeat and render the record again for more rows", t.repeat)
|
return nil, nil, fmt.Errorf("carries repeat %d, which composes its format into one string; a record projects columns instead — drop the repeat and render the record again for more rows", t.repeat)
|
||||||
}
|
}
|
||||||
columns := recordColumns(t)
|
names := recordColumns(t)
|
||||||
if len(columns) == 0 {
|
if len(names) == 0 {
|
||||||
return nil, nil, errors.New("has no fields, so no columns")
|
return nil, nil, errors.New("has no fields, so no columns")
|
||||||
}
|
}
|
||||||
if err := checkColumnRefs(t, columns); err != nil {
|
if err := checkColumnRefs(t, names); err != nil {
|
||||||
return nil, nil, err
|
return nil, nil, err
|
||||||
}
|
}
|
||||||
|
columns := make([]Column, len(names))
|
||||||
|
for i, name := range names {
|
||||||
|
datatype, _ := columnDatatype(t.fields[name]) // compile refused a column whose items disagree
|
||||||
|
columns[i] = Column{Name: name, DataType: datatype}
|
||||||
|
}
|
||||||
return t, columns, nil
|
return t, columns, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -262,11 +294,16 @@ func columnRefs(t *template, columns []string) ([]columnRef, error) {
|
|||||||
|
|
||||||
// renderRecord draws each column once, in the name order recordOf fixed, over one
|
// renderRecord draws each column once, in the name order recordOf fixed, over one
|
||||||
// reference scope shared across them.
|
// reference scope shared across them.
|
||||||
func renderRecord(s *session, t *template, columns []string) *Record {
|
func renderRecord(s *session, t *template, columns []Column) *Record {
|
||||||
scope := &draws{variant: map[string]node{}, value: map[string]string{}}
|
scope := &draws{variant: map[string]node{}, value: map[string]string{}}
|
||||||
r := &Record{columns: make([]Column, len(columns))}
|
r := &Record{columns: append([]Column(nil), columns...)}
|
||||||
for i, name := range columns {
|
for i := range r.columns {
|
||||||
r.columns[i] = Column{Name: name, Value: render(s, t.fields[name], scope)}
|
n := drawn(s, t.fields[r.columns[i].Name])
|
||||||
|
if _, isNull := n.(*null); isNull {
|
||||||
|
r.columns[i].Null = true
|
||||||
|
} else {
|
||||||
|
r.columns[i].Value = render(s, n, scope)
|
||||||
|
}
|
||||||
}
|
}
|
||||||
return r
|
return r
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -56,6 +56,8 @@ func render(s *session, n node, refScope *draws) string {
|
|||||||
switch n := n.(type) {
|
switch n := n.(type) {
|
||||||
case *choice:
|
case *choice:
|
||||||
return render(s, pick(s, n), refScope)
|
return render(s, pick(s, n), refScope)
|
||||||
|
case *null:
|
||||||
|
return ""
|
||||||
case *template:
|
case *template:
|
||||||
if n.repeat == 1 {
|
if n.repeat == 1 {
|
||||||
if n.fixed {
|
if n.fixed {
|
||||||
|
|||||||
+403
@@ -0,0 +1,403 @@
|
|||||||
|
package fejkdata
|
||||||
|
|
||||||
|
import (
|
||||||
|
"slices"
|
||||||
|
"strconv"
|
||||||
|
"strings"
|
||||||
|
"unicode/utf8"
|
||||||
|
)
|
||||||
|
|
||||||
|
// grammar is a deterministic automaton over a scalar's text: state 0 is dead, 1 the
|
||||||
|
// start, and each state lists the runes that leave it and where they lead.
|
||||||
|
type grammar [][]arc
|
||||||
|
|
||||||
|
type arc struct {
|
||||||
|
on string
|
||||||
|
to int
|
||||||
|
}
|
||||||
|
|
||||||
|
func (g *grammar) run(q int, s string) int {
|
||||||
|
for _, r := range s {
|
||||||
|
if q = g.step(q, r); q == 0 {
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return q
|
||||||
|
}
|
||||||
|
|
||||||
|
func (g *grammar) step(q int, r rune) int {
|
||||||
|
for _, a := range (*g)[q] {
|
||||||
|
if strings.ContainsRune(a.on, r) {
|
||||||
|
return a.to
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
const (
|
||||||
|
decimalDigits = "0123456789"
|
||||||
|
nonZeroDigits = "123456789"
|
||||||
|
)
|
||||||
|
|
||||||
|
// numberGrammar reads a JSON number. States: 2 "-", 3 "0", 4 more integer digits, 5 ".",
|
||||||
|
// 6 fraction digits, 7 "e", 8 its sign, 9 exponent digits.
|
||||||
|
var numberGrammar = &grammar{
|
||||||
|
nil,
|
||||||
|
{{"-", 2}, {"0", 3}, {nonZeroDigits, 4}},
|
||||||
|
{{"0", 3}, {nonZeroDigits, 4}},
|
||||||
|
{{".", 5}, {"eE", 7}},
|
||||||
|
{{decimalDigits, 4}, {".", 5}, {"eE", 7}},
|
||||||
|
{{decimalDigits, 6}},
|
||||||
|
{{decimalDigits, 6}, {"eE", 7}},
|
||||||
|
{{"+-", 8}, {decimalDigits, 9}},
|
||||||
|
{{decimalDigits, 9}},
|
||||||
|
{{decimalDigits, 9}},
|
||||||
|
}
|
||||||
|
|
||||||
|
const (
|
||||||
|
integerAccept uint32 = 1<<3 | 1<<4
|
||||||
|
numberAccept = integerAccept | 1<<6 | 1<<9
|
||||||
|
)
|
||||||
|
|
||||||
|
var booleanGrammar = &grammar{
|
||||||
|
nil,
|
||||||
|
{{"t", 2}, {"f", 6}},
|
||||||
|
{{"r", 3}}, {{"u", 4}}, {{"e", 5}}, nil,
|
||||||
|
{{"a", 7}}, {{"l", 8}}, {{"s", 9}}, {{"e", 10}}, nil,
|
||||||
|
}
|
||||||
|
|
||||||
|
const booleanAccept uint32 = 1<<5 | 1<<10
|
||||||
|
|
||||||
|
// decimalGrammar reads what a calc operand must render to be proven finite: a sign,
|
||||||
|
// digits and at most one dot. Past the sign, states 4–9 are positive and 10–15 their
|
||||||
|
// negatives: 4 zero digits, 5 a nonzero integer, 6 a leading dot, 7 zero with a dot,
|
||||||
|
// 8 a nonzero integer with a zero fraction, 9 a nonzero fraction.
|
||||||
|
var decimalGrammar = &grammar{
|
||||||
|
nil,
|
||||||
|
{{"+", 2}, {"-", 3}, {"0", 4}, {nonZeroDigits, 5}, {".", 6}},
|
||||||
|
{{"0", 4}, {nonZeroDigits, 5}, {".", 6}},
|
||||||
|
{{"0", 10}, {nonZeroDigits, 11}, {".", 12}},
|
||||||
|
{{"0", 4}, {nonZeroDigits, 5}, {".", 7}},
|
||||||
|
{{decimalDigits, 5}, {".", 8}},
|
||||||
|
{{"0", 7}, {nonZeroDigits, 9}},
|
||||||
|
{{"0", 7}, {nonZeroDigits, 9}},
|
||||||
|
{{"0", 8}, {nonZeroDigits, 9}},
|
||||||
|
{{decimalDigits, 9}},
|
||||||
|
{{"0", 10}, {nonZeroDigits, 11}, {".", 13}},
|
||||||
|
{{decimalDigits, 11}, {".", 14}},
|
||||||
|
{{"0", 13}, {nonZeroDigits, 15}},
|
||||||
|
{{"0", 13}, {nonZeroDigits, 15}},
|
||||||
|
{{"0", 14}, {nonZeroDigits, 15}},
|
||||||
|
{{decimalDigits, 15}},
|
||||||
|
}
|
||||||
|
|
||||||
|
const (
|
||||||
|
decimalAccept uint32 = 1<<4 | 1<<5 | 1<<7 | 1<<8 | 1<<9 | 1<<10 | 1<<11 | 1<<13 | 1<<14 | 1<<15
|
||||||
|
decimalNegative uint32 = 0xfc00
|
||||||
|
decimalZero uint32 = 1<<4 | 1<<7 | 1<<10 | 1<<13
|
||||||
|
decimalFractional uint32 = 1<<9 | 1<<15
|
||||||
|
)
|
||||||
|
|
||||||
|
// relation is what a node's renders do to a grammar: from each state, the states a
|
||||||
|
// render can end in, and one render reaching each.
|
||||||
|
type relation struct {
|
||||||
|
g *grammar
|
||||||
|
to []uint32
|
||||||
|
w []witness // w[from*len(to)+to]
|
||||||
|
}
|
||||||
|
|
||||||
|
// witness is one render, cut past witnessCap bytes, and why it can occur when the text
|
||||||
|
// alone does not say.
|
||||||
|
type witness struct {
|
||||||
|
text string
|
||||||
|
cut bool
|
||||||
|
why string
|
||||||
|
}
|
||||||
|
|
||||||
|
const witnessCap = 60
|
||||||
|
|
||||||
|
func (w witness) then(next witness) witness {
|
||||||
|
if w.why == "" {
|
||||||
|
w.why = next.why
|
||||||
|
}
|
||||||
|
if w.cut {
|
||||||
|
return w
|
||||||
|
}
|
||||||
|
w.text += next.text
|
||||||
|
w.cut = next.cut
|
||||||
|
if len(w.text) > witnessCap {
|
||||||
|
end := witnessCap
|
||||||
|
for !utf8.RuneStart(w.text[end]) {
|
||||||
|
end--
|
||||||
|
}
|
||||||
|
w.text, w.cut = w.text[:end], true
|
||||||
|
}
|
||||||
|
return w
|
||||||
|
}
|
||||||
|
|
||||||
|
func (w witness) String() string {
|
||||||
|
if w.cut {
|
||||||
|
return strconv.Quote(w.text + "…")
|
||||||
|
}
|
||||||
|
return strconv.Quote(w.text)
|
||||||
|
}
|
||||||
|
|
||||||
|
func newRelation(g *grammar) *relation {
|
||||||
|
n := len(*g)
|
||||||
|
return &relation{g: g, to: make([]uint32, n), w: make([]witness, n*n)}
|
||||||
|
}
|
||||||
|
|
||||||
|
func (r *relation) add(from, to int, w witness) {
|
||||||
|
if r.to[from]&(1<<to) == 0 {
|
||||||
|
r.to[from] |= 1 << to
|
||||||
|
r.w[from*len(r.to)+to] = w
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// textRelation is the relation of a render that is always s.
|
||||||
|
func textRelation(g *grammar, s, why string) *relation {
|
||||||
|
r := newRelation(g)
|
||||||
|
w := witness{why: why}.then(witness{text: s})
|
||||||
|
for q := range r.to {
|
||||||
|
r.add(q, g.run(q, s), w)
|
||||||
|
}
|
||||||
|
return r
|
||||||
|
}
|
||||||
|
|
||||||
|
// union is the renders of either relation; a nil relation has none.
|
||||||
|
func union(a, b *relation) *relation {
|
||||||
|
if a == nil {
|
||||||
|
return b
|
||||||
|
}
|
||||||
|
if b == nil {
|
||||||
|
return a
|
||||||
|
}
|
||||||
|
u := newRelation(a.g)
|
||||||
|
for _, r := range []*relation{a, b} {
|
||||||
|
for from, ends := range r.to {
|
||||||
|
for to := range r.to {
|
||||||
|
if ends&(1<<to) != 0 {
|
||||||
|
u.add(from, to, r.w[from*len(r.to)+to])
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return u
|
||||||
|
}
|
||||||
|
|
||||||
|
// then is a render of r followed by a render of next.
|
||||||
|
func (r *relation) then(next *relation) *relation {
|
||||||
|
c := newRelation(r.g)
|
||||||
|
n := len(r.to)
|
||||||
|
for from, mids := range r.to {
|
||||||
|
for mid := 0; mid < n; mid++ {
|
||||||
|
if mids&(1<<mid) == 0 {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
for to := 0; to < n; to++ {
|
||||||
|
if next.to[mid]&(1<<to) != 0 && c.to[from]&(1<<to) == 0 {
|
||||||
|
c.add(from, to, r.w[from*n+mid].then(next.w[mid*n+to]))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return c
|
||||||
|
}
|
||||||
|
|
||||||
|
// power is k renders of r in a row, k at least 1, composed by squaring.
|
||||||
|
func (r *relation) power(k int) *relation {
|
||||||
|
var out *relation
|
||||||
|
for base := r; ; base = base.then(base) {
|
||||||
|
if k&1 == 1 {
|
||||||
|
if out == nil {
|
||||||
|
out = base
|
||||||
|
} else {
|
||||||
|
out = out.then(base)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if k >>= 1; k == 0 {
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// closure is any number of renders of r in a row, where r includes the empty render.
|
||||||
|
func (r *relation) closure() *relation {
|
||||||
|
for {
|
||||||
|
next := r.then(r)
|
||||||
|
if slices.Equal(next.to, r.to) {
|
||||||
|
return r
|
||||||
|
}
|
||||||
|
r = next
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// escape finds a render from the start that ends outside accept, preferring one that
|
||||||
|
// carries a reason.
|
||||||
|
func (r *relation) escape(accept uint32) (witness, bool) {
|
||||||
|
var found witness
|
||||||
|
escapes := false
|
||||||
|
for to := range r.to {
|
||||||
|
if (r.to[1]&^accept)&(1<<to) == 0 {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if w := r.w[len(r.to)+to]; !escapes || found.why == "" && w.why != "" {
|
||||||
|
found, escapes = w, true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return found, escapes
|
||||||
|
}
|
||||||
|
|
||||||
|
// textShape is the text a builtin can emit: alternatives, each a sequence of runs.
|
||||||
|
type textShape [][]charRun
|
||||||
|
|
||||||
|
// charRun is between min and max characters, each one of chars; max -1 is unbounded.
|
||||||
|
// chars is ASCII, so a run of k characters is k bytes.
|
||||||
|
type charRun struct {
|
||||||
|
chars string
|
||||||
|
min, max int
|
||||||
|
}
|
||||||
|
|
||||||
|
// textLanguage is what one grammar makes of the renders a check reads, worked out once
|
||||||
|
// per node and fold.
|
||||||
|
type textLanguage struct {
|
||||||
|
g *grammar
|
||||||
|
proof *calcProof
|
||||||
|
memo map[languageKey]*relation
|
||||||
|
empty *relation
|
||||||
|
}
|
||||||
|
|
||||||
|
type languageKey struct {
|
||||||
|
n node
|
||||||
|
fold string
|
||||||
|
}
|
||||||
|
|
||||||
|
// fold is the transforms a render passes through before the grammar reads it, innermost
|
||||||
|
// first. Each rewrites rune by rune, so folding a render is folding each of its pieces.
|
||||||
|
type fold []string
|
||||||
|
|
||||||
|
func (f fold) apply(s string) string {
|
||||||
|
for _, name := range f {
|
||||||
|
s = transforms[name](s)
|
||||||
|
}
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
|
||||||
|
func newTextLanguage(g *grammar, proof *calcProof) *textLanguage {
|
||||||
|
return &textLanguage{g: g, proof: proof, memo: map[languageKey]*relation{}, empty: textRelation(g, "", "")}
|
||||||
|
}
|
||||||
|
|
||||||
|
func (l *textLanguage) node(n node, f fold) *relation {
|
||||||
|
key := languageKey{n, strings.Join(f, ",")}
|
||||||
|
if r, done := l.memo[key]; done {
|
||||||
|
return r
|
||||||
|
}
|
||||||
|
r := l.empty // a null renders ""
|
||||||
|
switch n := n.(type) {
|
||||||
|
case *choice:
|
||||||
|
r = nil
|
||||||
|
for _, it := range n.items {
|
||||||
|
r = union(r, l.node(it, f))
|
||||||
|
}
|
||||||
|
case *template:
|
||||||
|
r = l.format(n, f)
|
||||||
|
if n.repeat > 1 {
|
||||||
|
r = r.then(l.text(f.apply(n.separator)).then(r).power(n.repeat - 1))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
l.memo[key] = r
|
||||||
|
return r
|
||||||
|
}
|
||||||
|
|
||||||
|
func (l *textLanguage) text(s string) *relation { return textRelation(l.g, s, "") }
|
||||||
|
|
||||||
|
// format reads a template's format the way expand renders it: literal runs and tokens
|
||||||
|
// in turn.
|
||||||
|
func (l *textLanguage) format(t *template, f fold) *relation {
|
||||||
|
r := l.empty
|
||||||
|
var lit strings.Builder
|
||||||
|
_ = eachToken(t.format, func(tok ftoken) error {
|
||||||
|
if tok.kind == 'l' {
|
||||||
|
lit.WriteRune(tok.r)
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
r = r.then(l.text(f.apply(lit.String()))).then(l.token(t, tok.body, f))
|
||||||
|
lit.Reset()
|
||||||
|
return nil
|
||||||
|
})
|
||||||
|
return r.then(l.text(f.apply(lit.String())))
|
||||||
|
}
|
||||||
|
|
||||||
|
// token reads one {…} token: a field read, a transform over one, a calc, or what a
|
||||||
|
// builtin emits.
|
||||||
|
func (l *textLanguage) token(t *template, body string, f fold) *relation {
|
||||||
|
name, args, isFunc := funcCall(body)
|
||||||
|
if !isFunc {
|
||||||
|
var r *relation
|
||||||
|
for _, a := range splitArms(body, t.refs) {
|
||||||
|
r = union(r, l.read(t, a, f))
|
||||||
|
}
|
||||||
|
return r
|
||||||
|
}
|
||||||
|
if _, isTransform := transforms[name]; isTransform {
|
||||||
|
leaf, chain, _ := unwrapTransform(args[0])
|
||||||
|
inner := slices.Clone(chain)
|
||||||
|
slices.Reverse(inner)
|
||||||
|
return l.read(t, splitArm(leaf, t.refs), append(append(inner, name), f...))
|
||||||
|
}
|
||||||
|
if name == "calc" {
|
||||||
|
return l.calc(t, args, f)
|
||||||
|
}
|
||||||
|
return l.shape(builtins[name].emits(args), f)
|
||||||
|
}
|
||||||
|
|
||||||
|
// read is one arm of a token: every node its path can land on.
|
||||||
|
func (l *textLanguage) read(t *template, a arm, f fold) *relation {
|
||||||
|
var r *relation
|
||||||
|
for _, leaf := range pathLeaves(t.fields[a.key], a.tail) {
|
||||||
|
r = union(r, l.node(leaf, f))
|
||||||
|
}
|
||||||
|
return r
|
||||||
|
}
|
||||||
|
|
||||||
|
func (l *textLanguage) calc(t *template, args []string, f fold) *relation {
|
||||||
|
b, d := l.proof.call(t, args)
|
||||||
|
if d != nil {
|
||||||
|
return textRelation(l.g, f.apply(d.render), d.why)
|
||||||
|
}
|
||||||
|
return l.shape(printedFloat(b.lo, b.hi, calcDecimals(args), b.integral), f)
|
||||||
|
}
|
||||||
|
|
||||||
|
func (l *textLanguage) shape(s textShape, f fold) *relation {
|
||||||
|
var r *relation
|
||||||
|
for _, alt := range s {
|
||||||
|
seq := l.empty
|
||||||
|
for _, run := range alt {
|
||||||
|
seq = seq.then(l.run(run, f))
|
||||||
|
}
|
||||||
|
r = union(r, seq)
|
||||||
|
}
|
||||||
|
return r
|
||||||
|
}
|
||||||
|
|
||||||
|
// run reads a charRun: min characters, then up to max-min more.
|
||||||
|
func (l *textLanguage) run(c charRun, f fold) *relation {
|
||||||
|
one := newRelation(l.g)
|
||||||
|
for from := range one.to {
|
||||||
|
for _, ch := range c.chars {
|
||||||
|
s := f.apply(string(ch))
|
||||||
|
one.add(from, l.g.run(from, s), witness{text: s})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
more := l.empty
|
||||||
|
switch optional := union(one, l.empty); {
|
||||||
|
case c.max < 0:
|
||||||
|
more = optional.closure()
|
||||||
|
case c.max > c.min:
|
||||||
|
more = optional.power(c.max - c.min)
|
||||||
|
}
|
||||||
|
if c.min == 0 {
|
||||||
|
return more
|
||||||
|
}
|
||||||
|
return one.power(c.min).then(more)
|
||||||
|
}
|
||||||
@@ -72,6 +72,9 @@ type builtin struct {
|
|||||||
// operands names the fields the call reads, which expand renders for it; nil
|
// operands names the fields the call reads, which expand renders for it; nil
|
||||||
// for a builtin that reads none.
|
// for a builtin that reads none.
|
||||||
operands func(args []string) []string
|
operands func(args []string) []string
|
||||||
|
// emits is the text a call can print, for the datatype check; nil for calc and the
|
||||||
|
// transforms, whose text the check derives from what they read.
|
||||||
|
emits func(args []string) textShape
|
||||||
}
|
}
|
||||||
|
|
||||||
// funcCall splits a "{token}" body shaped name(args) into its parts; ok is false
|
// funcCall splits a "{token}" body shaped name(args) into its parts; ok is false
|
||||||
|
|||||||
@@ -6,11 +6,6 @@ The record API lands first, so the data update can use it.
|
|||||||
|
|
||||||
### Record API
|
### Record API
|
||||||
|
|
||||||
- Typed columns — a column declares its type, so `json` writes `42` rather than
|
|
||||||
`"42"` and `sql` an unquoted literal: string, integer, number, boolean, and a
|
|
||||||
way to write null. A template that can render a value its type rejects is a
|
|
||||||
load error. The option key is reserved from then on, so a common column name
|
|
||||||
like `type` is a poor pick.
|
|
||||||
- Struct-filling — fill a Go struct from `fake:"…"` tags holding a path or an
|
- Struct-filling — fill a Go struct from `fake:"…"` tags holding a path or an
|
||||||
inline template, for parity with gofakeit and go-faker. The field's Go type is
|
inline template, for parity with gofakeit and go-faker. The field's Go type is
|
||||||
the column type, through the same conversion and load checks as typed columns,
|
the column type, through the same conversion and load checks as typed columns,
|
||||||
|
|||||||
Reference in New Issue
Block a user