Typed columns and null #13

Merged
lilleman merged 10 commits from typed-columns into main 2026-09-15 15:19:39 +02:00
11 changed files with 379 additions and 844 deletions
Showing only changes of commit 5615ab6d89 - Show all commits
+38 -19
View File
@@ -271,25 +271,31 @@ string:
Writes e.g. `{"id":1,"paid":true,"total":59.97}`. A column is a field of the top-level Writes e.g. `{"id":1,"paid":true,"total":59.97}`. A column is a field of the top-level
template, or an item of a choice standing in for one; `datatype` anywhere else is a template, or an item of a choice standing in for one; `datatype` anywhere else is a
load error. So is a column that can render text its datatype rejects — `integer` takes load error. A typed column holds one value, alone in its format: a literal, one
`-?(0|[1-9][0-9]*)`, `number` a JSON number, `boolean` `true` or `false` — and the `{int()}`, `{float()}`, `{seq()}` or `{calc()}` call, or a read that lands only on such
error shows such a render: values. `integer` is an int64 written `-?(0|[1-9][0-9]*)` — `{float()}` prints one at
`0` decimals — `number` a JSON number, `boolean` `true` or `false`. A value its
datatype cannot hold is a load error naming it:
```text ```text
order.id: datatype integer, but it can render "000", which is not an integer order.id: datatype integer: {digits(3)} prints text, not an integer
order.id: datatype integer: "1{digits(2)}" is not one value; write one literal or one {int()}, {float()}, {seq()} or {calc()}, or read one
``` ```
A `{calc()}` fills an `integer` or `number` column only where it provably prints no A typed column's `{calc()}` must be proven to print a number: each operand a number
`NaN` or `Inf`: each operand is a plain decimal — a sign, digits, one dot — of at most literal, an `{int()}`, `{float()}`, `{seq()}` or `{digits()}` call, a calc, or a read of
300 bytes, or a field holding only such a calc, and no divisor can be zero. An such values, whose bounds keep every divisor from zero and the result within `1e300`.
`integer` column also needs a decimals count of `0`, or integer operands and no `/`. What the bounds cannot show is refused — `{calc(a / b)}: divides by b, which is not
proven nonzero`. The calc fills an `integer` column at `0` decimals, or over whole
operands with no `/`.
### Null ### Null
A `null` item draws a record column as null: `json` writes `null`, `sql` `NULL`, and A `null` item draws a record column as null: `json` writes `null`, `sql` `NULL`, and
`csv` an empty field, with an empty string written `""` so PostgreSQL's `COPY … CSV` `csv` an empty field, with an empty string written `""` — the convention PostgreSQL's
reads both back. `Fake` renders a null as `""`. The other items' weights skew its `COPY … CSV` reads. A record of one null column is a blank line, which `COPY` reads as
odds: null but most CSV readers skip, so write such a record as `json` or `sql`. `Fake`
renders a null as `""`. The other items' weights skew its odds:
```json ```json
{ "format": "", "deleted_at": null, "middle": [null, { "format": "{n}", "n": ["Ann", "Eva"], "weight": 3 }] } { "format": "", "deleted_at": null, "middle": [null, { "format": "{n}", "n": ["Ann", "Eva"], "weight": 3 }] }
@@ -363,7 +369,7 @@ Renders e.g. `19.99 x 3 = 59.97`. An operand that can never be a number (`"abc"`
or a choice of such) is rejected at load, as is a division by a constant zero or a choice of such) is rejected at load, as is a division by a constant zero
(`1/0`, or a fixed `"0"` field); an operand that sometimes is not a number yields (`1/0`, or a fixed `"0"` field); an operand that sometimes is not a number yields
`NaN`, and a division by one that is not constant `Inf` — both print rather than `NaN`, and a division by one that is not constant `Inf` — both print rather than
fail. fail, except in a [typed column](#datatype), which must prove neither happens.
### Transforms ### Transforms
@@ -556,11 +562,13 @@ tokens add cost in proportion to the output.
- **64-bit targets only.** The gate builds amd64, and the buffer sizing a render - **64-bit targets only.** The gate builds amd64, and the buffer sizing a render
pre-computes (renders × bytes) assumes a 64-bit int; on a 32-bit target it could pre-computes (renders × bytes) assumes a 64-bit int; on a 32-bit target it could
overflow and panic. overflow and panic.
- **A constant zero divisor is a load error; a divisor that is not constant prints - **A constant zero divisor is a load error; in a string column a divisor that is not
`Inf`.** `1/0` and a fixed `"0"` field are decidable, so they join the constant prints `Inf`.** `1/0` and a fixed `"0"` field are decidable, so they join
never-numeric operand as a load error; the fold stops where an operand varies, the never-numeric operand as a load error; the fold stops where an operand varies,
so `a/(b*c)` with `b` fixed at `0` and `c` varying loads and prints `Inf` every so `a/(b*c)` with `b` fixed at `0` and `c` varying loads and prints `Inf` every
draw — catching it needs zero-absorbing algebra for a shape nobody writes. draw — catching it needs zero-absorbing algebra for a shape nobody writes. A
[typed column](#datatype) bounds its operands instead and refuses a divisor it
cannot keep from zero.
- **In data, a default written out and a constant spelled as a sample are load - **In data, a default written out and a constant spelled as a sample are load
errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `datatype: "string"`, errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `datatype: "string"`,
`int(5,5)`, `float(1,1,2)`, `int(5,5)`, `float(1,1,2)`,
@@ -605,6 +613,17 @@ tokens add cost in proportion to the output.
- **Null is a `null` item, not a rate.** A null is one more outcome of a column's - **Null is a `null` item, not a rate.** A null is one more outcome of a column's
draw, so a choice's weights skew it like any other; a null-rate option would be a draw, so a choice's weights skew it like any other; a null-rate option would be a
second way to state odds. second way to state odds.
- **A typed column holds one value, not composed text.** Its bounds come from a
literal or a call's arguments, so a load error names a real value, a range check is
one comparison, and `1{digits(2)}` is a second spelling of `{int(100,199)}`.
- **A typed column's calc is refused unless proven.** Operand bounds must keep each
divisor from zero and the result finite; what they cannot show is refused rather
than trusted, since a bare `NaN` breaks the JSON and SQL it lands in.
- **`Column` carries text, not a Go value.** `Value` is the rendered string beside
`DataType` and `Null`, which each serializer writes as the load check proved it; a
`Value any` would hand every caller a type switch.
- **The package stays flat.** Go ties a package to one directory, so folders would
split the API into packages.
- **The performance gate asserts allocations, not wall-clock time.** `AllocsPerRun` - **The performance gate asserts allocations, not wall-clock time.** `AllocsPerRun`
is deterministic across machines, so a ±10% ceiling does not flake under CI load, is deterministic across machines, so a ±10% ceiling does not flake under CI load,
while time varies with the machine and its neighbours. A rendering slowdown while time varies with the machine and its neighbours. A rendering slowdown
@@ -660,9 +679,9 @@ hold.go the hold: one draw per expansion for paths and operands, and its
reference.go reference sigils, and binding references across the tree reference.go reference sigils, and binding references across the tree
graph.go the render graph: edges, cycles, the repeat bound, tree walks graph.go the render graph: edges, cycles, the repeat bound, tree walks
builtins.go the {name()} function registry and its implementations builtins.go the {name()} function registry and its implementations
calc.go the {calc()} arithmetic evaluator: parser, eval, validation, and the proof a typed column's calc is finite calc.go the {calc()} arithmetic evaluator: parser, eval, validation
datatype.go column datatypes: DataType, where datatype and null may sit, and the load check every typed render passes datatype.go column datatypes: DataType, where datatype and null may sit, a column's datatype
renderlang.go what text a node can render, as relations over a scalar's grammar value.go the value proof: what a typed column or calc operand holds, checked at load
data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge
cmd/fejkdata/ the fejkdata CLI cmd/fejkdata/ the fejkdata CLI
data/ shipped data (JSON), embedded at build: locale folders + a misc folder data/ shipped data (JSON), embedded at build: locale folders + a misc folder
+29 -101
View File
@@ -5,7 +5,6 @@ import (
"errors" "errors"
"fmt" "fmt"
"math" "math"
"slices"
"strconv" "strconv"
"strings" "strings"
"unicode" "unicode"
@@ -25,36 +24,42 @@ const (
// samples read only the rng. A time-based id (uuid v7, ulid) draws its timestamp // samples read only the rng. A time-based id (uuid v7, ulid) draws its timestamp
// from the rng, not the wall clock, so seeded output stays reproducible. // from the rng, not the wall clock, so seeded output stays reproducible.
var builtins = map[string]builtin{ var builtins = map[string]builtin{
"luhn": {arity: 0, prep: derive(func(e string) string { return string(rune('0' + luhnCheck(e))) }), emits: always(textShape{{{decimalDigits, 1, 1}}})}, "luhn": {arity: 0, prep: derive(func(e string) string { return string(rune('0' + luhnCheck(e))) })},
"mod11": {arity: 0, prep: derive(mod11Check), emits: always(textShape{{{decimalDigits + "X", 1, 1}}})}, "mod11": {arity: 0, prep: derive(mod11Check)},
"ean": {arity: 0, prep: derive(eanCheck), emits: always(textShape{{{decimalDigits, 1, 1}}})}, "ean": {arity: 0, prep: derive(eanCheck)},
"uuid": {arity: 0, prep: sample(uuidV7), emits: always(uuidShape)}, "uuid": {arity: 0, prep: sample(uuidV7)},
"ulid": {arity: 0, prep: sample(ulid), emits: always(textShape{{{crockford[:8], 1, 1}, {crockford, 25, 25}}})}, "ulid": {arity: 0, prep: sample(ulid)},
"nanoid": sampleOf(nanoidAlphabet), "nanoid": {arity: 1, check: posIntArg, prep: chars(nanoidAlphabet)},
"hex": sampleOf(hexDigits), "hex": {arity: 1, check: posIntArg, prep: chars(hexDigits)},
"digits": sampleOf(decimalDigits), "digits": {arity: 1, check: posIntArg, prep: chars("0123456789"), number: func(a []string) (proven, DataType) {
"upper": sampleOf("ABCDEFGHIJKLMNOPQRSTUVWXYZ"), return bounded(0, math.Pow(10, float64(atoi(a[0])))-1, true), DataTypeString
"lower": sampleOf("abcdefghijklmnopqrstuvwxyz"), }},
"upper": {arity: 1, check: posIntArg, prep: chars("ABCDEFGHIJKLMNOPQRSTUVWXYZ")},
"lower": {arity: 1, check: posIntArg, prep: chars("abcdefghijklmnopqrstuvwxyz")},
"base64": {arity: 1, check: posIntArg, prep: func(a []string) callFn { "base64": {arity: 1, check: posIntArg, prep: func(a []string) callFn {
n := atoi(a[0]) n := atoi(a[0])
return func(s *session, _ string, _ []string) string { return func(s *session, _ string, _ []string) string {
return base64.StdEncoding.EncodeToString(randBytes(s, n)) return base64.StdEncoding.EncodeToString(randBytes(s, n))
} }
}, emits: base64Shape}, }},
"int": {arity: 2, check: intRangeArgs, prep: func(a []string) callFn { "int": {arity: 2, check: intRangeArgs, prep: func(a []string) callFn {
lo, span := atoi(a[0]), atoi(a[1])-atoi(a[0])+1 lo, span := atoi(a[0]), atoi(a[1])-atoi(a[0])+1
return func(s *session, _ string, _ []string) string { return strconv.Itoa(lo + s.IntN(span)) } return func(s *session, _ string, _ []string) string { return strconv.Itoa(lo + s.IntN(span)) }
}, emits: intShape}, }, number: func(a []string) (proven, DataType) {
return bounded(float64(atoi(a[0])), float64(atoi(a[1])), true), DataTypeInteger
}},
"float": {arity: 3, check: floatArgs, prep: func(a []string) callFn { "float": {arity: 3, check: floatArgs, prep: func(a []string) callFn {
lo, hi, dp := atof(a[0]), atof(a[1]), atoi(a[2]) lo, hi, dp := atof(a[0]), atof(a[1]), atoi(a[2])
return func(s *session, _ string, _ []string) string { return func(s *session, _ string, _ []string) string {
return strconv.FormatFloat(lo+s.Float64()*(hi-lo), 'f', dp, 64) return strconv.FormatFloat(lo+s.Float64()*(hi-lo), 'f', dp, 64)
} }
}, emits: func(a []string) textShape { return printedFloat(atof(a[0]), atof(a[1]), atoi(a[2]), false) }}, }, number: func(a []string) (proven, DataType) {
return printedNumber(bounded(atof(a[0]), atof(a[1]), false), atoi(a[2]))
}},
"iban": {arity: 1, check: ibanArg, prep: func(a []string) callFn { "iban": {arity: 1, check: ibanArg, prep: func(a []string) callFn {
cc := a[0] cc := a[0]
return func(s *session, _ string, _ []string) string { return iban(s, cc) } return func(s *session, _ string, _ []string) string { return iban(s, cc) }
}, emits: ibanShape}, }},
"calc": {arity: -1, check: checkCalc, prep: calcPrep, operands: calcOperands}, "calc": {arity: -1, check: checkCalc, prep: calcPrep, operands: calcOperands},
"lowercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToLower), operands: transformOperand}, "lowercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToLower), operands: transformOperand},
"uppercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToUpper), operands: transformOperand}, "uppercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToUpper), operands: transformOperand},
@@ -70,7 +75,9 @@ var builtins = map[string]builtin{
return func(s *session, _ string, _ []string) string { return func(s *session, _ string, _ []string) string {
return strconv.FormatUint(s.next(key), 10) return strconv.FormatUint(s.next(key), 10)
} }
}, emits: always(textShape{{{nonZeroDigits, 1, 1}, {decimalDigits, 0, 19}}})}, }, number: func([]string) (proven, DataType) {
return bounded(1, math.MaxInt64, true), DataTypeInteger
}},
} }
// derive and sample are the two argument-free builtin shapes: a derivation reads // derive and sample are the two argument-free builtin shapes: a derivation reads
@@ -95,82 +102,6 @@ func chars(alphabet string) func([]string) callFn {
} }
} }
// sampleOf is the builtin that draws n characters from an alphabet.
func sampleOf(alphabet string) builtin {
return builtin{arity: 1, check: posIntArg, prep: chars(alphabet), emits: func(a []string) textShape {
n := atoi(a[0])
return textShape{{{alphabet, n, n}}}
}}
}
// always is the emits of a builtin whose args do not change what it can print.
func always(s textShape) func([]string) textShape {
return func([]string) textShape { return s }
}
var uuidShape = textShape{{{hexDigits, 8, 8}, {"-", 1, 1}, {hexDigits, 4, 4}, {"-", 1, 1}, {"7", 1, 1}, {hexDigits, 3, 3}, {"-", 1, 1}, {"89ab", 1, 1}, {hexDigits, 3, 3}, {"-", 1, 1}, {hexDigits, 12, 12}}}
const base64Alphabet = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/"
func base64Shape(a []string) textShape {
n := atoi(a[0])
pad := (3 - n%3) % 3
size := 4*((n+2)/3) - pad
return textShape{{{base64Alphabet, size, size}, {"=", pad, pad}}}
}
// intShape is what int prints: a sign only below zero, and no leading zero.
func intShape(a []string) textShape {
lo, hi := atoi(a[0]), atoi(a[1])
var s textShape
if lo <= 0 && hi >= 0 {
s = append(s, []charRun{{"0", 1, 1}})
}
if hi > 0 {
s = append(s, []charRun{{nonZeroDigits, 1, 1}, {decimalDigits, 0, len(a[1]) - 1}})
}
if lo < 0 {
s = append(s, []charRun{{"-", 1, 1}, {nonZeroDigits, 1, 1}, {decimalDigits, 0, len(a[0]) - 2}})
}
return s
}
func ibanShape(a []string) textShape {
cc, digits := a[0], ibanLen[a[0]]-2
return textShape{{{cc[:1], 1, 1}, {cc[1:], 1, 1}, {decimalDigits, digits, digits}}}
}
// shortestFraction bounds the fraction FormatFloat's shortest form prints: at most 17
// significant digits after up to 323 zeros.
const shortestFraction = 340
// printedFloat is what strconv.FormatFloat(v, 'f', dp, 64) prints for a v in [lo, hi]
// that is whole when integral.
func printedFloat(lo, hi float64, dp int, integral bool) textShape {
digits := len(strconv.FormatFloat(math.Floor(math.Max(math.Abs(lo), math.Abs(hi))), 'f', 0, 64)) + 1 // one more for a rounding carry
wholes := [][]charRun{{{"0", 1, 1}}, {{nonZeroDigits, 1, 1}, {decimalDigits, 0, digits - 1}}}
fractions := [][]charRun{nil}
switch {
case dp > 0:
fractions = [][]charRun{{{".", 1, 1}, {decimalDigits, dp, dp}}}
case dp < 0 && !integral:
fractions = append(fractions, []charRun{{".", 1, 1}, {decimalDigits, 1, shortestFraction}})
}
signs := [][]charRun{nil}
if lo < 0 || math.Signbit(lo) {
signs = append(signs, []charRun{{"-", 1, 1}})
}
var s textShape
for _, sign := range signs {
for _, whole := range wholes {
for _, fraction := range fractions {
s = append(s, slices.Concat(sign, whole, fraction))
}
}
}
return s
}
const hexDigits = "0123456789abcdef" const hexDigits = "0123456789abcdef"
// transforms are the builtins that rewrite one operand's value; they nest, so // transforms are the builtins that rewrite one operand's value; they nest, so
@@ -183,19 +114,20 @@ var transforms = map[string]func(string) string{
// unwrapTransform peels nested transform calls off an operand arg, returning the // unwrapTransform peels nested transform calls off an operand arg, returning the
// field it finally names and the transforms to apply, innermost last. // field it finally names and the transforms to apply, innermost last.
func unwrapTransform(arg string) (leaf string, chain []string, err error) { func unwrapTransform(arg string) (leaf string, chain []func(string) string, err error) {
for { for {
name, args, isCall := funcCall(arg) name, args, isCall := funcCall(arg)
if !isCall { if !isCall {
return arg, chain, nil return arg, chain, nil
} }
if _, isTransform := transforms[name]; !isTransform { fn, isTransform := transforms[name]
if !isTransform {
return "", nil, fmt.Errorf("%s(%s) is not a transform, so it cannot be an operand", name, strings.Join(args, ",")) return "", nil, fmt.Errorf("%s(%s) is not a transform, so it cannot be an operand", name, strings.Join(args, ","))
} }
if len(args) != 1 { if len(args) != 1 {
return "", nil, fmt.Errorf("%s takes 1 arg, got %d", name, len(args)) return "", nil, fmt.Errorf("%s takes 1 arg, got %d", name, len(args))
} }
chain = append(chain, name) chain = append(chain, fn)
arg = args[0] arg = args[0]
} }
} }
@@ -226,14 +158,10 @@ func transformPrep(outer func(string) string) func([]string) callFn {
if err != nil { if err != nil {
panic(fmt.Sprintf("fejkdata: transform arg %q reached prep unvalidated: %v", a[0], err)) panic(fmt.Sprintf("fejkdata: transform arg %q reached prep unvalidated: %v", a[0], err))
} }
fns := make([]func(string) string, len(chain))
for i, name := range chain {
fns[i] = transforms[name]
}
return func(_ *session, _ string, operands []string) string { return func(_ *session, _ string, operands []string) string {
v := operands[0] v := operands[0]
for i := len(fns) - 1; i >= 0; i-- { for i := len(chain) - 1; i >= 0; i-- {
v = fns[i](v) v = chain[i](v)
} }
return outer(v) return outer(v)
} }
+8 -254
View File
@@ -6,7 +6,6 @@ import (
"strconv" "strconv"
"strings" "strings"
"unicode" "unicode"
"unicode/utf8"
) )
// calcNode is a parsed expression node. It evaluates over the operand values expand // calcNode is a parsed expression node. It evaluates over the operand values expand
@@ -200,6 +199,14 @@ func calcPrep(args []string) callFn {
} }
} }
// calcDecimals is a calc's decimals count, or -1 for the shortest form.
func calcDecimals(args []string) int {
if len(args) == 2 {
return atoi(args[1])
}
return -1
}
// indexVars replaces each operand name with its position in the values expand reads. // indexVars replaces each operand name with its position in the values expand reads.
// Both sides take that order from calcVars, so they cannot drift. // Both sides take that order from calcVars, so they cannot drift.
func indexVars(n calcNode, at map[string]int) calcNode { func indexVars(n calcNode, at map[string]int) calcNode {
@@ -384,256 +391,3 @@ func contains(bs []byte, b byte) bool {
} }
return false return false
} }
// calcDecimals is a calc's decimals count, or -1 for the shortest form.
func calcDecimals(args []string) int {
if len(args) == 2 {
return atoi(args[1])
}
return -1
}
// calcLimit is the largest magnitude a proof accepts as finite, far enough below
// math.MaxFloat64 that rounding in the bounds cannot hide an overflow.
const calcLimit = 1e300
// maxOperandLen is the longest operand text a proof bounds by its length, so that
// bound, 10^maxOperandLen, stays within calcLimit.
const maxOperandLen = 300
// calcBound is what a proof knows of every value a calc can take: it lies in [lo, hi],
// is at least nonZero from zero unless nonZero is 0, and is whole when integral.
type calcBound struct {
lo, hi, nonZero float64
integral bool
}
func magnitude(b calcBound) float64 { return math.Max(math.Abs(b.lo), math.Abs(b.hi)) }
// doubt is why a proof could not show a calc finite, and the render that shows it.
type doubt struct{ render, why string }
type bounded struct {
b calcBound
d *doubt
}
// calcProof bounds a typed column's calcs from their operands' renders, to show each
// prints a number rather than NaN or Inf.
type calcProof struct {
decimal *textLanguage
operands map[node]bounded
lengths map[node]int
}
func newCalcProof() *calcProof {
p := &calcProof{operands: map[node]bounded{}, lengths: map[node]int{}}
p.decimal = newTextLanguage(decimalGrammar, p)
return p
}
// call bounds one calc token of t.
func (p *calcProof) call(t *template, args []string) (calcBound, *doubt) {
expr, err := parseCalc(args[0])
if err != nil {
panic(fmt.Sprintf("fejkdata: calc(%q) reached a proof unparsed: %v", args[0], err))
}
b, d := p.expr(expr, t.fields)
if d != nil {
return b, &doubt{d.render, fmt.Sprintf("{calc(%s)}: %s", strings.Join(args, ", "), d.why)}
}
return b, nil
}
func (p *calcProof) expr(n calcNode, fields map[string]node) (calcBound, *doubt) {
switch n := n.(type) {
case calcNum:
v := float64(n)
return calcBound{v, v, v, v == math.Trunc(v)}, nil
case calcVar:
return p.operand(string(n), fields[string(n)])
case calcNeg:
b, d := p.expr(n.x, fields)
return calcBound{-b.hi, -b.lo, b.nonZero, b.integral}, d
case calcBin:
l, d := p.expr(n.l, fields)
if d != nil {
return l, d
}
r, d := p.expr(n.r, fields)
if d != nil {
return r, d
}
return combine(n, l, r)
}
panic(fmt.Sprintf("fejkdata: calc node %T has no bound", n))
}
// combine bounds one operation from the bounds of its sides.
func combine(n calcBin, l, r calcBound) (calcBound, *doubt) {
b := calcBound{integral: l.integral && r.integral}
switch n.op {
case '+':
b.lo, b.hi = l.lo+r.lo, l.hi+r.hi
case '-':
b.lo, b.hi = l.lo-r.hi, l.hi-r.lo
case '*':
b.lo = min(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi)
b.hi = max(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi)
b.nonZero = l.nonZero * r.nonZero
default:
if r.nonZero == 0 {
return b, &doubt{"+Inf", fmt.Sprintf("divides by %s, which can be zero", calcText(n.r))}
}
m := magnitude(l) / r.nonZero
b = calcBound{lo: -m, hi: m, nonZero: l.nonZero / magnitude(r)}
}
if b.lo > 0 || b.hi < 0 {
b.nonZero = math.Max(b.nonZero, math.Min(math.Abs(b.lo), math.Abs(b.hi)))
}
if !(magnitude(b) <= calcLimit) {
return b, &doubt{"+Inf", calcText(n) + " can overflow"}
}
return b, nil
}
// operand bounds a calc operand, once per node.
func (p *calcProof) operand(name string, n node) (calcBound, *doubt) {
if seen, done := p.operands[n]; done {
return seen.b, seen.d
}
b, d := p.measure(name, n)
p.operands[n] = bounded{b, d}
return b, d
}
// measure bounds an operand through the calc it renders when that is all it renders,
// and otherwise from its text: a plain decimal of at most maxOperandLen bytes.
func (p *calcProof) measure(name string, n node) (calcBound, *doubt) {
if t, ok := n.(*template); ok {
if args, isCalc := soleCalc(t); isCalc {
b, d := p.call(t, args)
return rounded(b, calcDecimals(args)), d
}
}
text := p.decimal.node(n, nil)
if w, escapes := text.escape(decimalAccept); escapes {
why := fmt.Sprintf("operand %q can render %s, which is not a plain decimal", name, w)
if w.why != "" {
why += ": " + w.why
}
return calcBound{}, &doubt{"NaN", why}
}
size := p.length(n)
if size > maxOperandLen {
return calcBound{}, &doubt{"NaN", fmt.Sprintf("operand %q can render more than %d bytes, too many to bound", name, maxOperandLen)}
}
ends, m := text.to[1], math.Pow(10, float64(size))
b := calcBound{hi: m, nonZero: 1 / m, integral: ends&decimalFractional == 0}
if ends&decimalNegative != 0 {
b.lo = -m
}
if ends&decimalZero != 0 {
b.nonZero = 0
}
return b, nil
}
// soleCalc reports a template that renders one calc and nothing else, with its args.
func soleCalc(t *template) ([]string, bool) {
if t.repeat != 1 || len(t.ops) != 1 || t.ops[0].kind != 'b' {
return nil, false
}
name, args, _ := funcCall(t.format[1 : len(t.format)-1])
return args, name == "calc"
}
// rounded is b once printed to dp decimals, which moves a value by up to half a unit.
func rounded(b calcBound, dp int) calcBound {
if dp < 0 {
return b
}
half := math.Pow(10, -float64(dp)) / 2
return calcBound{b.lo - half, b.hi + half, math.Max(0, b.nonZero-half), b.integral || dp == 0}
}
// length is the most bytes a render of n can take, anything past maxOperandLen
// reported as maxOperandLen+1.
func (p *calcProof) length(n node) int {
if size, done := p.lengths[n]; done {
return size
}
size := 0
switch n := n.(type) {
case *choice:
for _, it := range n.items {
size = max(size, p.length(it))
}
case *template:
size = p.formatLength(n)*n.repeat + len(n.separator)*(n.repeat-1)
}
size = min(size, maxOperandLen+1)
p.lengths[n] = size
return size
}
func (p *calcProof) formatLength(t *template) int {
size := 0
_ = eachToken(t.format, func(tok ftoken) error {
if tok.kind == 'l' {
size += utf8.RuneLen(tok.r)
} else {
size += p.tokenLength(t, tok.body)
}
size = min(size, maxOperandLen+1)
return nil
})
return size
}
// tokenLength is the most bytes one token can print. A transform never lengthens a
// render that reads as a decimal: it maps each non-ASCII rune, two bytes or more, to at
// most two ASCII letters.
func (p *calcProof) tokenLength(t *template, body string) int {
name, args, isFunc := funcCall(body)
var arms []arm
switch _, isTransform := transforms[name]; {
case !isFunc:
arms = splitArms(body, t.refs)
case isTransform:
leaf, _, _ := unwrapTransform(args[0])
arms = []arm{splitArm(leaf, t.refs)}
case name == "calc":
b, d := p.call(t, args)
if d != nil {
return len(d.render)
}
return shapeLength(printedFloat(b.lo, b.hi, calcDecimals(args), b.integral))
default:
return shapeLength(builtins[name].emits(args))
}
size := 0
for _, a := range arms {
for _, leaf := range pathLeaves(t.fields[a.key], a.tail) {
size = max(size, p.length(leaf))
}
}
return size
}
// shapeLength is the most bytes a shape can emit, anything past maxOperandLen reported
// as maxOperandLen+1.
func shapeLength(s textShape) int {
longest := 0
for _, alt := range s {
size := 0
for _, run := range alt {
if run.max < 0 {
return maxOperandLen + 1
}
size += run.max
}
longest = max(longest, size)
}
return min(longest, maxOperandLen+1)
}
+23 -56
View File
@@ -16,7 +16,10 @@ const (
DataTypeBoolean DataTypeBoolean
) )
var dataTypeNames = [...]string{"string", "integer", "number", "boolean"} var (
dataTypeNames = [...]string{"string", "integer", "number", "boolean"}
dataTypeNouns = [...]string{"text", "an integer", "a number", "a boolean"}
)
// String is the datatype as data spells it. // String is the datatype as data spells it.
func (d DataType) String() string { func (d DataType) String() string {
@@ -32,7 +35,7 @@ type position int
const ( const (
inFormat position = iota // rendered by a format, so neither inFormat position = iota // rendered by a format, so neither
atTop // a category or an inline template, whose fields are the columns atTop // a category or an inline template, whose fields may be columns
inColumn // a column, or a choice item standing in for one inColumn // a column, or a choice item standing in for one
) )
@@ -64,7 +67,7 @@ func datatypeOf(m map[string]any, pos position) (DataType, error) {
// columnDatatype is the datatype a column's items declare. They must agree, since a // columnDatatype is the datatype a column's items declare. They must agree, since a
// column holds one; a column only ever null is a string. // column holds one; a column only ever null is a string.
func columnDatatype(n node) (DataType, error) { func columnDatatype(n node) (DataType, error) {
var declared []DataType var items []*template
var collect func(node) var collect func(node)
collect = func(n node) { collect = func(n node) {
switch n := n.(type) { switch n := n.(type) {
@@ -73,68 +76,32 @@ func columnDatatype(n node) (DataType, error) {
collect(it) collect(it)
} }
case *template: case *template:
declared = append(declared, n.datatype) items = append(items, n)
} }
} }
collect(n) collect(n)
if len(declared) == 0 { if len(items) == 0 {
return DataTypeString, nil return DataTypeString, nil
} }
for _, d := range declared { for _, t := range items[1:] {
if d != declared[0] { if t.datatype != items[0].datatype {
return declared[0], fmt.Errorf("its items declare %s and %s; a column holds one datatype, so give every item the same", declared[0], d) return items[0].datatype, disagreement(items[0], t)
} }
} }
return declared[0], nil return items[0].datatype, nil
} }
// datatypeSpec is what a datatype's text must satisfy: a grammar, the states a render // disagreement names the fix for two items of one column declaring different datatypes.
// may end in, and how an error names the datatype. func disagreement(a, b *template) error {
type datatypeSpec struct { typed, bare := a, b
grammar *grammar if typed.datatype == DataTypeString {
accept uint32 typed, bare = b, a
noun string
} }
switch {
var datatypeSpecs = map[DataType]datatypeSpec{ case bare.datatype != DataTypeString:
DataTypeInteger: {numberGrammar, integerAccept, "an integer"}, return fmt.Errorf("its items declare %s and %s; a column holds one datatype", a.datatype, b.datatype)
DataTypeNumber: {numberGrammar, numberAccept, "a number"}, case len(bare.fields) == 0 && bare.repeat == 1:
DataTypeBoolean: {booleanGrammar, booleanAccept, "a boolean"}, return fmt.Errorf(`item %q declares no datatype, and a column holds one; write it as {"format":%q,"datatype":%q}`, bare.format, bare.format, typed.datatype)
} }
return fmt.Errorf(`an item declares no datatype beside one declaring %s; a column holds one, so give it "datatype": %q`, typed.datatype, typed.datatype)
// datatypeCheck proves every render of a typed column is text its datatype takes. One
// check covers a scope, so a node several columns reach is read once per grammar.
type datatypeCheck struct {
languages map[*grammar]*textLanguage
proof *calcProof
}
func (c *datatypeCheck) check(path string, n node) error {
t, ok := n.(*template)
if !ok || t.datatype == DataTypeString {
return nil
}
spec := datatypeSpecs[t.datatype]
w, escapes := c.language(spec.grammar).node(t, nil).escape(spec.accept)
if !escapes {
return nil
}
msg := fmt.Sprintf("%s: datatype %s, but it can render %s, which is not %s", path, t.datatype, w, spec.noun)
if w.why != "" {
msg += ": " + w.why
}
return errors.New(msg)
}
func (c *datatypeCheck) language(g *grammar) *textLanguage {
if c.proof == nil {
c.proof = newCalcProof()
c.languages = map[*grammar]*textLanguage{}
}
l, made := c.languages[g]
if !made {
l = newTextLanguage(g, c.proof)
c.languages[g] = l
}
return l
} }
+1 -1
View File
@@ -189,7 +189,7 @@ func checkScope(s nodeScope) error {
if err := s(heldCheck); err != nil { if err := s(heldCheck); err != nil {
return err return err
} }
return s((&datatypeCheck{}).check) return s((&valueProof{}).checkDatatype)
} }
type reachMemo map[node]int type reachMemo map[node]int
+1 -1
View File
@@ -229,7 +229,7 @@ func compileTemplate(m map[string]any, pos position) (node, error) {
return nil, err return nil, err
} }
fieldPos := inFormat fieldPos := inFormat
if pos == atTop && o.repeat == 1 { if pos == atTop && projectsColumns(o.repeat) {
fieldPos = inColumn fieldPos = inColumn
} }
fields, err := compileFields(m, fieldPos) fields, err := compileFields(m, fieldPos)
+8 -6
View File
@@ -63,16 +63,14 @@ func (r *Record) CSVHeader() string {
} }
// CSVLine renders the column values as one CSV row: a null column an empty field and an // CSVLine renders the column values as one CSV row: a null column an empty field and an
// empty string "", the convention PostgreSQL's COPY reads a null by. // empty string "", the convention PostgreSQL's COPY reads a null by. A record of one null
// column is a blank line, which COPY reads as null and most CSV readers skip.
func (r *Record) CSVLine() string { func (r *Record) CSVLine() string {
fields := make([]string, len(r.columns)) fields := make([]string, len(r.columns))
for i, c := range r.columns { for i, c := range r.columns {
fields[i] = literal(c, csvField, "") fields[i] = literal(c, csvField, "")
} }
if line := strings.Join(fields, ","); line != "" { return strings.Join(fields, ",")
return line
}
return `""` // a blank line is a row every CSV reader drops
} }
func csvField(s string) string { func csvField(s string) string {
@@ -207,7 +205,7 @@ func recordOf(n node) (*template, []Column, error) {
if !ok { if !ok {
return nil, nil, errors.New("names a choice, not a template; a record is a template whose fields are its columns") return nil, nil, errors.New("names a choice, not a template; a record is a template whose fields are its columns")
} }
if t.repeat != 1 { if !projectsColumns(t.repeat) {
return nil, nil, fmt.Errorf("carries repeat %d, which composes its format into one string; a record projects columns instead — drop the repeat and render the record again for more rows", t.repeat) return nil, nil, fmt.Errorf("carries repeat %d, which composes its format into one string; a record projects columns instead — drop the repeat and render the record again for more rows", t.repeat)
} }
names := recordColumns(t) names := recordColumns(t)
@@ -225,6 +223,10 @@ func recordOf(n node) (*template, []Column, error) {
return t, columns, nil return t, columns, nil
} }
// projectsColumns reports whether a category or inline template with this repeat is a
// record, its fields the columns; a repeat composes the format into one string instead.
func projectsColumns(repeat int) bool { return repeat == 1 }
// checkColumnRefs rejects the reference reads a record's shared draw cannot answer // checkColumnRefs rejects the reference reads a record's shared draw cannot answer
// for: one column rendering a level another reads a path into, and a column // for: one column rendering a level another reads a path into, and a column
// reading the record back through its own path. // reading the record back through its own path.
-403
View File
@@ -1,403 +0,0 @@
package fejkdata
import (
"slices"
"strconv"
"strings"
"unicode/utf8"
)
// grammar is a deterministic automaton over a scalar's text: state 0 is dead, 1 the
// start, and each state lists the runes that leave it and where they lead.
type grammar [][]arc
type arc struct {
on string
to int
}
func (g *grammar) run(q int, s string) int {
for _, r := range s {
if q = g.step(q, r); q == 0 {
return 0
}
}
return q
}
func (g *grammar) step(q int, r rune) int {
for _, a := range (*g)[q] {
if strings.ContainsRune(a.on, r) {
return a.to
}
}
return 0
}
const (
decimalDigits = "0123456789"
nonZeroDigits = "123456789"
)
// numberGrammar reads a JSON number. States: 2 "-", 3 "0", 4 more integer digits, 5 ".",
// 6 fraction digits, 7 "e", 8 its sign, 9 exponent digits.
var numberGrammar = &grammar{
nil,
{{"-", 2}, {"0", 3}, {nonZeroDigits, 4}},
{{"0", 3}, {nonZeroDigits, 4}},
{{".", 5}, {"eE", 7}},
{{decimalDigits, 4}, {".", 5}, {"eE", 7}},
{{decimalDigits, 6}},
{{decimalDigits, 6}, {"eE", 7}},
{{"+-", 8}, {decimalDigits, 9}},
{{decimalDigits, 9}},
{{decimalDigits, 9}},
}
const (
integerAccept uint32 = 1<<3 | 1<<4
numberAccept = integerAccept | 1<<6 | 1<<9
)
var booleanGrammar = &grammar{
nil,
{{"t", 2}, {"f", 6}},
{{"r", 3}}, {{"u", 4}}, {{"e", 5}}, nil,
{{"a", 7}}, {{"l", 8}}, {{"s", 9}}, {{"e", 10}}, nil,
}
const booleanAccept uint32 = 1<<5 | 1<<10
// decimalGrammar reads what a calc operand must render to be proven finite: a sign,
// digits and at most one dot. Past the sign, states 4–9 are positive and 10–15 their
// negatives: 4 zero digits, 5 a nonzero integer, 6 a leading dot, 7 zero with a dot,
// 8 a nonzero integer with a zero fraction, 9 a nonzero fraction.
var decimalGrammar = &grammar{
nil,
{{"+", 2}, {"-", 3}, {"0", 4}, {nonZeroDigits, 5}, {".", 6}},
{{"0", 4}, {nonZeroDigits, 5}, {".", 6}},
{{"0", 10}, {nonZeroDigits, 11}, {".", 12}},
{{"0", 4}, {nonZeroDigits, 5}, {".", 7}},
{{decimalDigits, 5}, {".", 8}},
{{"0", 7}, {nonZeroDigits, 9}},
{{"0", 7}, {nonZeroDigits, 9}},
{{"0", 8}, {nonZeroDigits, 9}},
{{decimalDigits, 9}},
{{"0", 10}, {nonZeroDigits, 11}, {".", 13}},
{{decimalDigits, 11}, {".", 14}},
{{"0", 13}, {nonZeroDigits, 15}},
{{"0", 13}, {nonZeroDigits, 15}},
{{"0", 14}, {nonZeroDigits, 15}},
{{decimalDigits, 15}},
}
const (
decimalAccept uint32 = 1<<4 | 1<<5 | 1<<7 | 1<<8 | 1<<9 | 1<<10 | 1<<11 | 1<<13 | 1<<14 | 1<<15
decimalNegative uint32 = 0xfc00
decimalZero uint32 = 1<<4 | 1<<7 | 1<<10 | 1<<13
decimalFractional uint32 = 1<<9 | 1<<15
)
// relation is what a node's renders do to a grammar: from each state, the states a
// render can end in, and one render reaching each.
type relation struct {
g *grammar
to []uint32
w []witness // w[from*len(to)+to]
}
// witness is one render, cut past witnessCap bytes, and why it can occur when the text
// alone does not say.
type witness struct {
text string
cut bool
why string
}
const witnessCap = 60
func (w witness) then(next witness) witness {
if w.why == "" {
w.why = next.why
}
if w.cut {
return w
}
w.text += next.text
w.cut = next.cut
if len(w.text) > witnessCap {
end := witnessCap
for !utf8.RuneStart(w.text[end]) {
end--
}
w.text, w.cut = w.text[:end], true
}
return w
}
func (w witness) String() string {
if w.cut {
return strconv.Quote(w.text + "…")
}
return strconv.Quote(w.text)
}
func newRelation(g *grammar) *relation {
n := len(*g)
return &relation{g: g, to: make([]uint32, n), w: make([]witness, n*n)}
}
func (r *relation) add(from, to int, w witness) {
if r.to[from]&(1<<to) == 0 {
r.to[from] |= 1 << to
r.w[from*len(r.to)+to] = w
}
}
// textRelation is the relation of a render that is always s.
func textRelation(g *grammar, s, why string) *relation {
r := newRelation(g)
w := witness{why: why}.then(witness{text: s})
for q := range r.to {
r.add(q, g.run(q, s), w)
}
return r
}
// union is the renders of either relation; a nil relation has none.
func union(a, b *relation) *relation {
if a == nil {
return b
}
if b == nil {
return a
}
u := newRelation(a.g)
for _, r := range []*relation{a, b} {
for from, ends := range r.to {
for to := range r.to {
if ends&(1<<to) != 0 {
u.add(from, to, r.w[from*len(r.to)+to])
}
}
}
}
return u
}
// then is a render of r followed by a render of next.
func (r *relation) then(next *relation) *relation {
c := newRelation(r.g)
n := len(r.to)
for from, mids := range r.to {
for mid := 0; mid < n; mid++ {
if mids&(1<<mid) == 0 {
continue
}
for to := 0; to < n; to++ {
if next.to[mid]&(1<<to) != 0 && c.to[from]&(1<<to) == 0 {
c.add(from, to, r.w[from*n+mid].then(next.w[mid*n+to]))
}
}
}
}
return c
}
// power is k renders of r in a row, k at least 1, composed by squaring.
func (r *relation) power(k int) *relation {
var out *relation
for base := r; ; base = base.then(base) {
if k&1 == 1 {
if out == nil {
out = base
} else {
out = out.then(base)
}
}
if k >>= 1; k == 0 {
return out
}
}
}
// closure is any number of renders of r in a row, where r includes the empty render.
func (r *relation) closure() *relation {
for {
next := r.then(r)
if slices.Equal(next.to, r.to) {
return r
}
r = next
}
}
// escape finds a render from the start that ends outside accept, preferring one that
// carries a reason.
func (r *relation) escape(accept uint32) (witness, bool) {
var found witness
escapes := false
for to := range r.to {
if (r.to[1]&^accept)&(1<<to) == 0 {
continue
}
if w := r.w[len(r.to)+to]; !escapes || found.why == "" && w.why != "" {
found, escapes = w, true
}
}
return found, escapes
}
// textShape is the text a builtin can emit: alternatives, each a sequence of runs.
type textShape [][]charRun
// charRun is between min and max characters, each one of chars; max -1 is unbounded.
// chars is ASCII, so a run of k characters is k bytes.
type charRun struct {
chars string
min, max int
}
// textLanguage is what one grammar makes of the renders a check reads, worked out once
// per node and fold.
type textLanguage struct {
g *grammar
proof *calcProof
memo map[languageKey]*relation
empty *relation
}
type languageKey struct {
n node
fold string
}
// fold is the transforms a render passes through before the grammar reads it, innermost
// first. Each rewrites rune by rune, so folding a render is folding each of its pieces.
type fold []string
func (f fold) apply(s string) string {
for _, name := range f {
s = transforms[name](s)
}
return s
}
func newTextLanguage(g *grammar, proof *calcProof) *textLanguage {
return &textLanguage{g: g, proof: proof, memo: map[languageKey]*relation{}, empty: textRelation(g, "", "")}
}
func (l *textLanguage) node(n node, f fold) *relation {
key := languageKey{n, strings.Join(f, ",")}
if r, done := l.memo[key]; done {
return r
}
r := l.empty // a null renders ""
switch n := n.(type) {
case *choice:
r = nil
for _, it := range n.items {
r = union(r, l.node(it, f))
}
case *template:
r = l.format(n, f)
if n.repeat > 1 {
r = r.then(l.text(f.apply(n.separator)).then(r).power(n.repeat - 1))
}
}
l.memo[key] = r
return r
}
func (l *textLanguage) text(s string) *relation { return textRelation(l.g, s, "") }
// format reads a template's format the way expand renders it: literal runs and tokens
// in turn.
func (l *textLanguage) format(t *template, f fold) *relation {
r := l.empty
var lit strings.Builder
_ = eachToken(t.format, func(tok ftoken) error {
if tok.kind == 'l' {
lit.WriteRune(tok.r)
return nil
}
r = r.then(l.text(f.apply(lit.String()))).then(l.token(t, tok.body, f))
lit.Reset()
return nil
})
return r.then(l.text(f.apply(lit.String())))
}
// token reads one {…} token: a field read, a transform over one, a calc, or what a
// builtin emits.
func (l *textLanguage) token(t *template, body string, f fold) *relation {
name, args, isFunc := funcCall(body)
if !isFunc {
var r *relation
for _, a := range splitArms(body, t.refs) {
r = union(r, l.read(t, a, f))
}
return r
}
if _, isTransform := transforms[name]; isTransform {
leaf, chain, _ := unwrapTransform(args[0])
inner := slices.Clone(chain)
slices.Reverse(inner)
return l.read(t, splitArm(leaf, t.refs), append(append(inner, name), f...))
}
if name == "calc" {
return l.calc(t, args, f)
}
return l.shape(builtins[name].emits(args), f)
}
// read is one arm of a token: every node its path can land on.
func (l *textLanguage) read(t *template, a arm, f fold) *relation {
var r *relation
for _, leaf := range pathLeaves(t.fields[a.key], a.tail) {
r = union(r, l.node(leaf, f))
}
return r
}
func (l *textLanguage) calc(t *template, args []string, f fold) *relation {
b, d := l.proof.call(t, args)
if d != nil {
return textRelation(l.g, f.apply(d.render), d.why)
}
return l.shape(printedFloat(b.lo, b.hi, calcDecimals(args), b.integral), f)
}
func (l *textLanguage) shape(s textShape, f fold) *relation {
var r *relation
for _, alt := range s {
seq := l.empty
for _, run := range alt {
seq = seq.then(l.run(run, f))
}
r = union(r, seq)
}
return r
}
// run reads a charRun: min characters, then up to max-min more.
func (l *textLanguage) run(c charRun, f fold) *relation {
one := newRelation(l.g)
for from := range one.to {
for _, ch := range c.chars {
s := f.apply(string(ch))
one.add(from, l.g.run(from, s), witness{text: s})
}
}
more := l.empty
switch optional := union(one, l.empty); {
case c.max < 0:
more = optional.closure()
case c.max > c.min:
more = optional.power(c.max - c.min)
}
if c.min == 0 {
return more
}
return one.power(c.min).then(more)
}
+3 -3
View File
@@ -72,9 +72,9 @@ type builtin struct {
// operands names the fields the call reads, which expand renders for it; nil // operands names the fields the call reads, which expand renders for it; nil
// for a builtin that reads none. // for a builtin that reads none.
operands func(args []string) []string operands func(args []string) []string
// emits is the text a call can print, for the datatype check; nil for calc and the // number bounds the number a call prints and names the datatype its text is; nil
// transforms, whose text the check derives from what they read. // for a builtin that prints text.
emits func(args []string) textShape number func(args []string) (proven, DataType)
} }
// funcCall splits a "{token}" body shaped name(args) into its parts; ok is false // funcCall splits a "{token}" body shaped name(args) into its parts; ok is false
+4
View File
@@ -22,6 +22,10 @@ The record API lands first, so the data update can use it.
- `code` and `symbol` sibling fields reading `currency`, as `{code} {symbol}` → a matching pair - `code` and `symbol` sibling fields reading `currency`, as `{code} {symbol}` → a matching pair
- `{a} & {b}`, each reading `person` → one person, or two when `a` and `b` name different groups - `{a} & {b}`, each reading `person` → one person, or two when `a` and `b` name different groups
- two bare `{/sv_SE.word}` → two words - two bare `{/sv_SE.word}` → two words
- Reference inheritance — settle whether a column that is exactly one reference to
another record's column, like `{/src.score}`, takes that column's datatype and
null. Today a null there writes `""`, and a typed column reading it is refused.
Settle before draw groups and the data update.
### Data ### Data
+264
View File
@@ -0,0 +1,264 @@
package fejkdata
import (
"fmt"
"math"
"regexp"
"strconv"
"strings"
)
// proven is what a proof knows of every render of a node: bounds on the number each
// reads as, and per datatype why some render's text is not one ("" when none).
type proven struct {
lo, hi float64
nonZero float64 // every value is at least this far from zero; 0 when one can be zero
integral bool
notNumber string // why some render reads as no finite number, the way calc reads it
not [len(dataTypeNames)]string
}
// valueProof proves what typed columns and their calc operands hold, each node once per
// scope. A typed column holds one value: a literal, one value builtin, one calc, or a
// read of such values.
type valueProof struct {
memo map[node]proven
}
// checkDatatype rejects a typed column some render of which is not text of its datatype.
func (p *valueProof) checkDatatype(path string, n node) error {
t, ok := n.(*template)
if !ok || t.datatype == DataTypeString {
return nil
}
if err := p.prove(t, t.datatype); err != nil {
return fmt.Errorf("%s: %w", path, err)
}
return nil
}
// prove reports why some render of n is not text of datatype d.
func (p *valueProof) prove(n node, d DataType) error {
if reason := p.of(n).not[d]; reason != "" {
return fmt.Errorf("datatype %s: %s", d, reason)
}
return nil
}
func (p *valueProof) of(n node) proven {
if v, done := p.memo[n]; done {
return v
}
if p.memo == nil {
p.memo = map[node]proven{}
}
var v proven
switch n := n.(type) {
case *choice:
v = p.unite(n.items)
case *template:
v = p.template(n)
default:
v = unproven(`it reads a null, which renders "" outside its own column`)
}
p.memo[n] = v
return v
}
func (p *valueProof) unite(nodes []node) proven {
v := p.of(nodes[0])
for _, n := range nodes[1:] {
w := p.of(n)
v.lo, v.hi, v.nonZero = min(v.lo, w.lo), max(v.hi, w.hi), min(v.nonZero, w.nonZero)
v.integral = v.integral && w.integral
if v.notNumber == "" {
v.notNumber = w.notNumber
}
for d := range v.not {
if v.not[d] == "" {
v.not[d] = w.not[d]
}
}
}
return v
}
// template proves a template that renders one value: fixed text, or a format that is
// one token alone.
func (p *valueProof) template(t *template) proven {
switch {
case t.repeat != 1:
return unproven(fmt.Sprintf("%q carries a repeat, which composes text rather than one value", t.format))
case t.fixed:
return literalValue(t.lit)
case len(t.ops) != 1:
return unproven(fmt.Sprintf("%q is not one value; write one literal or one {int()}, {float()}, {seq()} or {calc()}, or read one", t.format))
}
body := t.format[1 : len(t.format)-1]
name, args, isFunc := funcCall(body)
switch _, isTransform := transforms[name]; {
case !isFunc:
var leaves []node
for _, a := range splitArms(body, t.refs) {
leaves = append(leaves, pathLeaves(t.fields[a.key], a.tail)...)
}
return p.unite(leaves)
case name == "calc":
return p.calc(t, body, args)
case builtins[name].number != nil:
v, prints := builtins[name].number(args)
return printing(body, prints, v)
case isTransform:
return unproven(fmt.Sprintf("{%s} rewrites text rather than printing a value; write the values it would print", body))
}
return printing(body, DataTypeString, proven{notNumber: fmt.Sprintf("{%s} prints text, not a number", body)})
}
func (p *valueProof) calc(t *template, body string, args []string) proven {
expr, err := parseCalc(args[0])
if err != nil {
panic(fmt.Sprintf("fejkdata: calc(%q) reached a proof unparsed: %v", args[0], err))
}
v, doubt := p.expr(expr, t.fields)
if doubt == "" && !(magnitude(v) <= calcLimit) {
doubt = calcText(expr) + " is not proven within 1e300"
}
if doubt != "" {
return unproven(fmt.Sprintf("{%s}: %s", body, doubt))
}
v, prints := printedNumber(v, calcDecimals(args))
return printing(body, prints, v)
}
// calcLimit is the largest magnitude a proof accepts as finite, far enough below
// math.MaxFloat64 that rounding in the bounds cannot hide an overflow.
const calcLimit = 1e300
// expr bounds a calc expression from its operands, or says why it cannot.
func (p *valueProof) expr(n calcNode, fields map[string]node) (proven, string) {
switch n := n.(type) {
case calcNum:
v := float64(n)
return bounded(v, v, v == math.Trunc(v)), ""
case calcVar:
v := p.of(fields[string(n)])
if v.notNumber != "" {
return proven{}, fmt.Sprintf("operand %q: %s", string(n), v.notNumber)
}
return proven{lo: v.lo, hi: v.hi, nonZero: v.nonZero, integral: v.integral}, ""
case calcNeg:
v, doubt := p.expr(n.x, fields)
v.lo, v.hi = -v.hi, -v.lo
return v, doubt
case calcBin:
l, doubt := p.expr(n.l, fields)
if doubt != "" {
return l, doubt
}
r, doubt := p.expr(n.r, fields)
if doubt != "" {
return r, doubt
}
return combine(n, l, r)
}
panic(fmt.Sprintf("fejkdata: calc node %T has no bound", n))
}
// combine bounds one operation from the bounds of its sides.
func combine(n calcBin, l, r proven) (proven, string) {
var v proven
integral := l.integral && r.integral
switch n.op {
case '+':
v = bounded(l.lo+r.lo, l.hi+r.hi, integral)
case '-':
v = bounded(l.lo-r.hi, l.hi-r.lo, integral)
case '*':
v = bounded(min(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi), max(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi), integral)
v.nonZero = max(v.nonZero, l.nonZero*r.nonZero)
default:
if r.nonZero == 0 {
return v, fmt.Sprintf("divides by %s, which is not proven nonzero", calcText(n.r))
}
m := magnitude(l) / r.nonZero
v = proven{lo: -m, hi: m, nonZero: l.nonZero / magnitude(r)}
}
if !(magnitude(v) <= calcLimit) {
return v, calcText(n) + " is not proven within 1e300"
}
return v, ""
}
// bounded is a number in [lo, hi], its distance from zero read off the bounds.
func bounded(lo, hi float64, integral bool) proven {
v := proven{lo: lo, hi: hi, integral: integral}
switch {
case lo > 0:
v.nonZero = lo
case hi < 0:
v.nonZero = -hi
}
return v
}
func magnitude(v proven) float64 { return math.Max(math.Abs(v.lo), math.Abs(v.hi)) }
// printedNumber is v once strconv.FormatFloat prints it to dp decimals, and the datatype
// that text is: an integer when whole and within int64, else a number.
func printedNumber(v proven, dp int) (proven, DataType) {
if dp >= 0 {
half := math.Pow(10, -float64(dp)) / 2
v = proven{lo: v.lo - half, hi: v.hi + half, nonZero: math.Max(0, v.nonZero-half), integral: v.integral || dp == 0}
}
if (dp == 0 || dp < 0 && v.integral) && magnitude(v) < math.MaxInt64 {
return v, DataTypeInteger
}
return v, DataTypeNumber
}
// printing is v for a token whose every render is text of datatype prints, with a reason
// against each datatype that text is not.
func printing(token string, prints DataType, v proven) proven {
for d := DataTypeInteger; d <= DataTypeBoolean; d++ {
if prints != d && !(prints == DataTypeInteger && d == DataTypeNumber) {
v.not[d] = fmt.Sprintf("{%s} prints %s, not %s", token, dataTypeNouns[prints], dataTypeNouns[d])
}
}
return v
}
// unproven is a render no datatype and no calc can take, for why.
func unproven(why string) proven {
v := proven{notNumber: why}
for d := DataTypeInteger; d <= DataTypeBoolean; d++ {
v.not[d] = why
}
return v
}
var (
integerText = regexp.MustCompile(`^-?(0|[1-9][0-9]*)$`)
numberText = regexp.MustCompile(`^-?(0|[1-9][0-9]*)(\.[0-9]+)?([eE][+-]?[0-9]+)?$`)
)
// literalValue proves fixed text: the number calc reads it as, and each datatype it is.
func literalValue(text string) proven {
var v proven
if f, err := strconv.ParseFloat(strings.TrimSpace(text), 64); err != nil || math.IsNaN(f) || math.IsInf(f, 0) {
v.notNumber = fmt.Sprintf("%q is not a number", text)
} else {
v = bounded(f, f, f == math.Trunc(f))
}
if _, err := strconv.ParseInt(text, 10, 64); !integerText.MatchString(text) {
v.not[DataTypeInteger] = fmt.Sprintf("%q is not an integer", text)
} else if err != nil {
v.not[DataTypeInteger] = fmt.Sprintf("%q is past the int64 range", text)
}
if v.notNumber != "" || !numberText.MatchString(text) {
v.not[DataTypeNumber] = fmt.Sprintf("%q is not a number", text)
}
if text != "true" && text != "false" {
v.not[DataTypeBoolean] = fmt.Sprintf("%q is not a boolean", text)
}
return v
}