Check a typed column's value, not its text: one literal, value call, calc or read, bounded by its arguments
Tests / vet + fmt + tests (pull_request) Successful in 59s

This commit is contained in:
2026-09-15 14:33:04 +02:00
parent 770dca507a
commit 5615ab6d89
11 changed files with 379 additions and 844 deletions
+38 -19
View File
@@ -271,25 +271,31 @@ string:
Writes e.g. `{"id":1,"paid":true,"total":59.97}`. A column is a field of the top-level
template, or an item of a choice standing in for one; `datatype` anywhere else is a
load error. So is a column that can render text its datatype rejects — `integer` takes
`-?(0|[1-9][0-9]*)`, `number` a JSON number, `boolean` `true` or `false` — and the
error shows such a render:
load error. A typed column holds one value, alone in its format: a literal, one
`{int()}`, `{float()}`, `{seq()}` or `{calc()}` call, or a read that lands only on such
values. `integer` is an int64 written `-?(0|[1-9][0-9]*)` — `{float()}` prints one at
`0` decimals — `number` a JSON number, `boolean` `true` or `false`. A value its
datatype cannot hold is a load error naming it:
```text
order.id: datatype integer, but it can render "000", which is not an integer
order.id: datatype integer: {digits(3)} prints text, not an integer
order.id: datatype integer: "1{digits(2)}" is not one value; write one literal or one {int()}, {float()}, {seq()} or {calc()}, or read one
```
A `{calc()}` fills an `integer` or `number` column only where it provably prints no
`NaN` or `Inf`: each operand is a plain decimal — a sign, digits, one dot — of at most
300 bytes, or a field holding only such a calc, and no divisor can be zero. An
`integer` column also needs a decimals count of `0`, or integer operands and no `/`.
A typed column's `{calc()}` must be proven to print a number: each operand a number
literal, an `{int()}`, `{float()}`, `{seq()}` or `{digits()}` call, a calc, or a read of
such values, whose bounds keep every divisor from zero and the result within `1e300`.
What the bounds cannot show is refused — `{calc(a / b)}: divides by b, which is not
proven nonzero`. The calc fills an `integer` column at `0` decimals, or over whole
operands with no `/`.
### Null
A `null` item draws a record column as null: `json` writes `null`, `sql` `NULL`, and
`csv` an empty field, with an empty string written `""` so PostgreSQL's `COPY … CSV`
reads both back. `Fake` renders a null as `""`. The other items' weights skew its
odds:
`csv` an empty field, with an empty string written `""` — the convention PostgreSQL's
`COPY … CSV` reads. A record of one null column is a blank line, which `COPY` reads as
null but most CSV readers skip, so write such a record as `json` or `sql`. `Fake`
renders a null as `""`. The other items' weights skew its odds:
```json
{ "format": "", "deleted_at": null, "middle": [null, { "format": "{n}", "n": ["Ann", "Eva"], "weight": 3 }] }
@@ -363,7 +369,7 @@ Renders e.g. `19.99 x 3 = 59.97`. An operand that can never be a number (`"abc"`
or a choice of such) is rejected at load, as is a division by a constant zero
(`1/0`, or a fixed `"0"` field); an operand that sometimes is not a number yields
`NaN`, and a division by one that is not constant `Inf` — both print rather than
fail.
fail, except in a [typed column](#datatype), which must prove neither happens.
### Transforms
@@ -556,11 +562,13 @@ tokens add cost in proportion to the output.
- **64-bit targets only.** The gate builds amd64, and the buffer sizing a render
pre-computes (renders × bytes) assumes a 64-bit int; on a 32-bit target it could
overflow and panic.
- **A constant zero divisor is a load error; a divisor that is not constant prints
`Inf`.** `1/0` and a fixed `"0"` field are decidable, so they join the
never-numeric operand as a load error; the fold stops where an operand varies,
- **A constant zero divisor is a load error; in a string column a divisor that is not
constant prints `Inf`.** `1/0` and a fixed `"0"` field are decidable, so they join
the never-numeric operand as a load error; the fold stops where an operand varies,
so `a/(b*c)` with `b` fixed at `0` and `c` varying loads and prints `Inf` every
draw — catching it needs zero-absorbing algebra for a shape nobody writes.
draw — catching it needs zero-absorbing algebra for a shape nobody writes. A
[typed column](#datatype) bounds its operands instead and refuses a divisor it
cannot keep from zero.
- **In data, a default written out and a constant spelled as a sample are load
errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `datatype: "string"`,
`int(5,5)`, `float(1,1,2)`,
@@ -605,6 +613,17 @@ tokens add cost in proportion to the output.
- **Null is a `null` item, not a rate.** A null is one more outcome of a column's
draw, so a choice's weights skew it like any other; a null-rate option would be a
second way to state odds.
- **A typed column holds one value, not composed text.** Its bounds come from a
literal or a call's arguments, so a load error names a real value, a range check is
one comparison, and `1{digits(2)}` is a second spelling of `{int(100,199)}`.
- **A typed column's calc is refused unless proven.** Operand bounds must keep each
divisor from zero and the result finite; what they cannot show is refused rather
than trusted, since a bare `NaN` breaks the JSON and SQL it lands in.
- **`Column` carries text, not a Go value.** `Value` is the rendered string beside
`DataType` and `Null`, which each serializer writes as the load check proved it; a
`Value any` would hand every caller a type switch.
- **The package stays flat.** Go ties a package to one directory, so folders would
split the API into packages.
- **The performance gate asserts allocations, not wall-clock time.** `AllocsPerRun`
is deterministic across machines, so a ±10% ceiling does not flake under CI load,
while time varies with the machine and its neighbours. A rendering slowdown
@@ -660,9 +679,9 @@ hold.go the hold: one draw per expansion for paths and operands, and its
reference.go reference sigils, and binding references across the tree
graph.go the render graph: edges, cycles, the repeat bound, tree walks
builtins.go the {name()} function registry and its implementations
calc.go the {calc()} arithmetic evaluator: parser, eval, validation, and the proof a typed column's calc is finite
datatype.go column datatypes: DataType, where datatype and null may sit, and the load check every typed render passes
renderlang.go what text a node can render, as relations over a scalar's grammar
calc.go the {calc()} arithmetic evaluator: parser, eval, validation
datatype.go column datatypes: DataType, where datatype and null may sit, a column's datatype
value.go the value proof: what a typed column or calc operand holds, checked at load
data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge
cmd/fejkdata/ the fejkdata CLI
data/ shipped data (JSON), embedded at build: locale folders + a misc folder
+29 -101
View File
@@ -5,7 +5,6 @@ import (
"errors"
"fmt"
"math"
"slices"
"strconv"
"strings"
"unicode"
@@ -25,36 +24,42 @@ const (
// samples read only the rng. A time-based id (uuid v7, ulid) draws its timestamp
// from the rng, not the wall clock, so seeded output stays reproducible.
var builtins = map[string]builtin{
"luhn": {arity: 0, prep: derive(func(e string) string { return string(rune('0' + luhnCheck(e))) }), emits: always(textShape{{{decimalDigits, 1, 1}}})},
"mod11": {arity: 0, prep: derive(mod11Check), emits: always(textShape{{{decimalDigits + "X", 1, 1}}})},
"ean": {arity: 0, prep: derive(eanCheck), emits: always(textShape{{{decimalDigits, 1, 1}}})},
"uuid": {arity: 0, prep: sample(uuidV7), emits: always(uuidShape)},
"ulid": {arity: 0, prep: sample(ulid), emits: always(textShape{{{crockford[:8], 1, 1}, {crockford, 25, 25}}})},
"nanoid": sampleOf(nanoidAlphabet),
"hex": sampleOf(hexDigits),
"digits": sampleOf(decimalDigits),
"upper": sampleOf("ABCDEFGHIJKLMNOPQRSTUVWXYZ"),
"lower": sampleOf("abcdefghijklmnopqrstuvwxyz"),
"luhn": {arity: 0, prep: derive(func(e string) string { return string(rune('0' + luhnCheck(e))) })},
"mod11": {arity: 0, prep: derive(mod11Check)},
"ean": {arity: 0, prep: derive(eanCheck)},
"uuid": {arity: 0, prep: sample(uuidV7)},
"ulid": {arity: 0, prep: sample(ulid)},
"nanoid": {arity: 1, check: posIntArg, prep: chars(nanoidAlphabet)},
"hex": {arity: 1, check: posIntArg, prep: chars(hexDigits)},
"digits": {arity: 1, check: posIntArg, prep: chars("0123456789"), number: func(a []string) (proven, DataType) {
return bounded(0, math.Pow(10, float64(atoi(a[0])))-1, true), DataTypeString
}},
"upper": {arity: 1, check: posIntArg, prep: chars("ABCDEFGHIJKLMNOPQRSTUVWXYZ")},
"lower": {arity: 1, check: posIntArg, prep: chars("abcdefghijklmnopqrstuvwxyz")},
"base64": {arity: 1, check: posIntArg, prep: func(a []string) callFn {
n := atoi(a[0])
return func(s *session, _ string, _ []string) string {
return base64.StdEncoding.EncodeToString(randBytes(s, n))
}
}, emits: base64Shape},
}},
"int": {arity: 2, check: intRangeArgs, prep: func(a []string) callFn {
lo, span := atoi(a[0]), atoi(a[1])-atoi(a[0])+1
return func(s *session, _ string, _ []string) string { return strconv.Itoa(lo + s.IntN(span)) }
}, emits: intShape},
}, number: func(a []string) (proven, DataType) {
return bounded(float64(atoi(a[0])), float64(atoi(a[1])), true), DataTypeInteger
}},
"float": {arity: 3, check: floatArgs, prep: func(a []string) callFn {
lo, hi, dp := atof(a[0]), atof(a[1]), atoi(a[2])
return func(s *session, _ string, _ []string) string {
return strconv.FormatFloat(lo+s.Float64()*(hi-lo), 'f', dp, 64)
}
}, emits: func(a []string) textShape { return printedFloat(atof(a[0]), atof(a[1]), atoi(a[2]), false) }},
}, number: func(a []string) (proven, DataType) {
return printedNumber(bounded(atof(a[0]), atof(a[1]), false), atoi(a[2]))
}},
"iban": {arity: 1, check: ibanArg, prep: func(a []string) callFn {
cc := a[0]
return func(s *session, _ string, _ []string) string { return iban(s, cc) }
}, emits: ibanShape},
}},
"calc": {arity: -1, check: checkCalc, prep: calcPrep, operands: calcOperands},
"lowercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToLower), operands: transformOperand},
"uppercase": {arity: 1, check: transformArg, prep: transformPrep(strings.ToUpper), operands: transformOperand},
@@ -70,7 +75,9 @@ var builtins = map[string]builtin{
return func(s *session, _ string, _ []string) string {
return strconv.FormatUint(s.next(key), 10)
}
}, emits: always(textShape{{{nonZeroDigits, 1, 1}, {decimalDigits, 0, 19}}})},
}, number: func([]string) (proven, DataType) {
return bounded(1, math.MaxInt64, true), DataTypeInteger
}},
}
// derive and sample are the two argument-free builtin shapes: a derivation reads
@@ -95,82 +102,6 @@ func chars(alphabet string) func([]string) callFn {
}
}
// sampleOf is the builtin that draws n characters from an alphabet.
func sampleOf(alphabet string) builtin {
return builtin{arity: 1, check: posIntArg, prep: chars(alphabet), emits: func(a []string) textShape {
n := atoi(a[0])
return textShape{{{alphabet, n, n}}}
}}
}
// always is the emits of a builtin whose args do not change what it can print.
func always(s textShape) func([]string) textShape {
return func([]string) textShape { return s }
}
var uuidShape = textShape{{{hexDigits, 8, 8}, {"-", 1, 1}, {hexDigits, 4, 4}, {"-", 1, 1}, {"7", 1, 1}, {hexDigits, 3, 3}, {"-", 1, 1}, {"89ab", 1, 1}, {hexDigits, 3, 3}, {"-", 1, 1}, {hexDigits, 12, 12}}}
const base64Alphabet = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/"
func base64Shape(a []string) textShape {
n := atoi(a[0])
pad := (3 - n%3) % 3
size := 4*((n+2)/3) - pad
return textShape{{{base64Alphabet, size, size}, {"=", pad, pad}}}
}
// intShape is what int prints: a sign only below zero, and no leading zero.
func intShape(a []string) textShape {
lo, hi := atoi(a[0]), atoi(a[1])
var s textShape
if lo <= 0 && hi >= 0 {
s = append(s, []charRun{{"0", 1, 1}})
}
if hi > 0 {
s = append(s, []charRun{{nonZeroDigits, 1, 1}, {decimalDigits, 0, len(a[1]) - 1}})
}
if lo < 0 {
s = append(s, []charRun{{"-", 1, 1}, {nonZeroDigits, 1, 1}, {decimalDigits, 0, len(a[0]) - 2}})
}
return s
}
func ibanShape(a []string) textShape {
cc, digits := a[0], ibanLen[a[0]]-2
return textShape{{{cc[:1], 1, 1}, {cc[1:], 1, 1}, {decimalDigits, digits, digits}}}
}
// shortestFraction bounds the fraction FormatFloat's shortest form prints: at most 17
// significant digits after up to 323 zeros.
const shortestFraction = 340
// printedFloat is what strconv.FormatFloat(v, 'f', dp, 64) prints for a v in [lo, hi]
// that is whole when integral.
func printedFloat(lo, hi float64, dp int, integral bool) textShape {
digits := len(strconv.FormatFloat(math.Floor(math.Max(math.Abs(lo), math.Abs(hi))), 'f', 0, 64)) + 1 // one more for a rounding carry
wholes := [][]charRun{{{"0", 1, 1}}, {{nonZeroDigits, 1, 1}, {decimalDigits, 0, digits - 1}}}
fractions := [][]charRun{nil}
switch {
case dp > 0:
fractions = [][]charRun{{{".", 1, 1}, {decimalDigits, dp, dp}}}
case dp < 0 && !integral:
fractions = append(fractions, []charRun{{".", 1, 1}, {decimalDigits, 1, shortestFraction}})
}
signs := [][]charRun{nil}
if lo < 0 || math.Signbit(lo) {
signs = append(signs, []charRun{{"-", 1, 1}})
}
var s textShape
for _, sign := range signs {
for _, whole := range wholes {
for _, fraction := range fractions {
s = append(s, slices.Concat(sign, whole, fraction))
}
}
}
return s
}
const hexDigits = "0123456789abcdef"
// transforms are the builtins that rewrite one operand's value; they nest, so
@@ -183,19 +114,20 @@ var transforms = map[string]func(string) string{
// unwrapTransform peels nested transform calls off an operand arg, returning the
// field it finally names and the transforms to apply, innermost last.
func unwrapTransform(arg string) (leaf string, chain []string, err error) {
func unwrapTransform(arg string) (leaf string, chain []func(string) string, err error) {
for {
name, args, isCall := funcCall(arg)
if !isCall {
return arg, chain, nil
}
if _, isTransform := transforms[name]; !isTransform {
fn, isTransform := transforms[name]
if !isTransform {
return "", nil, fmt.Errorf("%s(%s) is not a transform, so it cannot be an operand", name, strings.Join(args, ","))
}
if len(args) != 1 {
return "", nil, fmt.Errorf("%s takes 1 arg, got %d", name, len(args))
}
chain = append(chain, name)
chain = append(chain, fn)
arg = args[0]
}
}
@@ -226,14 +158,10 @@ func transformPrep(outer func(string) string) func([]string) callFn {
if err != nil {
panic(fmt.Sprintf("fejkdata: transform arg %q reached prep unvalidated: %v", a[0], err))
}
fns := make([]func(string) string, len(chain))
for i, name := range chain {
fns[i] = transforms[name]
}
return func(_ *session, _ string, operands []string) string {
v := operands[0]
for i := len(fns) - 1; i >= 0; i-- {
v = fns[i](v)
for i := len(chain) - 1; i >= 0; i-- {
v = chain[i](v)
}
return outer(v)
}
+8 -254
View File
@@ -6,7 +6,6 @@ import (
"strconv"
"strings"
"unicode"
"unicode/utf8"
)
// calcNode is a parsed expression node. It evaluates over the operand values expand
@@ -200,6 +199,14 @@ func calcPrep(args []string) callFn {
}
}
// calcDecimals is a calc's decimals count, or -1 for the shortest form.
func calcDecimals(args []string) int {
if len(args) == 2 {
return atoi(args[1])
}
return -1
}
// indexVars replaces each operand name with its position in the values expand reads.
// Both sides take that order from calcVars, so they cannot drift.
func indexVars(n calcNode, at map[string]int) calcNode {
@@ -384,256 +391,3 @@ func contains(bs []byte, b byte) bool {
}
return false
}
// calcDecimals is a calc's decimals count, or -1 for the shortest form.
func calcDecimals(args []string) int {
if len(args) == 2 {
return atoi(args[1])
}
return -1
}
// calcLimit is the largest magnitude a proof accepts as finite, far enough below
// math.MaxFloat64 that rounding in the bounds cannot hide an overflow.
const calcLimit = 1e300
// maxOperandLen is the longest operand text a proof bounds by its length, so that
// bound, 10^maxOperandLen, stays within calcLimit.
const maxOperandLen = 300
// calcBound is what a proof knows of every value a calc can take: it lies in [lo, hi],
// is at least nonZero from zero unless nonZero is 0, and is whole when integral.
type calcBound struct {
lo, hi, nonZero float64
integral bool
}
func magnitude(b calcBound) float64 { return math.Max(math.Abs(b.lo), math.Abs(b.hi)) }
// doubt is why a proof could not show a calc finite, and the render that shows it.
type doubt struct{ render, why string }
type bounded struct {
b calcBound
d *doubt
}
// calcProof bounds a typed column's calcs from their operands' renders, to show each
// prints a number rather than NaN or Inf.
type calcProof struct {
decimal *textLanguage
operands map[node]bounded
lengths map[node]int
}
func newCalcProof() *calcProof {
p := &calcProof{operands: map[node]bounded{}, lengths: map[node]int{}}
p.decimal = newTextLanguage(decimalGrammar, p)
return p
}
// call bounds one calc token of t.
func (p *calcProof) call(t *template, args []string) (calcBound, *doubt) {
expr, err := parseCalc(args[0])
if err != nil {
panic(fmt.Sprintf("fejkdata: calc(%q) reached a proof unparsed: %v", args[0], err))
}
b, d := p.expr(expr, t.fields)
if d != nil {
return b, &doubt{d.render, fmt.Sprintf("{calc(%s)}: %s", strings.Join(args, ", "), d.why)}
}
return b, nil
}
func (p *calcProof) expr(n calcNode, fields map[string]node) (calcBound, *doubt) {
switch n := n.(type) {
case calcNum:
v := float64(n)
return calcBound{v, v, v, v == math.Trunc(v)}, nil
case calcVar:
return p.operand(string(n), fields[string(n)])
case calcNeg:
b, d := p.expr(n.x, fields)
return calcBound{-b.hi, -b.lo, b.nonZero, b.integral}, d
case calcBin:
l, d := p.expr(n.l, fields)
if d != nil {
return l, d
}
r, d := p.expr(n.r, fields)
if d != nil {
return r, d
}
return combine(n, l, r)
}
panic(fmt.Sprintf("fejkdata: calc node %T has no bound", n))
}
// combine bounds one operation from the bounds of its sides.
func combine(n calcBin, l, r calcBound) (calcBound, *doubt) {
b := calcBound{integral: l.integral && r.integral}
switch n.op {
case '+':
b.lo, b.hi = l.lo+r.lo, l.hi+r.hi
case '-':
b.lo, b.hi = l.lo-r.hi, l.hi-r.lo
case '*':
b.lo = min(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi)
b.hi = max(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi)
b.nonZero = l.nonZero * r.nonZero
default:
if r.nonZero == 0 {
return b, &doubt{"+Inf", fmt.Sprintf("divides by %s, which can be zero", calcText(n.r))}
}
m := magnitude(l) / r.nonZero
b = calcBound{lo: -m, hi: m, nonZero: l.nonZero / magnitude(r)}
}
if b.lo > 0 || b.hi < 0 {
b.nonZero = math.Max(b.nonZero, math.Min(math.Abs(b.lo), math.Abs(b.hi)))
}
if !(magnitude(b) <= calcLimit) {
return b, &doubt{"+Inf", calcText(n) + " can overflow"}
}
return b, nil
}
// operand bounds a calc operand, once per node.
func (p *calcProof) operand(name string, n node) (calcBound, *doubt) {
if seen, done := p.operands[n]; done {
return seen.b, seen.d
}
b, d := p.measure(name, n)
p.operands[n] = bounded{b, d}
return b, d
}
// measure bounds an operand through the calc it renders when that is all it renders,
// and otherwise from its text: a plain decimal of at most maxOperandLen bytes.
func (p *calcProof) measure(name string, n node) (calcBound, *doubt) {
if t, ok := n.(*template); ok {
if args, isCalc := soleCalc(t); isCalc {
b, d := p.call(t, args)
return rounded(b, calcDecimals(args)), d
}
}
text := p.decimal.node(n, nil)
if w, escapes := text.escape(decimalAccept); escapes {
why := fmt.Sprintf("operand %q can render %s, which is not a plain decimal", name, w)
if w.why != "" {
why += ": " + w.why
}
return calcBound{}, &doubt{"NaN", why}
}
size := p.length(n)
if size > maxOperandLen {
return calcBound{}, &doubt{"NaN", fmt.Sprintf("operand %q can render more than %d bytes, too many to bound", name, maxOperandLen)}
}
ends, m := text.to[1], math.Pow(10, float64(size))
b := calcBound{hi: m, nonZero: 1 / m, integral: ends&decimalFractional == 0}
if ends&decimalNegative != 0 {
b.lo = -m
}
if ends&decimalZero != 0 {
b.nonZero = 0
}
return b, nil
}
// soleCalc reports a template that renders one calc and nothing else, with its args.
func soleCalc(t *template) ([]string, bool) {
if t.repeat != 1 || len(t.ops) != 1 || t.ops[0].kind != 'b' {
return nil, false
}
name, args, _ := funcCall(t.format[1 : len(t.format)-1])
return args, name == "calc"
}
// rounded is b once printed to dp decimals, which moves a value by up to half a unit.
func rounded(b calcBound, dp int) calcBound {
if dp < 0 {
return b
}
half := math.Pow(10, -float64(dp)) / 2
return calcBound{b.lo - half, b.hi + half, math.Max(0, b.nonZero-half), b.integral || dp == 0}
}
// length is the most bytes a render of n can take, anything past maxOperandLen
// reported as maxOperandLen+1.
func (p *calcProof) length(n node) int {
if size, done := p.lengths[n]; done {
return size
}
size := 0
switch n := n.(type) {
case *choice:
for _, it := range n.items {
size = max(size, p.length(it))
}
case *template:
size = p.formatLength(n)*n.repeat + len(n.separator)*(n.repeat-1)
}
size = min(size, maxOperandLen+1)
p.lengths[n] = size
return size
}
func (p *calcProof) formatLength(t *template) int {
size := 0
_ = eachToken(t.format, func(tok ftoken) error {
if tok.kind == 'l' {
size += utf8.RuneLen(tok.r)
} else {
size += p.tokenLength(t, tok.body)
}
size = min(size, maxOperandLen+1)
return nil
})
return size
}
// tokenLength is the most bytes one token can print. A transform never lengthens a
// render that reads as a decimal: it maps each non-ASCII rune, two bytes or more, to at
// most two ASCII letters.
func (p *calcProof) tokenLength(t *template, body string) int {
name, args, isFunc := funcCall(body)
var arms []arm
switch _, isTransform := transforms[name]; {
case !isFunc:
arms = splitArms(body, t.refs)
case isTransform:
leaf, _, _ := unwrapTransform(args[0])
arms = []arm{splitArm(leaf, t.refs)}
case name == "calc":
b, d := p.call(t, args)
if d != nil {
return len(d.render)
}
return shapeLength(printedFloat(b.lo, b.hi, calcDecimals(args), b.integral))
default:
return shapeLength(builtins[name].emits(args))
}
size := 0
for _, a := range arms {
for _, leaf := range pathLeaves(t.fields[a.key], a.tail) {
size = max(size, p.length(leaf))
}
}
return size
}
// shapeLength is the most bytes a shape can emit, anything past maxOperandLen reported
// as maxOperandLen+1.
func shapeLength(s textShape) int {
longest := 0
for _, alt := range s {
size := 0
for _, run := range alt {
if run.max < 0 {
return maxOperandLen + 1
}
size += run.max
}
longest = max(longest, size)
}
return min(longest, maxOperandLen+1)
}
+23 -56
View File
@@ -16,7 +16,10 @@ const (
DataTypeBoolean
)
var dataTypeNames = [...]string{"string", "integer", "number", "boolean"}
var (
dataTypeNames = [...]string{"string", "integer", "number", "boolean"}
dataTypeNouns = [...]string{"text", "an integer", "a number", "a boolean"}
)
// String is the datatype as data spells it.
func (d DataType) String() string {
@@ -32,7 +35,7 @@ type position int
const (
inFormat position = iota // rendered by a format, so neither
atTop // a category or an inline template, whose fields are the columns
atTop // a category or an inline template, whose fields may be columns
inColumn // a column, or a choice item standing in for one
)
@@ -64,7 +67,7 @@ func datatypeOf(m map[string]any, pos position) (DataType, error) {
// columnDatatype is the datatype a column's items declare. They must agree, since a
// column holds one; a column only ever null is a string.
func columnDatatype(n node) (DataType, error) {
var declared []DataType
var items []*template
var collect func(node)
collect = func(n node) {
switch n := n.(type) {
@@ -73,68 +76,32 @@ func columnDatatype(n node) (DataType, error) {
collect(it)
}
case *template:
declared = append(declared, n.datatype)
items = append(items, n)
}
}
collect(n)
if len(declared) == 0 {
if len(items) == 0 {
return DataTypeString, nil
}
for _, d := range declared {
if d != declared[0] {
return declared[0], fmt.Errorf("its items declare %s and %s; a column holds one datatype, so give every item the same", declared[0], d)
for _, t := range items[1:] {
if t.datatype != items[0].datatype {
return items[0].datatype, disagreement(items[0], t)
}
}
return declared[0], nil
return items[0].datatype, nil
}
// datatypeSpec is what a datatype's text must satisfy: a grammar, the states a render
// may end in, and how an error names the datatype.
type datatypeSpec struct {
grammar *grammar
accept uint32
noun string
}
var datatypeSpecs = map[DataType]datatypeSpec{
DataTypeInteger: {numberGrammar, integerAccept, "an integer"},
DataTypeNumber: {numberGrammar, numberAccept, "a number"},
DataTypeBoolean: {booleanGrammar, booleanAccept, "a boolean"},
}
// datatypeCheck proves every render of a typed column is text its datatype takes. One
// check covers a scope, so a node several columns reach is read once per grammar.
type datatypeCheck struct {
languages map[*grammar]*textLanguage
proof *calcProof
}
func (c *datatypeCheck) check(path string, n node) error {
t, ok := n.(*template)
if !ok || t.datatype == DataTypeString {
return nil
// disagreement names the fix for two items of one column declaring different datatypes.
func disagreement(a, b *template) error {
typed, bare := a, b
if typed.datatype == DataTypeString {
typed, bare = b, a
}
spec := datatypeSpecs[t.datatype]
w, escapes := c.language(spec.grammar).node(t, nil).escape(spec.accept)
if !escapes {
return nil
switch {
case bare.datatype != DataTypeString:
return fmt.Errorf("its items declare %s and %s; a column holds one datatype", a.datatype, b.datatype)
case len(bare.fields) == 0 && bare.repeat == 1:
return fmt.Errorf(`item %q declares no datatype, and a column holds one; write it as {"format":%q,"datatype":%q}`, bare.format, bare.format, typed.datatype)
}
msg := fmt.Sprintf("%s: datatype %s, but it can render %s, which is not %s", path, t.datatype, w, spec.noun)
if w.why != "" {
msg += ": " + w.why
}
return errors.New(msg)
}
func (c *datatypeCheck) language(g *grammar) *textLanguage {
if c.proof == nil {
c.proof = newCalcProof()
c.languages = map[*grammar]*textLanguage{}
}
l, made := c.languages[g]
if !made {
l = newTextLanguage(g, c.proof)
c.languages[g] = l
}
return l
return fmt.Errorf(`an item declares no datatype beside one declaring %s; a column holds one, so give it "datatype": %q`, typed.datatype, typed.datatype)
}
+1 -1
View File
@@ -189,7 +189,7 @@ func checkScope(s nodeScope) error {
if err := s(heldCheck); err != nil {
return err
}
return s((&datatypeCheck{}).check)
return s((&valueProof{}).checkDatatype)
}
type reachMemo map[node]int
+1 -1
View File
@@ -229,7 +229,7 @@ func compileTemplate(m map[string]any, pos position) (node, error) {
return nil, err
}
fieldPos := inFormat
if pos == atTop && o.repeat == 1 {
if pos == atTop && projectsColumns(o.repeat) {
fieldPos = inColumn
}
fields, err := compileFields(m, fieldPos)
+8 -6
View File
@@ -63,16 +63,14 @@ func (r *Record) CSVHeader() string {
}
// CSVLine renders the column values as one CSV row: a null column an empty field and an
// empty string "", the convention PostgreSQL's COPY reads a null by.
// empty string "", the convention PostgreSQL's COPY reads a null by. A record of one null
// column is a blank line, which COPY reads as null and most CSV readers skip.
func (r *Record) CSVLine() string {
fields := make([]string, len(r.columns))
for i, c := range r.columns {
fields[i] = literal(c, csvField, "")
}
if line := strings.Join(fields, ","); line != "" {
return line
}
return `""` // a blank line is a row every CSV reader drops
return strings.Join(fields, ",")
}
func csvField(s string) string {
@@ -207,7 +205,7 @@ func recordOf(n node) (*template, []Column, error) {
if !ok {
return nil, nil, errors.New("names a choice, not a template; a record is a template whose fields are its columns")
}
if t.repeat != 1 {
if !projectsColumns(t.repeat) {
return nil, nil, fmt.Errorf("carries repeat %d, which composes its format into one string; a record projects columns instead — drop the repeat and render the record again for more rows", t.repeat)
}
names := recordColumns(t)
@@ -225,6 +223,10 @@ func recordOf(n node) (*template, []Column, error) {
return t, columns, nil
}
// projectsColumns reports whether a category or inline template with this repeat is a
// record, its fields the columns; a repeat composes the format into one string instead.
func projectsColumns(repeat int) bool { return repeat == 1 }
// checkColumnRefs rejects the reference reads a record's shared draw cannot answer
// for: one column rendering a level another reads a path into, and a column
// reading the record back through its own path.
-403
View File
@@ -1,403 +0,0 @@
package fejkdata
import (
"slices"
"strconv"
"strings"
"unicode/utf8"
)
// grammar is a deterministic automaton over a scalar's text: state 0 is dead, 1 the
// start, and each state lists the runes that leave it and where they lead.
type grammar [][]arc
type arc struct {
on string
to int
}
func (g *grammar) run(q int, s string) int {
for _, r := range s {
if q = g.step(q, r); q == 0 {
return 0
}
}
return q
}
func (g *grammar) step(q int, r rune) int {
for _, a := range (*g)[q] {
if strings.ContainsRune(a.on, r) {
return a.to
}
}
return 0
}
const (
decimalDigits = "0123456789"
nonZeroDigits = "123456789"
)
// numberGrammar reads a JSON number. States: 2 "-", 3 "0", 4 more integer digits, 5 ".",
// 6 fraction digits, 7 "e", 8 its sign, 9 exponent digits.
var numberGrammar = &grammar{
nil,
{{"-", 2}, {"0", 3}, {nonZeroDigits, 4}},
{{"0", 3}, {nonZeroDigits, 4}},
{{".", 5}, {"eE", 7}},
{{decimalDigits, 4}, {".", 5}, {"eE", 7}},
{{decimalDigits, 6}},
{{decimalDigits, 6}, {"eE", 7}},
{{"+-", 8}, {decimalDigits, 9}},
{{decimalDigits, 9}},
{{decimalDigits, 9}},
}
const (
integerAccept uint32 = 1<<3 | 1<<4
numberAccept = integerAccept | 1<<6 | 1<<9
)
var booleanGrammar = &grammar{
nil,
{{"t", 2}, {"f", 6}},
{{"r", 3}}, {{"u", 4}}, {{"e", 5}}, nil,
{{"a", 7}}, {{"l", 8}}, {{"s", 9}}, {{"e", 10}}, nil,
}
const booleanAccept uint32 = 1<<5 | 1<<10
// decimalGrammar reads what a calc operand must render to be proven finite: a sign,
// digits and at most one dot. Past the sign, states 4–9 are positive and 10–15 their
// negatives: 4 zero digits, 5 a nonzero integer, 6 a leading dot, 7 zero with a dot,
// 8 a nonzero integer with a zero fraction, 9 a nonzero fraction.
var decimalGrammar = &grammar{
nil,
{{"+", 2}, {"-", 3}, {"0", 4}, {nonZeroDigits, 5}, {".", 6}},
{{"0", 4}, {nonZeroDigits, 5}, {".", 6}},
{{"0", 10}, {nonZeroDigits, 11}, {".", 12}},
{{"0", 4}, {nonZeroDigits, 5}, {".", 7}},
{{decimalDigits, 5}, {".", 8}},
{{"0", 7}, {nonZeroDigits, 9}},
{{"0", 7}, {nonZeroDigits, 9}},
{{"0", 8}, {nonZeroDigits, 9}},
{{decimalDigits, 9}},
{{"0", 10}, {nonZeroDigits, 11}, {".", 13}},
{{decimalDigits, 11}, {".", 14}},
{{"0", 13}, {nonZeroDigits, 15}},
{{"0", 13}, {nonZeroDigits, 15}},
{{"0", 14}, {nonZeroDigits, 15}},
{{decimalDigits, 15}},
}
const (
decimalAccept uint32 = 1<<4 | 1<<5 | 1<<7 | 1<<8 | 1<<9 | 1<<10 | 1<<11 | 1<<13 | 1<<14 | 1<<15
decimalNegative uint32 = 0xfc00
decimalZero uint32 = 1<<4 | 1<<7 | 1<<10 | 1<<13
decimalFractional uint32 = 1<<9 | 1<<15
)
// relation is what a node's renders do to a grammar: from each state, the states a
// render can end in, and one render reaching each.
type relation struct {
g *grammar
to []uint32
w []witness // w[from*len(to)+to]
}
// witness is one render, cut past witnessCap bytes, and why it can occur when the text
// alone does not say.
type witness struct {
text string
cut bool
why string
}
const witnessCap = 60
func (w witness) then(next witness) witness {
if w.why == "" {
w.why = next.why
}
if w.cut {
return w
}
w.text += next.text
w.cut = next.cut
if len(w.text) > witnessCap {
end := witnessCap
for !utf8.RuneStart(w.text[end]) {
end--
}
w.text, w.cut = w.text[:end], true
}
return w
}
func (w witness) String() string {
if w.cut {
return strconv.Quote(w.text + "…")
}
return strconv.Quote(w.text)
}
func newRelation(g *grammar) *relation {
n := len(*g)
return &relation{g: g, to: make([]uint32, n), w: make([]witness, n*n)}
}
func (r *relation) add(from, to int, w witness) {
if r.to[from]&(1<<to) == 0 {
r.to[from] |= 1 << to
r.w[from*len(r.to)+to] = w
}
}
// textRelation is the relation of a render that is always s.
func textRelation(g *grammar, s, why string) *relation {
r := newRelation(g)
w := witness{why: why}.then(witness{text: s})
for q := range r.to {
r.add(q, g.run(q, s), w)
}
return r
}
// union is the renders of either relation; a nil relation has none.
func union(a, b *relation) *relation {
if a == nil {
return b
}
if b == nil {
return a
}
u := newRelation(a.g)
for _, r := range []*relation{a, b} {
for from, ends := range r.to {
for to := range r.to {
if ends&(1<<to) != 0 {
u.add(from, to, r.w[from*len(r.to)+to])
}
}
}
}
return u
}
// then is a render of r followed by a render of next.
func (r *relation) then(next *relation) *relation {
c := newRelation(r.g)
n := len(r.to)
for from, mids := range r.to {
for mid := 0; mid < n; mid++ {
if mids&(1<<mid) == 0 {
continue
}
for to := 0; to < n; to++ {
if next.to[mid]&(1<<to) != 0 && c.to[from]&(1<<to) == 0 {
c.add(from, to, r.w[from*n+mid].then(next.w[mid*n+to]))
}
}
}
}
return c
}
// power is k renders of r in a row, k at least 1, composed by squaring.
func (r *relation) power(k int) *relation {
var out *relation
for base := r; ; base = base.then(base) {
if k&1 == 1 {
if out == nil {
out = base
} else {
out = out.then(base)
}
}
if k >>= 1; k == 0 {
return out
}
}
}
// closure is any number of renders of r in a row, where r includes the empty render.
func (r *relation) closure() *relation {
for {
next := r.then(r)
if slices.Equal(next.to, r.to) {
return r
}
r = next
}
}
// escape finds a render from the start that ends outside accept, preferring one that
// carries a reason.
func (r *relation) escape(accept uint32) (witness, bool) {
var found witness
escapes := false
for to := range r.to {
if (r.to[1]&^accept)&(1<<to) == 0 {
continue
}
if w := r.w[len(r.to)+to]; !escapes || found.why == "" && w.why != "" {
found, escapes = w, true
}
}
return found, escapes
}
// textShape is the text a builtin can emit: alternatives, each a sequence of runs.
type textShape [][]charRun
// charRun is between min and max characters, each one of chars; max -1 is unbounded.
// chars is ASCII, so a run of k characters is k bytes.
type charRun struct {
chars string
min, max int
}
// textLanguage is what one grammar makes of the renders a check reads, worked out once
// per node and fold.
type textLanguage struct {
g *grammar
proof *calcProof
memo map[languageKey]*relation
empty *relation
}
type languageKey struct {
n node
fold string
}
// fold is the transforms a render passes through before the grammar reads it, innermost
// first. Each rewrites rune by rune, so folding a render is folding each of its pieces.
type fold []string
func (f fold) apply(s string) string {
for _, name := range f {
s = transforms[name](s)
}
return s
}
func newTextLanguage(g *grammar, proof *calcProof) *textLanguage {
return &textLanguage{g: g, proof: proof, memo: map[languageKey]*relation{}, empty: textRelation(g, "", "")}
}
func (l *textLanguage) node(n node, f fold) *relation {
key := languageKey{n, strings.Join(f, ",")}
if r, done := l.memo[key]; done {
return r
}
r := l.empty // a null renders ""
switch n := n.(type) {
case *choice:
r = nil
for _, it := range n.items {
r = union(r, l.node(it, f))
}
case *template:
r = l.format(n, f)
if n.repeat > 1 {
r = r.then(l.text(f.apply(n.separator)).then(r).power(n.repeat - 1))
}
}
l.memo[key] = r
return r
}
func (l *textLanguage) text(s string) *relation { return textRelation(l.g, s, "") }
// format reads a template's format the way expand renders it: literal runs and tokens
// in turn.
func (l *textLanguage) format(t *template, f fold) *relation {
r := l.empty
var lit strings.Builder
_ = eachToken(t.format, func(tok ftoken) error {
if tok.kind == 'l' {
lit.WriteRune(tok.r)
return nil
}
r = r.then(l.text(f.apply(lit.String()))).then(l.token(t, tok.body, f))
lit.Reset()
return nil
})
return r.then(l.text(f.apply(lit.String())))
}
// token reads one {…} token: a field read, a transform over one, a calc, or what a
// builtin emits.
func (l *textLanguage) token(t *template, body string, f fold) *relation {
name, args, isFunc := funcCall(body)
if !isFunc {
var r *relation
for _, a := range splitArms(body, t.refs) {
r = union(r, l.read(t, a, f))
}
return r
}
if _, isTransform := transforms[name]; isTransform {
leaf, chain, _ := unwrapTransform(args[0])
inner := slices.Clone(chain)
slices.Reverse(inner)
return l.read(t, splitArm(leaf, t.refs), append(append(inner, name), f...))
}
if name == "calc" {
return l.calc(t, args, f)
}
return l.shape(builtins[name].emits(args), f)
}
// read is one arm of a token: every node its path can land on.
func (l *textLanguage) read(t *template, a arm, f fold) *relation {
var r *relation
for _, leaf := range pathLeaves(t.fields[a.key], a.tail) {
r = union(r, l.node(leaf, f))
}
return r
}
func (l *textLanguage) calc(t *template, args []string, f fold) *relation {
b, d := l.proof.call(t, args)
if d != nil {
return textRelation(l.g, f.apply(d.render), d.why)
}
return l.shape(printedFloat(b.lo, b.hi, calcDecimals(args), b.integral), f)
}
func (l *textLanguage) shape(s textShape, f fold) *relation {
var r *relation
for _, alt := range s {
seq := l.empty
for _, run := range alt {
seq = seq.then(l.run(run, f))
}
r = union(r, seq)
}
return r
}
// run reads a charRun: min characters, then up to max-min more.
func (l *textLanguage) run(c charRun, f fold) *relation {
one := newRelation(l.g)
for from := range one.to {
for _, ch := range c.chars {
s := f.apply(string(ch))
one.add(from, l.g.run(from, s), witness{text: s})
}
}
more := l.empty
switch optional := union(one, l.empty); {
case c.max < 0:
more = optional.closure()
case c.max > c.min:
more = optional.power(c.max - c.min)
}
if c.min == 0 {
return more
}
return one.power(c.min).then(more)
}
+3 -3
View File
@@ -72,9 +72,9 @@ type builtin struct {
// operands names the fields the call reads, which expand renders for it; nil
// for a builtin that reads none.
operands func(args []string) []string
// emits is the text a call can print, for the datatype check; nil for calc and the
// transforms, whose text the check derives from what they read.
emits func(args []string) textShape
// number bounds the number a call prints and names the datatype its text is; nil
// for a builtin that prints text.
number func(args []string) (proven, DataType)
}
// funcCall splits a "{token}" body shaped name(args) into its parts; ok is false
+4
View File
@@ -22,6 +22,10 @@ The record API lands first, so the data update can use it.
- `code` and `symbol` sibling fields reading `currency`, as `{code} {symbol}` → a matching pair
- `{a} & {b}`, each reading `person` → one person, or two when `a` and `b` name different groups
- two bare `{/sv_SE.word}` → two words
- Reference inheritance — settle whether a column that is exactly one reference to
another record's column, like `{/src.score}`, takes that column's datatype and
null. Today a null there writes `""`, and a typed column reading it is refused.
Settle before draw groups and the data update.
### Data
+264
View File
@@ -0,0 +1,264 @@
package fejkdata
import (
"fmt"
"math"
"regexp"
"strconv"
"strings"
)
// proven is what a proof knows of every render of a node: bounds on the number each
// reads as, and per datatype why some render's text is not one ("" when none).
type proven struct {
lo, hi float64
nonZero float64 // every value is at least this far from zero; 0 when one can be zero
integral bool
notNumber string // why some render reads as no finite number, the way calc reads it
not [len(dataTypeNames)]string
}
// valueProof proves what typed columns and their calc operands hold, each node once per
// scope. A typed column holds one value: a literal, one value builtin, one calc, or a
// read of such values.
type valueProof struct {
memo map[node]proven
}
// checkDatatype rejects a typed column some render of which is not text of its datatype.
func (p *valueProof) checkDatatype(path string, n node) error {
t, ok := n.(*template)
if !ok || t.datatype == DataTypeString {
return nil
}
if err := p.prove(t, t.datatype); err != nil {
return fmt.Errorf("%s: %w", path, err)
}
return nil
}
// prove reports why some render of n is not text of datatype d.
func (p *valueProof) prove(n node, d DataType) error {
if reason := p.of(n).not[d]; reason != "" {
return fmt.Errorf("datatype %s: %s", d, reason)
}
return nil
}
func (p *valueProof) of(n node) proven {
if v, done := p.memo[n]; done {
return v
}
if p.memo == nil {
p.memo = map[node]proven{}
}
var v proven
switch n := n.(type) {
case *choice:
v = p.unite(n.items)
case *template:
v = p.template(n)
default:
v = unproven(`it reads a null, which renders "" outside its own column`)
}
p.memo[n] = v
return v
}
func (p *valueProof) unite(nodes []node) proven {
v := p.of(nodes[0])
for _, n := range nodes[1:] {
w := p.of(n)
v.lo, v.hi, v.nonZero = min(v.lo, w.lo), max(v.hi, w.hi), min(v.nonZero, w.nonZero)
v.integral = v.integral && w.integral
if v.notNumber == "" {
v.notNumber = w.notNumber
}
for d := range v.not {
if v.not[d] == "" {
v.not[d] = w.not[d]
}
}
}
return v
}
// template proves a template that renders one value: fixed text, or a format that is
// one token alone.
func (p *valueProof) template(t *template) proven {
switch {
case t.repeat != 1:
return unproven(fmt.Sprintf("%q carries a repeat, which composes text rather than one value", t.format))
case t.fixed:
return literalValue(t.lit)
case len(t.ops) != 1:
return unproven(fmt.Sprintf("%q is not one value; write one literal or one {int()}, {float()}, {seq()} or {calc()}, or read one", t.format))
}
body := t.format[1 : len(t.format)-1]
name, args, isFunc := funcCall(body)
switch _, isTransform := transforms[name]; {
case !isFunc:
var leaves []node
for _, a := range splitArms(body, t.refs) {
leaves = append(leaves, pathLeaves(t.fields[a.key], a.tail)...)
}
return p.unite(leaves)
case name == "calc":
return p.calc(t, body, args)
case builtins[name].number != nil:
v, prints := builtins[name].number(args)
return printing(body, prints, v)
case isTransform:
return unproven(fmt.Sprintf("{%s} rewrites text rather than printing a value; write the values it would print", body))
}
return printing(body, DataTypeString, proven{notNumber: fmt.Sprintf("{%s} prints text, not a number", body)})
}
func (p *valueProof) calc(t *template, body string, args []string) proven {
expr, err := parseCalc(args[0])
if err != nil {
panic(fmt.Sprintf("fejkdata: calc(%q) reached a proof unparsed: %v", args[0], err))
}
v, doubt := p.expr(expr, t.fields)
if doubt == "" && !(magnitude(v) <= calcLimit) {
doubt = calcText(expr) + " is not proven within 1e300"
}
if doubt != "" {
return unproven(fmt.Sprintf("{%s}: %s", body, doubt))
}
v, prints := printedNumber(v, calcDecimals(args))
return printing(body, prints, v)
}
// calcLimit is the largest magnitude a proof accepts as finite, far enough below
// math.MaxFloat64 that rounding in the bounds cannot hide an overflow.
const calcLimit = 1e300
// expr bounds a calc expression from its operands, or says why it cannot.
func (p *valueProof) expr(n calcNode, fields map[string]node) (proven, string) {
switch n := n.(type) {
case calcNum:
v := float64(n)
return bounded(v, v, v == math.Trunc(v)), ""
case calcVar:
v := p.of(fields[string(n)])
if v.notNumber != "" {
return proven{}, fmt.Sprintf("operand %q: %s", string(n), v.notNumber)
}
return proven{lo: v.lo, hi: v.hi, nonZero: v.nonZero, integral: v.integral}, ""
case calcNeg:
v, doubt := p.expr(n.x, fields)
v.lo, v.hi = -v.hi, -v.lo
return v, doubt
case calcBin:
l, doubt := p.expr(n.l, fields)
if doubt != "" {
return l, doubt
}
r, doubt := p.expr(n.r, fields)
if doubt != "" {
return r, doubt
}
return combine(n, l, r)
}
panic(fmt.Sprintf("fejkdata: calc node %T has no bound", n))
}
// combine bounds one operation from the bounds of its sides.
func combine(n calcBin, l, r proven) (proven, string) {
var v proven
integral := l.integral && r.integral
switch n.op {
case '+':
v = bounded(l.lo+r.lo, l.hi+r.hi, integral)
case '-':
v = bounded(l.lo-r.hi, l.hi-r.lo, integral)
case '*':
v = bounded(min(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi), max(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi), integral)
v.nonZero = max(v.nonZero, l.nonZero*r.nonZero)
default:
if r.nonZero == 0 {
return v, fmt.Sprintf("divides by %s, which is not proven nonzero", calcText(n.r))
}
m := magnitude(l) / r.nonZero
v = proven{lo: -m, hi: m, nonZero: l.nonZero / magnitude(r)}
}
if !(magnitude(v) <= calcLimit) {
return v, calcText(n) + " is not proven within 1e300"
}
return v, ""
}
// bounded is a number in [lo, hi], its distance from zero read off the bounds.
func bounded(lo, hi float64, integral bool) proven {
v := proven{lo: lo, hi: hi, integral: integral}
switch {
case lo > 0:
v.nonZero = lo
case hi < 0:
v.nonZero = -hi
}
return v
}
func magnitude(v proven) float64 { return math.Max(math.Abs(v.lo), math.Abs(v.hi)) }
// printedNumber is v once strconv.FormatFloat prints it to dp decimals, and the datatype
// that text is: an integer when whole and within int64, else a number.
func printedNumber(v proven, dp int) (proven, DataType) {
if dp >= 0 {
half := math.Pow(10, -float64(dp)) / 2
v = proven{lo: v.lo - half, hi: v.hi + half, nonZero: math.Max(0, v.nonZero-half), integral: v.integral || dp == 0}
}
if (dp == 0 || dp < 0 && v.integral) && magnitude(v) < math.MaxInt64 {
return v, DataTypeInteger
}
return v, DataTypeNumber
}
// printing is v for a token whose every render is text of datatype prints, with a reason
// against each datatype that text is not.
func printing(token string, prints DataType, v proven) proven {
for d := DataTypeInteger; d <= DataTypeBoolean; d++ {
if prints != d && !(prints == DataTypeInteger && d == DataTypeNumber) {
v.not[d] = fmt.Sprintf("{%s} prints %s, not %s", token, dataTypeNouns[prints], dataTypeNouns[d])
}
}
return v
}
// unproven is a render no datatype and no calc can take, for why.
func unproven(why string) proven {
v := proven{notNumber: why}
for d := DataTypeInteger; d <= DataTypeBoolean; d++ {
v.not[d] = why
}
return v
}
var (
integerText = regexp.MustCompile(`^-?(0|[1-9][0-9]*)$`)
numberText = regexp.MustCompile(`^-?(0|[1-9][0-9]*)(\.[0-9]+)?([eE][+-]?[0-9]+)?$`)
)
// literalValue proves fixed text: the number calc reads it as, and each datatype it is.
func literalValue(text string) proven {
var v proven
if f, err := strconv.ParseFloat(strings.TrimSpace(text), 64); err != nil || math.IsNaN(f) || math.IsInf(f, 0) {
v.notNumber = fmt.Sprintf("%q is not a number", text)
} else {
v = bounded(f, f, f == math.Trunc(f))
}
if _, err := strconv.ParseInt(text, 10, 64); !integerText.MatchString(text) {
v.not[DataTypeInteger] = fmt.Sprintf("%q is not an integer", text)
} else if err != nil {
v.not[DataTypeInteger] = fmt.Sprintf("%q is past the int64 range", text)
}
if v.notNumber != "" || !numberText.MatchString(text) {
v.not[DataTypeNumber] = fmt.Sprintf("%q is not a number", text)
}
if text != "true" && text != "false" {
v.not[DataTypeBoolean] = fmt.Sprintf("%q is not a boolean", text)
}
return v
}