diff --git a/README.md b/README.md index 02c7f22..2cda366 100644 --- a/README.md +++ b/README.md @@ -82,8 +82,8 @@ For structured output a record writes the row for you. A record is a template seen as columns: its fields are the columns, its `format` the whole. `--format json|ndjson|csv|sql` writes the records; the library's -`FakeRecord` (below) hands back the columns. Every column is a string — typed scalars -are on the release checklist, see [`todo.md`](todo.md). Save +`FakeRecord` (below) hands back the columns. A column is a string unless it declares a +[datatype](#datatype), and a [`null`](#null) item draws it as null. Save `mydata/users.json`: ```json @@ -164,7 +164,8 @@ r, err = f.FakeRecordTemplate(`{"format":"{x}","x":["a","b"]}`) // compile + ren | `WithDataFS(fsys)` | layer an `fs.FS`, such as your own `embed.FS` | | `WithoutShippedData()` | load only what you give | -A `*Record` carries its columns via `Columns()`, and serializes them with `JSON()` +A `*Record` carries its columns via `Columns()` — each a `Column` of `Name`, +`DataType`, rendered `Value` and `Null` — and serializes them with `JSON()` (one object), `CSVHeader()`/`CSVLine()`, or `SQLInsert(table)` — the shapes the CLI's `--format` writes. `FakeRecord` and `FakeRecordTemplate` take a record; a path or template that is not one — a bare string, a choice, or a folder — errors. @@ -255,10 +256,59 @@ Renders e.g. `bar foo baz`. Rejected at load: a `separator` without a `repeat`, a `separator` of `""` (the default), and a `repeat` that multiplies to more than 1 048 576 renders along any path of nested repeats. +### Datatype + +A record column may declare `datatype` — `integer`, `number` or `boolean` — so `json` +writes `42` rather than `"42"` and `sql` a bare literal; a column without one is a +string: + +```json +{ "format": "", + "id": { "format": "{seq()}", "datatype": "integer" }, + "paid": { "format": "{p}", "p": ["true", "false"], "datatype": "boolean" }, + "total": { "format": "{calc(net * qty, 2)}", "net": ["19.99", "5.00"], "qty": ["3", "7"], "datatype": "number" } } +``` + +Writes e.g. `{"id":1,"paid":true,"total":59.97}`. A column is a field of the top-level +template, or an item of a choice standing in for one; `datatype` anywhere else is a +load error. A typed column holds one value, alone in its format: a literal, one +`{int()}`, `{float()}`, `{seq()}` or `{calc()}` call, or a read that lands only on such +values. `integer` is an int64 written `0|-?[1-9][0-9]*` — `{float()}` prints one at +`0` decimals within int64 — `number` a JSON number, `boolean` `true` or `false`. A value its +datatype cannot hold is a load error naming it: + +```text +order.id: datatype integer: {digits(3)} prints text, not an integer +order.id: datatype integer: "1{digits(2)}" is not one value; write one literal or one {int()}, {float()}, {seq()} or {calc()}, or read one +``` + +A typed column's `{calc()}` must be proven to print a number: each operand a number +literal, an `{int()}`, `{float()}`, `{seq()}` or `{digits()}` call, a calc, or a read of +such values, whose bounds keep every divisor from zero and the result within `1e300`. +What the bounds cannot show is refused — `{calc(a / b)}: divides by b, which is not +proven nonzero`. The calc fills an `integer` column at `0` decimals, or over whole +operands with no `/`, while its bounds stay within int64. + +### Null + +A `null` item draws a record column as null: `json` writes `null`, `sql` `NULL`, and +`csv` an empty field, with an empty string written `""` — the convention PostgreSQL's +`COPY … CSV` reads. A record of one null column is a blank line, which `COPY` reads as +null but most CSV readers skip, so write such a record as `json` or `sql`. `Fake` +renders a null as `""`. The other items' weights skew its odds: + +```json +{ "format": "", "deleted_at": null, "middle": [null, { "format": "{n}", "n": ["Ann", "Eva"], "weight": 3 }] } +``` + +`deleted_at` is null every draw, `middle` a name three draws in four. Rejected at +load: `null` anywhere but a column, naming `""`, and a column whose items declare +different datatypes. + ### Options and fields -`format`, `weight`, `repeat` and `separator` are the only options; **any other -key is a field** (see [Decisions](#decisions)). An object that does nothing a +`format`, `weight`, `repeat`, `separator` and `datatype` are the only options; **any +other key is a field** (see [Decisions](#decisions)). An object that does nothing a string can't — only a `format` — is rejected naming the string, as is a one-item choice naming its item. @@ -315,11 +365,12 @@ hyphenated field can't be an operand. { "format": "{net} x {qty} = {calc(net * qty, 2)}", "net": ["19.99", "5.00"], "qty": ["3", "7"] } ``` -Renders e.g. `19.99 x 3 = 59.97`. An operand that can never be a number (`"abc"`, +Renders e.g. `19.99 x 3 = 59.97`. A result that rounds to zero prints unsigned — `0`, +`0.00` — as `{float()}`'s does. An operand that can never be a number (`"abc"`, or a choice of such) is rejected at load, as is a division by a constant zero (`1/0`, or a fixed `"0"` field); an operand that sometimes is not a number yields `NaN`, and a division by one that is not constant `Inf` — both print rather than -fail. +fail, except in a [typed column](#datatype), which must prove neither happens. ### Transforms @@ -437,8 +488,8 @@ tokens add cost in proportion to the output. ## Decisions -- **Options and fields share one namespace.** `format`, `weight`, `repeat` and - `separator` are reserved; every other key is a field. Nesting fields under a +- **Options and fields share one namespace.** `format`, `weight`, `repeat`, + `separator` and `datatype` are reserved; every other key is a field. Nesting fields under a key, or prefixing options, would tax every template to guard against a misspelt option. - **`{a|b}` stays beside nested choices.** `[[…], […]]` picks the same way, but @@ -512,13 +563,16 @@ tokens add cost in proportion to the output. - **64-bit targets only.** The gate builds amd64, and the buffer sizing a render pre-computes (renders × bytes) assumes a 64-bit int; on a 32-bit target it could overflow and panic. -- **A constant zero divisor is a load error; a divisor that is not constant prints - `Inf`.** `1/0` and a fixed `"0"` field are decidable, so they join the - never-numeric operand as a load error; the fold stops where an operand varies, +- **A constant zero divisor is a load error; in a string column a divisor that is not + constant prints `Inf`.** `1/0` and a fixed `"0"` field are decidable, so they join + the never-numeric operand as a load error; the fold stops where an operand varies, so `a/(b*c)` with `b` fixed at `0` and `c` varying loads and prints `Inf` every - draw — catching it needs zero-absorbing algebra for a shape nobody writes. + draw — catching it needs zero-absorbing algebra for a shape nobody writes. A + [typed column](#datatype) bounds its operands instead and refuses a divisor it + cannot keep from zero. - **In data, a default written out and a constant spelled as a sample are load - errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `int(5,5)`, `float(1,1,2)`, + errors.** `weight: 1`, `repeat: 1`, `separator: ""`, `datatype: "string"`, + `int(5,5)`, `float(1,1,2)`, `+5` and `05` each spell what a shorter form already spells, so each is rejected naming that form. The CLI's numbers follow the shell instead: `--seed 007` and `--repeat +3` are 7 and 3, as every command line reads them. @@ -557,6 +611,20 @@ tokens add cost in proportion to the output. row — would vary per draw. A fixed column set is what the CSV and `INSERT` contracts rest on, so the restriction holds even where a particular choice would happen to agree. +- **Null is a `null` item, not a rate.** A null is one more outcome of a column's + draw, so a choice's weights skew it like any other; a null-rate option would be a + second way to state odds. +- **A typed column holds one value, not composed text.** Its bounds come from a + literal or a call's arguments, so a load error names a real value, a range check is + one comparison, and `1{digits(2)}` is a second spelling of `{int(100,199)}`. +- **A typed column's calc is refused unless proven.** Operand bounds must keep each + divisor from zero and the result finite; what they cannot show is refused rather + than trusted, since a bare `NaN` breaks the JSON and SQL it lands in. +- **`Column` carries text, not a Go value.** `Value` is the rendered string beside + `DataType` and `Null`, which each serializer writes as the load check proved it; a + `Value any` would hand every caller a type switch. +- **The package stays flat.** Go ties a package to one directory, so folders would + split the API into packages. - **The performance gate asserts allocations, not wall-clock time.** `AllocsPerRun` is deterministic across machines, so a ±10% ceiling does not flake under CI load, while time varies with the machine and its neighbours. A rendering slowdown @@ -613,6 +681,8 @@ reference.go reference sigils, and binding references across the tree graph.go the render graph: edges, cycles, the repeat bound, tree walks builtins.go the {name()} function registry and its implementations calc.go the {calc()} arithmetic evaluator: parser, eval, validation +datatype.go column datatypes: DataType, where datatype and null may sit, a column's datatype +value.go the value proof: what a typed column or calc operand holds, checked at load data.go data loading: fs.FS folders/files -> namespace tree, multi-source merge cmd/fejkdata/ the fejkdata CLI data/ shipped data (JSON), embedded at build: locale folders + a misc folder diff --git a/builtins.go b/builtins.go index 8b9fec6..6075668 100644 --- a/builtins.go +++ b/builtins.go @@ -31,9 +31,11 @@ var builtins = map[string]builtin{ "ulid": {arity: 0, prep: sample(ulid)}, "nanoid": {arity: 1, check: posIntArg, prep: chars(nanoidAlphabet)}, "hex": {arity: 1, check: posIntArg, prep: chars(hexDigits)}, - "digits": {arity: 1, check: posIntArg, prep: chars("0123456789")}, - "upper": {arity: 1, check: posIntArg, prep: chars("ABCDEFGHIJKLMNOPQRSTUVWXYZ")}, - "lower": {arity: 1, check: posIntArg, prep: chars("abcdefghijklmnopqrstuvwxyz")}, + "digits": {arity: 1, check: posIntArg, prep: chars("0123456789"), number: func(token string, a []string) proven { + return printing(token, DataTypeString, bounded(0, math.Pow(10, float64(atoi(a[0])))-1, true)) + }}, + "upper": {arity: 1, check: posIntArg, prep: chars("ABCDEFGHIJKLMNOPQRSTUVWXYZ")}, + "lower": {arity: 1, check: posIntArg, prep: chars("abcdefghijklmnopqrstuvwxyz")}, "base64": {arity: 1, check: posIntArg, prep: func(a []string) callFn { n := atoi(a[0]) return func(s *session, _ string, _ []string) string { @@ -43,12 +45,16 @@ var builtins = map[string]builtin{ "int": {arity: 2, check: intRangeArgs, prep: func(a []string) callFn { lo, span := atoi(a[0]), atoi(a[1])-atoi(a[0])+1 return func(s *session, _ string, _ []string) string { return strconv.Itoa(lo + s.IntN(span)) } + }, number: func(token string, a []string) proven { + return printing(token, DataTypeInteger, bounded(float64(atoi(a[0])), float64(atoi(a[1])), true)) }}, "float": {arity: 3, check: floatArgs, prep: func(a []string) callFn { lo, hi, dp := atof(a[0]), atof(a[1]), atoi(a[2]) return func(s *session, _ string, _ []string) string { - return strconv.FormatFloat(lo+s.Float64()*(hi-lo), 'f', dp, 64) + return formatFloat(lo+s.Float64()*(hi-lo), dp) } + }, number: func(token string, a []string) proven { + return printedNumber(token, bounded(atof(a[0]), atof(a[1]), false), atoi(a[2])) }}, "iban": {arity: 1, check: ibanArg, prep: func(a []string) callFn { cc := a[0] @@ -69,6 +75,8 @@ var builtins = map[string]builtin{ return func(s *session, _ string, _ []string) string { return strconv.FormatUint(s.next(key), 10) } + }, number: func(token string, _ []string) proven { + return printing(token, DataTypeInteger, bounded(1, math.MaxInt64, true)) }}, } @@ -96,6 +104,15 @@ func chars(alphabet string) func([]string) callFn { const hexDigits = "0123456789abcdef" +// formatFloat prints v to dp decimals, -1 for the shortest form, and a zero unsigned. +func formatFloat(v float64, dp int) string { + s := strconv.FormatFloat(v, 'f', dp, 64) + if strings.HasPrefix(s, "-") && strings.Trim(s, "-0.") == "" { + return s[1:] + } + return s +} + // transforms are the builtins that rewrite one operand's value; they nest, so // {lowercase(ascii(x))} folds then lowers. var transforms = map[string]func(string) string{ diff --git a/builtins_test.go b/builtins_test.go index 7a63509..31b2f72 100644 --- a/builtins_test.go +++ b/builtins_test.go @@ -69,6 +69,9 @@ func TestBuiltinFloat(t *testing.T) { if got := mustRender(t, f, `"{float(1,2,3)}"`); !re.MatchString(got) { t.Fatalf("float(1,2,3) = %q, want d.ddd in [1,2]", got) } + if got := mustRender(t, f, `"{float(-1,1,0)}"`); got == "-0" { + t.Fatalf("float(-1,1,0) = %q, want a zero printed unsigned", got) + } } } diff --git a/calc.go b/calc.go index cd776b7..eb0c2c4 100644 --- a/calc.go +++ b/calc.go @@ -154,10 +154,12 @@ func calcText(n calcNode) string { return "?" } -// neverNumeric reports a node no render of which is a number: fixed text that does -// not parse, or a choice of only such items. text is one such render. +// neverNumeric reports a node no render of which is a number: a null, fixed text that +// does not parse, or a choice of only such items. text is one such render. func neverNumeric(n node) (text string, never bool) { switch n := n.(type) { + case *null: + return "", true case *template: if !n.fixed || n.repeat > 1 { return "", false @@ -191,15 +193,20 @@ func calcPrep(args []string) callFn { at[name] = i } placed := indexVars(expr, at) - dp := -1 - if len(args) == 2 { - dp = atoi(args[1]) - } + dp := calcDecimals(args) return func(_ *session, _ string, operands []string) string { - return strconv.FormatFloat(placed.eval(operands), 'f', dp, 64) + return formatFloat(placed.eval(operands), dp) } } +// calcDecimals is a calc's decimals count, or -1 for the shortest form. +func calcDecimals(args []string) int { + if len(args) == 2 { + return atoi(args[1]) + } + return -1 +} + // indexVars replaces each operand name with its position in the values expand reads. // Both sides take that order from calcVars, so they cannot drift. func indexVars(n calcNode, at map[string]int) calcNode { diff --git a/calc_test.go b/calc_test.go index 77adf77..7a8b7bd 100644 --- a/calc_test.go +++ b/calc_test.go @@ -36,6 +36,7 @@ func TestCalcAuto(t *testing.T) { `"{calc(10 / 3)}"`: "3.3333333333333335", `"{calc(6 / 2)}"`: "3", `"{calc(1 / 4)}"`: "0.25", + `"{calc(0 * -1)}"`: "0", } for tmpl, want := range cases { if got := mustRender(t, f, tmpl); got != want { @@ -52,6 +53,7 @@ func TestCalcDecimals(t *testing.T) { `"{calc(10 / 3, 2)}"`: "3.33", `"{calc(10 / 3, 0)}"`: "3", `"{calc(2 * 3, 2)}"`: "6.00", + `"{calc(-0.001, 2)}"`: "0.00", } for tmpl, want := range cases { if got := mustRender(t, f, tmpl); got != want { diff --git a/datatype.go b/datatype.go new file mode 100644 index 0000000..03bdb7f --- /dev/null +++ b/datatype.go @@ -0,0 +1,107 @@ +package fejkdata + +import ( + "errors" + "fmt" +) + +// DataType is what a record column holds, which decides how a record writes its value. +type DataType int + +// The datatypes a column declares with "datatype"; a column without one is a string. +const ( + DataTypeString DataType = iota + DataTypeInteger + DataTypeNumber + DataTypeBoolean +) + +var ( + dataTypeNames = [...]string{"string", "integer", "number", "boolean"} + dataTypeNouns = [...]string{"text", "an integer", "a number", "a boolean"} +) + +// String is the datatype as data spells it. +func (d DataType) String() string { + if d < 0 || int(d) >= len(dataTypeNames) { + return fmt.Sprintf("DataType(%d)", int(d)) + } + return dataTypeNames[d] +} + +// position is where a JSON value sits, which decides whether it may carry a datatype or +// be null. +type position int + +const ( + inFormat position = iota // rendered by a format, so neither + atTop // a category or an inline template, whose fields may be columns + inColumn // a column, or a choice item standing in for one +) + +// datatypeOf reads a template's "datatype" (default DataTypeString). +func datatypeOf(m map[string]any, pos position) (DataType, error) { + v, ok := m["datatype"] + if !ok { + return DataTypeString, nil + } + name, ok := v.(string) + if !ok { + return 0, fmt.Errorf("datatype must be a string, got %T", v) + } + if name == DataTypeString.String() { + return 0, fmt.Errorf("datatype %q is the default, so it has no effect; drop it", name) + } + for d := DataTypeInteger; d <= DataTypeBoolean; d++ { + if name != d.String() { + continue + } + if pos != inColumn { + return 0, errors.New("datatype only types a record column — a field of the top-level template — so it has no effect here") + } + return d, nil + } + return 0, fmt.Errorf(`datatype takes "integer", "number" or "boolean", got %q`, name) +} + +// columnDatatype is the datatype a column's items declare. They must agree, since a +// column holds one; a column only ever null is a string. +func columnDatatype(n node) (DataType, error) { + var items []*template + var collect func(node) + collect = func(n node) { + switch n := n.(type) { + case *choice: + for _, it := range n.items { + collect(it) + } + case *template: + items = append(items, n) + } + } + collect(n) + if len(items) == 0 { + return DataTypeString, nil + } + for _, t := range items[1:] { + if t.datatype != items[0].datatype { + return items[0].datatype, disagreement(items[0], t) + } + } + return items[0].datatype, nil +} + +// disagreement names the fix for two items of one column declaring different datatypes. +func disagreement(a, b *template) error { + typed, bare := a, b + if typed.datatype == DataTypeString { + typed, bare = b, a + } + switch { + case bare.datatype != DataTypeString: + return fmt.Errorf("its items declare %s and %s; a column holds one datatype", a.datatype, b.datatype) + case bare.fields == nil: // a JSON string; an object, which may carry a weight, has a fields map + return fmt.Errorf(`item %q declares no datatype, and a column holds one; write it as {"format":%q,"datatype":%q}`, bare.format, bare.format, typed.datatype) + } + return fmt.Errorf(`an item declares no datatype beside one declaring %s; a column holds one, so give it "datatype": %q`, typed.datatype, typed.datatype) +} diff --git a/datatype_test.go b/datatype_test.go new file mode 100644 index 0000000..a537991 --- /dev/null +++ b/datatype_test.go @@ -0,0 +1,187 @@ +package fejkdata + +import ( + "encoding/json" + "regexp" + "slices" + "strconv" + "strings" + "testing" +) + +func TestDatatypeAndNullSitOnlyInAColumn(t *testing.T) { + for _, src := range []string{ + `{"format":"","age":{"format":"{int(18,99)}","datatype":"integer"}}`, + `{"format":"","n":{"format":"42","datatype":"integer"}}`, + `{"format":"","gone":null}`, + `{"format":"","middle":[null,"Ann","Eva"]}`, + `{"format":"","age":[null,{"format":"{int(18,99)}","datatype":"integer","weight":9}]}`, + `{"format":"","pick":[[null,"a"],"b"]}`, + } { + if _, err := compile(parse(t, src)); err != nil { + t.Errorf("compile(%s) = %v, want a column to take a datatype and null", src, err) + } + } + for src, want := range map[string]string{ + `{"format":"","n":{"format":"1","datatype":"int"}}`: `datatype takes "integer", "number" or "boolean", got "int"`, + `{"format":"","n":{"format":"1","datatype":1}}`: "datatype must be a string", + `{"format":"{int(1,9)}","datatype":"integer"}`: "datatype only types a record column", + `[{"format":"1","datatype":"integer"},"x"]`: "datatype only types a record column", + `{"format":"{p}","p":{"format":"{n}","n":{"format":"1","datatype":"integer"}}}`: "datatype only types a record column", + `{"format":"{n}","repeat":2,"n":{"format":"1","datatype":"integer"}}`: "datatype only types a record column", + `null`: `so write ""`, + `{"format":"{p}","p":{"format":"{x}","x":[null,"a"]}}`: `so write ""`, + `{"format":"","c":[{"format":"1","datatype":"integer"},"x"]}`: `write it as {"format":"x","datatype":"integer"}`, + `{"format":"","c":[{"format":"1","datatype":"integer"},{"format":"2","weight":3}]}`: `give it "datatype": "integer"`, + `{"format":"","c":[{"format":"1","datatype":"integer"},{"format":"true","datatype":"boolean"}]}`: "a column holds one datatype", + } { + if _, err := compile(parse(t, src)); err == nil || !strings.Contains(err.Error(), want) { + t.Errorf("compile(%s) = %v, want an error containing %q", src, err, want) + } + } +} + +func TestDatatypeRejectsAValueItsTypeRejects(t *testing.T) { + tree := map[string]string{ + "cat": `[{"format":"{code}","code":"200"},{"format":"{code}","code":"2x"}]`, + "src": `{"format":"","score":[null,{"format":"{int(1,9)}","datatype":"integer"}]}`, + } + for _, c := range []struct{ name, column, want string }{ + {"a sample with leading zeros", `{"format":"{digits(3)}","datatype":"integer"}`, "{digits(3)} prints text, not an integer"}, + {"a fraction", `{"format":"{v}","v":["1","1.5"],"datatype":"integer"}`, `"1.5" is not an integer`}, + {"past int64", `{"format":"9223372036854775808","datatype":"integer"}`, "past the int64 range"}, + {"a float sample", `{"format":"{float(0,1,2)}","datatype":"integer"}`, "{float(0,1,2)} prints a number, not an integer"}, + {"composed digits", `{"format":"1{digits(2)}","datatype":"integer"}`, "is not one value"}, + {"a sign before a sample", `{"format":"-{int(1,9)}","datatype":"integer"}`, "is not one value"}, + {"a repeat", `{"format":"{int(1,9)}","repeat":2,"separator":",","datatype":"integer"}`, "carries a repeat"}, + {"through a reference", `{"format":"{/cat.code}","datatype":"integer"}`, `"2x" is not an integer`}, + {"a null through a reference", `{"format":"{/src.score}","datatype":"integer"}`, "reads a null"}, + {"a bare dot", `{"format":".5","datatype":"number"}`, `".5" is not a number`}, + {"a plus sign", `{"format":"+1","datatype":"number"}`, `"+1" is not a number`}, + {"a text sample", `{"format":"{hex(4)}","datatype":"number"}`, "{hex(4)} prints text, not a number"}, + {"a capital", `{"format":"{b}","b":["true","True"],"datatype":"boolean"}`, `"True" is not a boolean`}, + {"a transform", `{"format":"{lowercase(b)}","b":["TRUE","FALSE"],"datatype":"boolean"}`, "{lowercase(b)} rewrites text"}, + {"a number as a boolean", `{"format":"{int(0,1)}","datatype":"boolean"}`, "{int(0,1)} prints an integer, not a boolean"}, + {"an operand that is not always a number", `{"format":"{calc(a * 2)}","a":["1","x"],"datatype":"number"}`, `operand "a": "x" is not a number`}, + {"a divisor that can be zero", `{"format":"{calc(a / b)}","a":"{int(1,9)}","b":"{int(0,9)}","datatype":"number"}`, "divides by b, which is not proven nonzero"}, + {"an overflow", `{"format":"{calc(a * a)}","a":"{digits(200)}","datatype":"number"}`, "is not proven within 1e300"}, + {"a division in an integer column", `{"format":"{calc(a / b)}","a":"{int(1,9)}","b":"{int(1,9)}","datatype":"integer"}`, "prints a number, not an integer"}, + {"a composed operand", `{"format":"{calc(n * 2)}","n":"{int(1,99)}.{digits(2)}","datatype":"number"}`, "{seq()}, {digits()} or {calc()}"}, + {"a divisor that prints as zero", `{"format":"{calc(1 / b, 2)}","b":"{float(4.9999999999999994e-79,5e-79,78)}","datatype":"number"}`, "divides by b, which is not proven nonzero"}, + {"a divisor that rounds to zero", `{"format":"{calc(a / b)}","a":"{int(1,9)}","b":"{float(0.1,1,0)}","datatype":"number"}`, "divides by b, which is not proven nonzero"}, + {"a zero among a divisor's literals", `{"format":"{calc(a / b)}","a":"{int(1,9)}","b":["0","5"],"datatype":"number"}`, "divides by b, which is not proven nonzero"}, + {"a negated divisor crossing zero", `{"format":"{calc(a / (-b + 10))}","a":"{int(1,9)}","b":"{int(1,20)}","datatype":"number"}`, "which is not proven nonzero"}, + {"a subtracted divisor crossing zero", `{"format":"{calc(a / (10 - b))}","a":"{int(1,9)}","b":"{int(1,20)}","datatype":"number"}`, "which is not proven nonzero"}, + {"a quotient past the limit", `{"format":"{calc(a / b / b)}","a":"{digits(300)}","b":"{float(0.000001,1,6)}","datatype":"number"}`, "is not proven within 1e300"}, + {"an operand past the limit", `{"format":"{calc(a)}","a":"{digits(400)}","datatype":"number"}`, "is not proven within 1e300"}, + {"a whole calc past int64", `{"format":"{calc(a * 2)}","a":"{seq()}","datatype":"integer"}`, "{calc(a * 2)} is not proven within int64"}, + {"a whole float past int64", `{"format":"{float(0,1e19,0)}","datatype":"integer"}`, "{float(0,1e19,0)} is not proven within int64"}, + {"a signed zero integer", `{"format":"-0","datatype":"integer"}`, `"-0" is zero written with a sign; write "0"`}, + {"a signed zero number", `{"format":"-0.00","datatype":"number"}`, `"-0.00" is zero written with a sign; write "0.00"`}, + } { + row := `{"format":"","col":` + c.column + `}` + files := map[string]string{"row": row} + for name, body := range tree { + files[name] = body + } + _, err := New(WithoutShippedData(), WithDataPath(writeData(t, files))) + if err == nil || !strings.Contains(err.Error(), c.want) { + t.Errorf("%s: New = %v, want an error containing %q", c.name, err, c.want) + } + f := newGenerator(t, writeData(t, tree)) + if _, err := f.NewTemplate(row); err == nil || !strings.Contains(err.Error(), c.want) { + t.Errorf("%s: NewTemplate = %v, want the inline template refused the same way", c.name, err) + } + } +} + +var jsonInteger = regexp.MustCompile(`^(0|-?[1-9][0-9]*)$`) + +func TestDatatypeAcceptsAColumnThatAlwaysParses(t *testing.T) { + cat := `[{"format":"{code}","code":"200"},{"format":"{code}","code":"404"}]` + for _, column := range []string{ + `{"format":"{int(1,99)}","datatype":"integer"}`, + `{"format":"{int(-9,-1)}","datatype":"integer"}`, + `{"format":"{seq()}","datatype":"integer"}`, + `{"format":"{/cat.code}","datatype":"integer"}`, + `{"format":"{a|b}","a":"1","b":"{int(5,9)}","datatype":"integer"}`, + `{"format":"{float(-1.5,9.5,0)}","datatype":"integer"}`, + `{"format":"{float(-1,1,0)}","datatype":"integer"}`, + `{"format":"{calc(a * b)}","a":"{int(-9,-1)}","b":"{int(0,9)}","datatype":"integer"}`, + `{"format":"{calc(x + 1)}","x":"{float(0,9,0)}","datatype":"integer"}`, + `{"format":"{float(-1,1,2)}","datatype":"number"}`, + `{"format":"{v}","v":["1","2.5","6.022e23"],"datatype":"number"}`, + `{"format":"{v}","v":["-1e-400","0.5"],"datatype":"number"}`, + `{"format":"{b}","b":["true","false"],"datatype":"boolean"}`, + `{"format":"{calc(net * qty, 2)}","net":["19.99","5.00"],"qty":["3","7"],"datatype":"number"}`, + `{"format":"{calc(a + b)}","a":"{int(1,9)}","b":"{int(-9,9)}","datatype":"integer"}`, + `{"format":"{calc(x / (b - c), 2)}","x":"{int(1,9)}","b":"{int(10,20)}","c":"{int(1,5)}","datatype":"number"}`, + `{"format":"{calc(x * x * x * x, 2)}","x":"{float(0,1,80)}","datatype":"number"}`, + `{"format":"{calc(a / (b + 1), 2)}","a":"{int(1,9)}","b":"{digits(2)}","datatype":"number"}`, + `{"format":"{calc(a / b, 0)}","a":"{int(1,9)}","b":"{int(1,9)}","datatype":"integer"}`, + `{"format":"{calc(sub * 1.25, 2)}","sub":{"format":"{calc(a * b)}","a":"{int(1,9)}","b":"{float(0,5,2)}"},"datatype":"number"}`, + } { + row := `{"format":"","col":` + column + `}` + f, err := New(WithoutShippedData(), WithDataPath(writeData(t, map[string]string{"cat": cat, "row": row})), WithSeed(1)) + if err != nil { + t.Errorf("%s: New = %v, want it loaded", column, err) + continue + } + for i := 0; i < 200; i++ { + r, err := f.FakeRecord("row") + if err != nil { + t.Fatal(err) + } + var m map[string]any + if err := json.Unmarshal([]byte(r.JSON()), &m); err != nil { + t.Errorf("%s: JSON() = %s is not JSON: %v", column, r.JSON(), err) + break + } + c := r.Columns()[0] + _, isBool := m["col"].(bool) + _, isNumber := m["col"].(float64) + _, int64Err := strconv.ParseInt(c.Value, 10, 64) + if c.DataType == DataTypeBoolean && !isBool || c.DataType != DataTypeBoolean && !isNumber || c.DataType == DataTypeInteger && (!jsonInteger.MatchString(c.Value) || int64Err != nil) { + t.Errorf("%s: column %+v written as %s, want its datatype", column, c, r.JSON()) + break + } + } + } +} + +func TestNullColumn(t *testing.T) { + dir := writeData(t, map[string]string{ + "row": `{"format":"[{middle}]","gone":null,"middle":[null,"Ann"],"score":[null,{"format":"{int(1,9)}","datatype":"integer"}]}`, + }) + f := newGenerator(t, dir, WithSeed(1)) + drew := map[bool]bool{} + for i := 0; i < 100; i++ { + r, err := f.FakeRecord("row") + if err != nil { + t.Fatal(err) + } + gone, middle, score := r.Columns()[0], r.Columns()[1], r.Columns()[2] + if !gone.Null || gone.Value != "" || gone.DataType != DataTypeString { + t.Fatalf("gone = %+v, want a null string column every draw", gone) + } + if middle.Null == (middle.Value == "Ann") { + t.Fatalf("middle = %+v, want null or Ann", middle) + } + if score.DataType != DataTypeInteger { + t.Fatalf("score = %+v, want the integer its non-null item declares, null or not", score) + } + drew[middle.Null] = true + } + if len(drew) != 2 { + t.Errorf("middle drew only null=%v in 100 records, want both", drew) + } + if v := fake(t, f, "row"); v != "[]" && v != "[Ann]" { + t.Errorf("Fake(row) = %q, want a null to render as \"\"", v) + } + if v := fake(t, f, "row.gone"); v != "" { + t.Errorf("Fake(row.gone) = %q, want \"\"", v) + } + if !slices.Contains(f.List(), "row.gone") { + t.Errorf("List() = %v, want the null column row.gone, which Fake accepts", f.List()) + } +} diff --git a/fejkdata.go b/fejkdata.go index c558d9b..35cd73e 100644 --- a/fejkdata.go +++ b/fejkdata.go @@ -162,6 +162,8 @@ func paths(n node) []string { } } return out + case *null: + return []string{""} case *choice: out := []string{""} for p := range n.shared { diff --git a/graph.go b/graph.go index cf5cd5b..f8361e1 100644 --- a/graph.go +++ b/graph.go @@ -186,7 +186,10 @@ func checkScope(s nodeScope) error { if err := s(func(path string, n node) error { return repeatCheck(path, n, mem) }); err != nil { return err } - return s(heldCheck) + if err := s(heldCheck); err != nil { + return err + } + return s((&valueProof{}).checkDatatype) } type reachMemo map[node]int diff --git a/node.go b/node.go index 384a2c6..cef1bc7 100644 --- a/node.go +++ b/node.go @@ -32,6 +32,11 @@ type choice struct { func (*choice) isNode() {} +// null is a column's missing value, rendered ""; sized so two nulls are two map keys. +type null struct{ _ byte } + +func (*null) isNode() {} + // template renders a format string, substituting {tokens} from fields. A bare // JSON string is a template with no fields. repeat (default 1) renders that format // that many times and joins the results with separator (default ""), each render @@ -41,6 +46,7 @@ type template struct { fields map[string]node repeat int separator string + datatype DataType ops []op // format compiled once (see compileOps); what expand walks grow int // minimum output size, to size the render buffer fixed bool // no op varies, so every render is lit @@ -66,26 +72,37 @@ func (t *template) field(seg string) (node, bool) { return n, ok } -// compile converts parsed JSON into a node tree, validating structure up front. -// Only a choice's items carry a weight, so one here would be inert whatever its type. +// compile converts parsed JSON — a category or an inline template — into a node tree, +// validating structure up front. func compile(v any) (node, error) { + return compileAt(v, atTop) +} + +// compileAt compiles a node that is no choice's item. Only a choice's items carry a +// weight, so one here would be inert whatever its type. +func compileAt(v any, pos position) (node, error) { if m, ok := v.(map[string]any); ok { if _, weighted := m["weight"]; weighted { return nil, fmt.Errorf("weight only skews a choice's items, so it has no effect here; it is an option and can never be a field") } } - return compileItem(v) + return compileItem(v, pos) } // compileItem compiles one node, allowing the weight a choice item may carry. -func compileItem(v any) (node, error) { +func compileItem(v any, pos position) (node, error) { switch v := v.(type) { case string: return compileString(v) case []any: - return compileChoice(v) + return compileChoice(v, pos) case map[string]any: - return compileTemplate(v) + return compileTemplate(v, pos) + case nil: + if pos != inColumn { + return nil, fmt.Errorf(`null is a record column's value; here it only renders "", so write ""`) + } + return &null{}, nil default: return nil, fmt.Errorf("a template value must be a string, a list or an object, not %s", jsonKind(v)) } @@ -99,8 +116,6 @@ func jsonKind(v any) string { return "a number" case bool: return "a boolean" - case nil: - return "null" } return fmt.Sprintf("%T", v) } @@ -136,7 +151,11 @@ func (t *template) compileFormat() error { return checkNoRepeatedRead(t.format, c, t.refs) } -func compileChoice(items []any) (node, error) { +func compileChoice(items []any, pos position) (node, error) { + itemPos := inFormat + if pos == inColumn { + itemPos = inColumn + } if len(items) == 0 { return nil, fmt.Errorf("empty choice") } @@ -160,7 +179,7 @@ func compileChoice(items []any) (node, error) { } total += w cum[i] = total - n, err := compileItem(raw) + n, err := compileItem(raw, itemPos) if err != nil { return nil, err } @@ -196,6 +215,9 @@ func checkNoRepeatedItem(items []any) error { if s, isString := raw.(string); isString { return fmt.Errorf("choice item %q is repeated; skew the odds with a weight instead: { \"format\": %q, \"weight\": 2 }", s, s) } + if raw == nil { + return fmt.Errorf("choice item %d repeats null; a null takes no weight, so weight the other items instead", i) + } return fmt.Errorf("choice item %d repeats item %d; skew the odds with a weight on one of them instead", i, j) } seen[string(key)] = i @@ -203,22 +225,26 @@ func checkNoRepeatedItem(items []any) error { return nil } -func compileTemplate(m map[string]any) (node, error) { - o, err := readOptions(m) +func compileTemplate(m map[string]any, pos position) (node, error) { + o, err := readOptions(m, pos) if err != nil { return nil, err } - fields, err := compileFields(m) + fieldPos := inFormat + if pos == atTop && projectsColumns(o.repeat) { + fieldPos = inColumn + } + fields, err := compileFields(m, fieldPos) if err != nil { return nil, err } - if len(fields) == 0 && o.repeat == 1 && !o.weighted { + if len(fields) == 0 && o.repeat == 1 && !o.weighted && o.datatype == DataTypeString { return nil, fmt.Errorf("an object holding only a format is a string; write %q", o.format) } if err := checkTokens(o.format, fields); err != nil { return nil, err } - t := &template{format: o.format, fields: fields, repeat: o.repeat, separator: o.separator} + t := &template{format: o.format, fields: fields, repeat: o.repeat, separator: o.separator, datatype: o.datatype} if err := t.compileFormat(); err != nil { return nil, err } @@ -227,13 +253,14 @@ func compileTemplate(m map[string]any) (node, error) { // templateOptions is what a template object's option keys say. type templateOptions struct { + datatype DataType format string repeat int separator string weighted bool } -func readOptions(m map[string]any) (templateOptions, error) { +func readOptions(m map[string]any, pos position) (templateOptions, error) { var o templateOptions format, ok := m["format"].(string) if !ok { @@ -245,6 +272,9 @@ func readOptions(m map[string]any) (templateOptions, error) { return o, err } o.repeat = repeat + if o.datatype, err = datatypeOf(m, pos); err != nil { + return o, err + } if sv, ok := m["separator"]; ok { if o.separator, ok = sv.(string); !ok { return o, fmt.Errorf("separator must be a string, got %T", sv) @@ -262,7 +292,7 @@ func readOptions(m map[string]any) (templateOptions, error) { // compileFields compiles every non-option key of a template object, in name order // so which of several bad fields is reported does not vary. -func compileFields(m map[string]any) (map[string]node, error) { +func compileFields(m map[string]any, pos position) (map[string]node, error) { fields := make(map[string]node, len(m)) keys := make([]string, 0, len(m)) for k := range m { @@ -276,7 +306,10 @@ func compileFields(m map[string]any) (map[string]node, error) { if err := checkName(k); err != nil { return nil, fmt.Errorf("field %w", err) } - n, err := compile(m[k]) + n, err := compileAt(m[k], pos) + if err == nil && pos == inColumn { + _, err = columnDatatype(n) + } if err != nil { return nil, fmt.Errorf("field %q: %w", k, err) } @@ -360,10 +393,10 @@ func checkName(name string) error { } // isOption reports whether a template key configures the node instead of naming a -// field. These four names can never be fields. +// field. These names can never be fields. func isOption(name string) bool { switch name { - case "format", "repeat", "separator", "weight": + case "datatype", "format", "repeat", "separator", "weight": return true } return false diff --git a/one_spelling_test.go b/one_spelling_test.go index d5bbd8b..8253b1c 100644 --- a/one_spelling_test.go +++ b/one_spelling_test.go @@ -18,6 +18,7 @@ func TestRepeatedChoiceItemIsRejected(t *testing.T) { `["a", "a", "b"]`: `{ "format": "a", "weight": 2 }`, `[{"format":"{x}","x":"1"},{"format":"{x}","x":"1"}]`: "repeats item", `{"format":"{w}","w":["", "", "x"]}`: `{ "format": "", "weight": 2 }`, + `{"format":"","w":[null,null,"a"]}`: "a null takes no weight", } { if _, err := compile(parse(t, src)); err == nil || !strings.Contains(err.Error(), want) { t.Errorf("compile(%s) = %v, want an error naming %s", src, err, want) @@ -30,12 +31,13 @@ func TestRepeatedChoiceItemIsRejected(t *testing.T) { func TestInertObjectIsRejected(t *testing.T) { for src, want := range map[string]string{ - `{"format":"Malmö"}`: `write "Malmö"`, - `{"format":"{digits(3)}"}`: `write "{digits(3)}"`, - `[{"format":"a","weight":1},"b"]`: "weight 1", - `{"format":"{x}","x":"v","repeat":1}`: "repeat 1", - `{"format":"{x}","x":"v","separator":","}`: "separator", - `{"format":"{x}","x":"v","repeat":2,"separator":""}`: "default", + `{"format":"Malmö"}`: `write "Malmö"`, + `{"format":"{digits(3)}"}`: `write "{digits(3)}"`, + `[{"format":"a","weight":1},"b"]`: "weight 1", + `{"format":"{x}","x":"v","repeat":1}`: "repeat 1", + `{"format":"{x}","x":"v","separator":","}`: "separator", + `{"format":"{x}","x":"v","repeat":2,"separator":""}`: "default", + `{"format":"","n":{"format":"1","datatype":"string"}}`: `datatype "string" is the default`, } { if _, err := compile(parse(t, src)); err == nil || !strings.Contains(err.Error(), want) { t.Errorf("compile(%s) = %v, want an error mentioning %s", src, err, want) diff --git a/path.go b/path.go index ee9b1c5..687d9b9 100644 --- a/path.go +++ b/path.go @@ -60,7 +60,7 @@ func walkPath(n node, tail []string, w pathWalk) error { } return nil } - return fmt.Errorf("cannot descend into %T at %q", n, tail[0]) + return fmt.Errorf("no field %q", tail[0]) } // carriedByAll is the choice rule a path that must resolve on every call obeys: diff --git a/perf_test.go b/perf_test.go index a3b750d..0ee6dcb 100644 --- a/perf_test.go +++ b/perf_test.go @@ -68,6 +68,10 @@ func TestNoRecordAllocRegression(t *testing.T) { if allocs := testing.AllocsPerRun(10000, func() { f.FakeRecord("x") }); allocs > base*1.10 { t.Errorf("%s: %.1f allocs/op regressed past %.1f (baseline %.1f + 10%%); a record fence running per draw is the usual cause", s.name, allocs, base*1.10, base) } + r, _ := f.FakeRecord("x") + if allocs := testing.AllocsPerRun(10000, func() { _ = r.CSVLine() }); allocs > 2 { + t.Errorf("%s: CSVLine() makes %.1f allocs/op, want 2: the fields slice and the joined line", s.name, allocs) + } } } diff --git a/readme_test.go b/readme_test.go index 18790e9..c628709 100644 --- a/readme_test.go +++ b/readme_test.go @@ -1,6 +1,7 @@ package fejkdata import ( + "encoding/json" "os" "regexp" "strings" @@ -87,3 +88,43 @@ func TestReadmeSQLExampleOutput(t *testing.T) { t.Errorf("README SQL example with seed 1 = %q, README prints %q", got, want[1]) } } + +func TestReadmeDatatypeExample(t *testing.T) { + src := readme(t) + i := strings.Index(src, "### Datatype") + if i < 0 { + t.Fatal("README lost the Datatype section") + } + block := jsonBlock.FindStringSubmatch(src[i:]) + f, err := New(WithDataPath(writeData(t, map[string]string{"order": block[1]})), WithSeed(1)) + if err != nil { + t.Fatal(err) + } + r, err := f.FakeRecord("order") + if err != nil { + t.Fatal(err) + } + var m map[string]any + if err := json.Unmarshal([]byte(r.JSON()), &m); err != nil { + t.Fatalf("JSON() = %s: %v", r.JSON(), err) + } + shown := map[DataType]bool{} + for _, c := range r.Columns() { + shown[c.DataType] = true + var ok bool + switch c.DataType { + case DataTypeBoolean: + _, ok = m[c.Name].(bool) + case DataTypeInteger, DataTypeNumber: + _, ok = m[c.Name].(float64) + default: + _, ok = m[c.Name].(string) + } + if !ok { + t.Errorf("column %q, datatype %s, written as %s", c.Name, c.DataType, r.JSON()) + } + } + if !shown[DataTypeInteger] || !shown[DataTypeNumber] || !shown[DataTypeBoolean] { + t.Errorf("README Datatype example shows %v, want an integer, a number and a boolean column", shown) + } +} diff --git a/record.go b/record.go index 803e041..84bf8dd 100644 --- a/record.go +++ b/record.go @@ -1,18 +1,23 @@ package fejkdata import ( - "encoding/csv" "encoding/json" "errors" "fmt" "sort" "strings" + "unicode" + "unicode/utf8" ) -// Column is one rendered column of a record. +// Column is one rendered column of a record. Value is the rendered text, which a +// serializer quotes for DataTypeString and writes bare for any other datatype; a Null +// column has no Value. type Column struct { - Name string - Value string + Name string + DataType DataType + Value string + Null bool } // Record is one record rendered from a template: every direct field is a column, @@ -28,70 +33,90 @@ func (r *Record) Columns() []Column { return append([]Column(nil), r.columns...) } -// JSON renders the record as one JSON object, every column a string. +// JSON renders the record as one JSON object. func (r *Record) JSON() string { - m := make(map[string]string, len(r.columns)) - for _, c := range r.columns { - m[c.Name] = c.Value + var b strings.Builder + b.WriteByte('{') + for i, c := range r.columns { + if i > 0 { + b.WriteByte(',') + } + b.WriteString(jsonString(c.Name)) + b.WriteByte(':') + b.WriteString(literal(c, jsonString, "null")) } - b, _ := json.Marshal(m) + b.WriteByte('}') + return b.String() +} + +func jsonString(s string) string { + b, _ := json.Marshal(s) return string(b) } // CSVHeader renders the column names as one CSV header line. func (r *Record) CSVHeader() string { - return csvLine(r.names()) + fields := make([]string, len(r.columns)) + for i, c := range r.columns { + fields[i] = csvField(c.Name) + } + return strings.Join(fields, ",") } -// CSVLine renders the column values as one CSV row. +// CSVLine renders the column values as one CSV row: a null column an empty field and an +// empty string "", the convention PostgreSQL's COPY reads a null by. A record of one null +// column is a blank line, which COPY reads as null and most CSV readers skip. func (r *Record) CSVLine() string { - return csvLine(r.values()) -} - -func (r *Record) names() []string { - out := make([]string, len(r.columns)) + fields := make([]string, len(r.columns)) for i, c := range r.columns { - out[i] = c.Name + fields[i] = literal(c, csvField, "") } - return out + return strings.Join(fields, ",") } -func (r *Record) values() []string { - out := make([]string, len(r.columns)) - for i, c := range r.columns { - out[i] = c.Value +// csvField quotes a field where encoding/csv would, and an empty one too. +func csvField(s string) string { + first, _ := utf8.DecodeRuneInString(s) + switch { + case s == "": + return `""` + case s == `\.` || strings.ContainsAny(s, "\",\r\n") || unicode.IsSpace(first): + return `"` + strings.ReplaceAll(s, `"`, `""`) + `"` } - return out + return s } -func csvLine(cols []string) string { - var b strings.Builder - w := csv.NewWriter(&b) - _ = w.Write(cols) - w.Flush() - line := strings.TrimSuffix(b.String(), "\n") - if line == "" { - return `""` // a blank line is a row every CSV reader drops - } - return line -} - -// SQLInsert renders the record as one INSERT statement into table: identifiers in -// ANSI double quotes, every value a single-quoted string literal. +// SQLInsert renders the record as one INSERT statement into table, identifiers in ANSI +// double quotes. func (r *Record) SQLInsert(table string) string { cols := make([]string, len(r.columns)) vals := make([]string, len(r.columns)) for i, c := range r.columns { cols[i] = quoteIdent(c.Name) - vals[i] = "'" + strings.ReplaceAll(c.Value, "'", "''") + "'" + vals[i] = literal(c, sqlString, "NULL") } return fmt.Sprintf("INSERT INTO %s (%s) VALUES (%s);", quoteIdent(table), strings.Join(cols, ", "), strings.Join(vals, ", ")) } +func sqlString(s string) string { + return "'" + strings.ReplaceAll(s, "'", "''") + "'" +} + func quoteIdent(s string) string { return `"` + strings.ReplaceAll(s, `"`, `""`) + `"` } +// literal writes a string column quoted, a proven typed one bare, and a null as nullText. +func literal(c Column, quote func(string) string, nullText string) string { + switch { + case c.Null: + return nullText + case c.DataType == DataTypeString: + return quote(c.Value) + } + return c.Value +} + // FakeRecord renders a path as one record: the template it names, with each direct // field drawn as a column. Only a category-level template is a record — a path // that descends into a field, or that names a folder or a choice, is an error. @@ -116,7 +141,7 @@ func (f *Generator) FakeRecord(path string) (*Record, error) { // columns, or why it is not a record. type recordShape struct { t *template - columns []string + columns []Column err error } @@ -140,7 +165,7 @@ func (f *Generator) recordShapeOf(n node) recordShape { type RecordTemplate struct { g *Generator t *template - columns []string + columns []Column } // Fake renders the record with one draw. @@ -175,24 +200,33 @@ func (f *Generator) FakeRecordTemplate(input string) (*Record, error) { // recordOf is the fence both record entry points pass. The columns come back with // the template, fixed for every draw the caller goes on to make. -func recordOf(n node) (*template, []string, error) { +func recordOf(n node) (*template, []Column, error) { t, ok := n.(*template) if !ok { return nil, nil, errors.New("names a choice, not a template; a record is a template whose fields are its columns") } - if t.repeat != 1 { + if !projectsColumns(t.repeat) { return nil, nil, fmt.Errorf("carries repeat %d, which composes its format into one string; a record projects columns instead — drop the repeat and render the record again for more rows", t.repeat) } - columns := recordColumns(t) - if len(columns) == 0 { + names := recordColumns(t) + if len(names) == 0 { return nil, nil, errors.New("has no fields, so no columns") } - if err := checkColumnRefs(t, columns); err != nil { + if err := checkColumnRefs(t, names); err != nil { return nil, nil, err } + columns := make([]Column, len(names)) + for i, name := range names { + datatype, _ := columnDatatype(t.fields[name]) // compile refused a column whose items disagree + columns[i] = Column{Name: name, DataType: datatype} + } return t, columns, nil } +// projectsColumns reports whether a category or inline template with this repeat is a +// record, its fields the columns; a repeat composes the format into one string instead. +func projectsColumns(repeat int) bool { return repeat == 1 } + // checkColumnRefs rejects the reference reads a record's shared draw cannot answer // for: one column rendering a level another reads a path into, and a column // reading the record back through its own path. @@ -262,11 +296,16 @@ func columnRefs(t *template, columns []string) ([]columnRef, error) { // renderRecord draws each column once, in the name order recordOf fixed, over one // reference scope shared across them. -func renderRecord(s *session, t *template, columns []string) *Record { +func renderRecord(s *session, t *template, columns []Column) *Record { scope := &draws{variant: map[string]node{}, value: map[string]string{}} - r := &Record{columns: make([]Column, len(columns))} - for i, name := range columns { - r.columns[i] = Column{Name: name, Value: render(s, t.fields[name], scope)} + r := &Record{columns: append([]Column(nil), columns...)} + for i := range r.columns { + n := drawn(s, t.fields[r.columns[i].Name]) + if _, isNull := n.(*null); isNull { + r.columns[i].Null = true + } else { + r.columns[i].Value = render(s, n, scope) + } } return r } diff --git a/record_test.go b/record_test.go index ee2f754..0301804 100644 --- a/record_test.go +++ b/record_test.go @@ -3,6 +3,7 @@ package fejkdata import ( "encoding/csv" "encoding/json" + "reflect" "strings" "testing" ) @@ -130,6 +131,16 @@ func TestRecordCSV(t *testing.T) { if len(seen) != 4 { t.Fatalf("round-tripped %d of the 4 values; the comma, quote and newline shapes must each survive", len(seen)) } + for value, want := range map[string]string{`\.`: `"\."`, " x": `" x"`, " x": "\" x\"", "a\rb": "\"a\rb\""} { + body, _ := json.Marshal(map[string]string{"format": "", "v": value}) + r, err := newGenerator(t, writeData(t, map[string]string{"q": string(body)})).FakeRecord("q") + if err != nil { + t.Fatal(err) + } + if got := r.CSVLine(); got != want { + t.Errorf("CSVLine() of %q = %q, want %q, quoted where encoding/csv quotes", value, got, want) + } + } } func TestRecordCSVEmptyValueStaysARow(t *testing.T) { @@ -148,6 +159,46 @@ func TestRecordCSVEmptyValueStaysARow(t *testing.T) { } } +func TestRecordWritesTypedAndNullColumns(t *testing.T) { + dir := writeData(t, map[string]string{ + "row": `{"format":"","age":{"format":"42","datatype":"integer"},"gone":null,"name":"O'Brien","nick":"","paid":{"format":"true","datatype":"boolean"},"price":{"format":"19.99","datatype":"number"}}`, + }) + r, err := newGenerator(t, dir, WithSeed(1)).FakeRecord("row") + if err != nil { + t.Fatal(err) + } + want := []Column{ + {Name: "age", DataType: DataTypeInteger, Value: "42"}, + {Name: "gone", Null: true}, + {Name: "name", Value: "O'Brien"}, + {Name: "nick"}, + {Name: "paid", DataType: DataTypeBoolean, Value: "true"}, + {Name: "price", DataType: DataTypeNumber, Value: "19.99"}, + } + if got := r.Columns(); !reflect.DeepEqual(got, want) { + t.Errorf("Columns() = %+v, want %+v", got, want) + } + if got, want := r.JSON(), `{"age":42,"gone":null,"name":"O'Brien","nick":"","paid":true,"price":19.99}`; got != want { + t.Errorf("JSON() = %s, want %s", got, want) + } + if got, want := r.SQLInsert("t"), `INSERT INTO "t" ("age", "gone", "name", "nick", "paid", "price") VALUES (42, NULL, 'O''Brien', '', true, 19.99);`; got != want { + t.Errorf("SQLInsert() = %s, want %s", got, want) + } + if got, want := r.CSVLine(), `42,,O'Brien,"",true,19.99`; got != want { + t.Errorf("CSVLine() = %s, want %s: null an unquoted empty field, an empty string quoted", got, want) + } + if got := DataTypeNumber.String(); got != "number" { + t.Errorf("DataTypeNumber.String() = %q, want the data's spelling", got) + } + lone, err := newGenerator(t, writeData(t, map[string]string{"gone": `{"format":"","note":null}`})).FakeRecord("gone") + if err != nil { + t.Fatal(err) + } + if got := lone.CSVLine(); got != "" { + t.Errorf("a lone null column wrote CSVLine() = %q, want the blank line PostgreSQL's COPY reads as null", got) + } +} + func TestRecordRejectsOverlappingReferenceColumns(t *testing.T) { cat := `{"format":"","a":[{"format":"A={b}","b":"1"},{"format":"A={b}","b":"2"}]}` for _, c := range []struct{ name, row string }{ diff --git a/render.go b/render.go index 7a7af3a..330ed0f 100644 --- a/render.go +++ b/render.go @@ -56,6 +56,8 @@ func render(s *session, n node, refScope *draws) string { switch n := n.(type) { case *choice: return render(s, pick(s, n), refScope) + case *null: + return "" case *template: if n.repeat == 1 { if n.fixed { diff --git a/template.go b/template.go index 7ef9336..45bfcda 100644 --- a/template.go +++ b/template.go @@ -72,6 +72,9 @@ type builtin struct { // operands names the fields the call reads, which expand renders for it; nil // for a builtin that reads none. operands func(args []string) []string + // number proves what a call prints, token its body: the bounds of its number and the + // datatype of its text; nil for a builtin whose text is no number. + number func(token string, args []string) proven } // funcCall splits a "{token}" body shaped name(args) into its parts; ok is false diff --git a/todo.md b/todo.md index 0c2b28f..f252672 100644 --- a/todo.md +++ b/todo.md @@ -6,11 +6,6 @@ The record API lands first, so the data update can use it. ### Record API -- Typed columns — a column declares its type, so `json` writes `42` rather than - `"42"` and `sql` an unquoted literal: string, integer, number, boolean, and a - way to write null. A template that can render a value its type rejects is a - load error. The option key is reserved from then on, so a common column name - like `type` is a poor pick. - Struct-filling — fill a Go struct from `fake:"…"` tags holding a path or an inline template, for parity with gofakeit and go-faker. The field's Go type is the column type, through the same conversion and load checks as typed columns, @@ -27,6 +22,10 @@ The record API lands first, so the data update can use it. - `code` and `symbol` sibling fields reading `currency`, as `{code} {symbol}` → a matching pair - `{a} & {b}`, each reading `person` → one person, or two when `a` and `b` name different groups - two bare `{/sv_SE.word}` → two words +- Reference inheritance — settle whether a column that is exactly one reference to + another record's column, like `{/src.score}`, takes that column's datatype and + null. Today a null there writes `""`, and a typed column reading it is refused. + Settle before draw groups and the data update. ### Data diff --git a/value.go b/value.go new file mode 100644 index 0000000..d8aa48c --- /dev/null +++ b/value.go @@ -0,0 +1,284 @@ +package fejkdata + +import ( + "fmt" + "math" + "regexp" + "strconv" + "strings" +) + +// proven is what a proof knows of every render of a node: bounds on the number each +// reads as, and per datatype why some render's text is not one ("" when none). +type proven struct { + lo, hi float64 + nonZero float64 // every value is at least this far from zero; 0 when one can be zero + integral bool + notOperand string // why some render reads as no finite number, the way calc reads it + not [len(dataTypeNames)]string +} + +// valueProof proves what typed columns and their calc operands hold, each node once per scope. +type valueProof struct { + memo map[node]proven +} + +// checkDatatype rejects a typed column some render of which is not text of its datatype. +func (p *valueProof) checkDatatype(path string, n node) error { + t, ok := n.(*template) + if !ok || t.datatype == DataTypeString { + return nil + } + if err := p.prove(t, t.datatype); err != nil { + return fmt.Errorf("%s: %w", path, err) + } + return nil +} + +// prove reports why some render of n is not text of datatype d. +func (p *valueProof) prove(n node, d DataType) error { + if reason := p.of(n).not[d]; reason != "" { + return fmt.Errorf("datatype %s: %s", d, reason) + } + return nil +} + +func (p *valueProof) of(n node) proven { + if v, done := p.memo[n]; done { + return v + } + if p.memo == nil { + p.memo = map[node]proven{} + } + var v proven + switch n := n.(type) { + case *choice: + v = p.unite(n.items) + case *template: + v = p.template(n) + default: + v = unproven(`it reads a null, which renders "" outside its own column`) + } + p.memo[n] = v + return v +} + +func (p *valueProof) unite(nodes []node) proven { + v := p.of(nodes[0]) + for _, n := range nodes[1:] { + w := p.of(n) + v.lo, v.hi, v.nonZero = min(v.lo, w.lo), max(v.hi, w.hi), min(v.nonZero, w.nonZero) + v.integral = v.integral && w.integral + if v.notOperand == "" { + v.notOperand = w.notOperand + } + for d := range v.not { + if v.not[d] == "" { + v.not[d] = w.not[d] + } + } + } + return v +} + +// template proves a template that renders one value: fixed text, or a format that is +// one token alone. +func (p *valueProof) template(t *template) proven { + switch { + case t.repeat != 1: + return unproven(fmt.Sprintf("%q carries a repeat, which composes text rather than one value", t.format)) + case t.fixed: + return literalValue(t.lit) + case len(t.ops) != 1: + v := unproven(notOneValue(t.format, "{int()}, {float()}, {seq()} or {calc()}")) + v.notOperand = notOneValue(t.format, "{int()}, {float()}, {seq()}, {digits()} or {calc()}") + return v + } + body := t.format[1 : len(t.format)-1] + name, args, isFunc := funcCall(body) + switch _, isTransform := transforms[name]; { + case !isFunc: + var leaves []node + for _, a := range splitArms(body, t.refs) { + leaves = append(leaves, pathLeaves(t.fields[a.key], a.tail)...) + } + return p.unite(leaves) + case name == "calc": + return p.calc(t, body, args) + case builtins[name].number != nil: + return builtins[name].number(body, args) + case isTransform: + return unproven(fmt.Sprintf("{%s} rewrites text rather than printing a value; write the values it would print", body)) + } + return printing(body, DataTypeString, proven{notOperand: fmt.Sprintf("{%s} prints text, not a number", body)}) +} + +func (p *valueProof) calc(t *template, body string, args []string) proven { + expr, err := parseCalc(args[0]) + if err != nil { + panic(fmt.Sprintf("fejkdata: calc(%q) reached a proof unparsed: %v", args[0], err)) + } + v, doubt := p.expr(expr, t.fields) + if doubt == "" && !(magnitude(v) <= calcLimit) { + doubt = calcText(expr) + " is not proven within 1e300" + } + if doubt != "" { + return unproven(fmt.Sprintf("{%s}: %s", body, doubt)) + } + return printedNumber(body, v, calcDecimals(args)) +} + +// calcLimit is the largest magnitude a proof accepts as finite, far enough below +// math.MaxFloat64 that rounding in the bounds cannot hide an overflow. +const calcLimit = 1e300 + +// expr bounds a calc expression from its operands, or says why it cannot. +func (p *valueProof) expr(n calcNode, fields map[string]node) (proven, string) { + switch n := n.(type) { + case calcNum: + v := float64(n) + return bounded(v, v, v == math.Trunc(v)), "" + case calcVar: + v := p.of(fields[string(n)]) + if v.notOperand != "" { + return proven{}, fmt.Sprintf("operand %q: %s", string(n), v.notOperand) + } + return proven{lo: v.lo, hi: v.hi, nonZero: v.nonZero, integral: v.integral}, "" + case calcNeg: + v, doubt := p.expr(n.x, fields) + v.lo, v.hi = -v.hi, -v.lo + return v, doubt + case calcBin: + l, doubt := p.expr(n.l, fields) + if doubt != "" { + return l, doubt + } + r, doubt := p.expr(n.r, fields) + if doubt != "" { + return r, doubt + } + return combine(n, l, r) + } + panic(fmt.Sprintf("fejkdata: calc node %T has no bound", n)) +} + +// combine bounds one operation from the bounds of its sides. +func combine(n calcBin, l, r proven) (proven, string) { + var v proven + integral := l.integral && r.integral + switch n.op { + case '+': + v = bounded(l.lo+r.lo, l.hi+r.hi, integral) + case '-': + v = bounded(l.lo-r.hi, l.hi-r.lo, integral) + case '*': + v = bounded(min(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi), max(l.lo*r.lo, l.lo*r.hi, l.hi*r.lo, l.hi*r.hi), integral) + v.nonZero = max(v.nonZero, l.nonZero*r.nonZero) + default: + if r.nonZero == 0 { + return v, fmt.Sprintf("divides by %s, which is not proven nonzero", calcText(n.r)) + } + m := magnitude(l) / r.nonZero + v = proven{lo: -m, hi: m, nonZero: l.nonZero / magnitude(r)} + } + if !(magnitude(v) <= calcLimit) { + return v, calcText(n) + " is not proven within 1e300" + } + return v, "" +} + +// bounded is a number in [lo, hi], its distance from zero read off the bounds. +func bounded(lo, hi float64, integral bool) proven { + v := proven{lo: lo, hi: hi, integral: integral} + switch { + case lo > 0: + v.nonZero = lo + case hi < 0: + v.nonZero = -hi + } + return v +} + +func magnitude(v proven) float64 { return math.Max(math.Abs(v.lo), math.Abs(v.hi)) } + +// printedNumber is what a token printing v to dp decimals holds: an integer when whole +// and within int64, else a number. +func printedNumber(token string, v proven, dp int) proven { + if dp >= 0 { + half, _ := strconv.ParseFloat("5e-"+strconv.Itoa(dp+1), 64) + v = proven{lo: v.lo - half, hi: v.hi + half, nonZero: math.Max(0, v.nonZero-half), integral: v.integral || dp == 0} + } + if dp != 0 && !(dp < 0 && v.integral) { + return printing(token, DataTypeNumber, v) + } + v = printing(token, DataTypeInteger, v) + if !(magnitude(v) < math.MaxInt64) { + v.not[DataTypeInteger] = fmt.Sprintf("{%s} is not proven within int64", token) + } + return v +} + +// printing is v for a token whose every render is text of datatype prints, with a reason +// against each datatype that text is not. +func printing(token string, prints DataType, v proven) proven { + for d := DataTypeInteger; d <= DataTypeBoolean; d++ { + if prints != d && !(prints == DataTypeInteger && d == DataTypeNumber) { + v.not[d] = fmt.Sprintf("{%s} prints %s, not %s", token, dataTypeNouns[prints], dataTypeNouns[d]) + } + } + return v +} + +func notOneValue(format, calls string) string { + return fmt.Sprintf("%q is not one value; write one literal or one %s, or read one", format, calls) +} + +// unproven is a render no datatype and no calc can take, for why. +func unproven(why string) proven { + v := proven{notOperand: why} + for d := DataTypeInteger; d <= DataTypeBoolean; d++ { + v.not[d] = why + } + return v +} + +var ( + integerText = regexp.MustCompile(`^-?(0|[1-9][0-9]*)$`) + numberText = regexp.MustCompile(`^-?(0|[1-9][0-9]*)(\.[0-9]+)?([eE][+-]?[0-9]+)?$`) +) + +// literalValue proves fixed text: the number calc reads it as, and each datatype it is. +func literalValue(text string) proven { + var v proven + if f, err := strconv.ParseFloat(strings.TrimSpace(text), 64); err != nil || math.IsNaN(f) || math.IsInf(f, 0) { + v.notOperand = fmt.Sprintf("%q is not a number", text) + } else { + v = bounded(f, f, f == math.Trunc(f)) + } + if _, err := strconv.ParseInt(text, 10, 64); !integerText.MatchString(text) { + v.not[DataTypeInteger] = fmt.Sprintf("%q is not an integer", text) + } else if err != nil { + v.not[DataTypeInteger] = fmt.Sprintf("%q is past the int64 range", text) + } + if v.notOperand != "" || !numberText.MatchString(text) { + v.not[DataTypeNumber] = fmt.Sprintf("%q is not a number", text) + } + if text != "true" && text != "false" { + v.not[DataTypeBoolean] = fmt.Sprintf("%q is not a boolean", text) + } + return signedZero(text, v) +} + +// signedZero refuses a zero written with a sign as a typed value, naming it unsigned. +func signedZero(text string, v proven) proven { + mantissa, _, _ := strings.Cut(strings.ToLower(text), "e") + if !strings.HasPrefix(text, "-") || v.notOperand != "" || strings.Trim(mantissa, "-0.") != "" { + return v + } + for _, d := range []DataType{DataTypeInteger, DataTypeNumber} { + if v.not[d] == "" { + v.not[d] = fmt.Sprintf("%q is zero written with a sign; write %q", text, text[1:]) + } + } + return v +}