- Gooo typed path tiny v1
- Closed contract and resources
- Training observations
- Reserved bilingual TDD experiment
- Go usage
- Subsequent native compiler dogfood review
- Fresh native structural inference (2026-10-01 KST)
- Native integration is on main
- Prepared native Gooo paths (2026-10-01 KST)
- Incremental finite construction, 2026-10-01
- Closed contract and resources
Gooo typed path tiny v1
An independently trained bilingual model for ranking compiler-owned Gooo structural paths. The workflow is natural-language judgment → typed path ranking → finite tests → code assembly → recorded feedback. First proposal accuracy, finite completion, remaining candidates, failed cases and cost are reported separately.
Three training comparisons each include FP32, post-training ternary and
ternary-aware variants. No Laya weights are used. Two comparisons transfer the
hidden layer of our own public operator model; positioned-random is independently
initialized. Runtime, data generation, assembly and evaluation are Go. Python and
PyTorch are used only for offline MPS training.
Closed contract and resources
Five edits: local reference, assignment target, operand order, if branch-body
layout and declared assignment execution order. Identifiers, body contents and
alternatives are compiler-owned; the model returns a fixed label. Every offered
alternative and the combined body are type/scope checked. The structural ABI is
gooo/tiny-path-decision-model/v1, separate from the existing operation ABI.
256 features → 48 hidden units → eight labels: 12,728 parameters. FP32 weights occupy 50,912 bytes; ternary files 2,759 bytes. Five trits per byte use 1.6 bits per matrix weight plus float biases/scales. Resident tensor arrays are 50,912 bytes (FP32) or 12,896 bytes (ternary), plus eight bytes of matrix scales. The shared workspace is 1,248 bytes. Packed size is not process RAM or evidence of packed-arithmetic speedup.
uniform-transfer uses uniform hashed byte n-grams. Positioned models use
positioned_intent_ngrams_v1: four coarse intent positions and lower context
weight. Uniform features collide for some reordered assignment sentences.
Positioned features separate that pair, but did not improve global label
results here. Transferring hidden weights across this changed feature layout is
an initialization experiment.
Training observations
Public synthetic curriculum: 6,240 views (4,800 training, 480 calibration, 960 development evaluation). Templates and numeric configurations are partitioned; plain, fallback Gooo body and PROV-O vocabulary views stay grouped. Evaluation views represent 320 instructions and 160 program configurations. These 960 views informed representation development and are reused development evaluation, not an untouched final holdout.
| Comparison | FP32 global label /960 | PTQ /960 | QAT /960 |
|---|---|---|---|
| Uniform transfer | 609 | 576 | 643 |
| Positioned transfer | 482 | 399 | 566 |
| Positioned random | 594 | 599 | 605 |
All nine bundles retain global confidence threshold 1.0. Conservative Choose
abstains on every recorded development view and uses the fallback. This is not
successful learned body generation. Go/Python parity is separately verified on
96 saved predictions per comparison. The preserved first MPS failure occurred
before any optimizer step: empty repair indices promoted a batch index to
float64. Corrected training uses integral indices.
Positioned-random 60-epoch FP32/QAT loops took 1.363/1.212 seconds. Sampled MPS allocated peak was 5,717,248 bytes; Python lifetime peak RSS 478,281,728/482,066,432 bytes. These are short training observations, not Go inference requirements.
Reserved bilingual TDD experiment
After freezing checkpoints, reserved configurations 64..95 and new English/Korean templates produced 640 instructions, 320 program configurations and 1,920 views. The fixed positioned-random model ranks two eligible labels. Search evaluates at most two candidates using four authored cases; eight separate boundary inputs are checked after selection. Expectations use independent arithmetic.
| Arm | Initial typed label /1920 | Evaluated candidates | Final unseen-input cases | Median prediction | Median typed search |
|---|---|---|---|---|---|
| Deterministic order | 960 | 2,880 | 15,360/15,360 | No calls | 69.584 µs |
| FP32 ranking | 1,372 | 2,468 | 15,360/15,360 | 8.500 µs | 76.834 µs |
| PTQ ranking | 1,233 | 2,607 | 15,360/15,360 | 9.958 µs | 78.958 µs |
| QAT ranking | 1,300 | 2,540 | 15,360/15,360 | 10.041 µs | 79.333 µs |
FP32 saves 412 candidate evaluations (14.31%), while median search time increases. This is useful prioritization on a bounded workload, not wall-time speedup. Initial FP32 English/Korean choices are 671/960 and 701/960. All 1,920 global confidence abstentions per model remain recorded. TDD acceptance is a separate explicit mode. Conditional pair probabilities rank paths; they are not calibrated probabilities of functional completeness. No checkpoint is selected using this probe.
5,760 local predictions, zero external calls and zero native compiler calls. Whole audit: 1.23 s wall, 0.95 s CPU, 26,345,472-byte lifetime peak RSS. CPU/wall is approximately 77.24% of one core, not host utilization increase. Per-arm search timings exclude holdout evaluation and serialization. All per-view receipts are in the GitHub repository.
Go usage
Use gooo-neural-decision-experiments.
go run ./cmd/gooo-path-compose --plan plan.json
go run ./cmd/gooo-path-compose --plan plan.json --model positioned-random/models/fp32/model.json
go run ./cmd/gooo-path-compose --plan plan.json --model positioned-random/models/fp32/model.json --tests tests.json --max-attempts 2
Tests use gooo/typed-path-finite-tests/v1 and integer input/expected cases.
Without a model, assembly and TDD use deterministic declared order with no
inference/network calls. With TDD, one prediction per decision ranks candidates
before finite tests; there is no second model call after generation. Explicit
sampling binds plan/model/seed hashes. Deadlines, budgets, rejected types and
partial completion are recorded.
This is a Go structural assembly stage before native Gooo codegen. The native
compiler's --tiny-model still uses the operation ABI. Arbitrary body synthesis,
native in-process structural inference, OWL reasoning and production online
learning are not established. Finite synthetic cases do not prove general
natural-language intent completion. PROV-O views condition vocabulary; provenance
can support later learning without becoming an acceptance authority.
Publication copies a fixed synthetic-only allowlist and verifies digests. Credentials, host paths, environments, private source and upstream large weights are excluded. GitHub stores source, failed attempts and full raw evidence; this HF bundle stores weights, public data and compact reports.
Subsequent native compiler dogfood review
The reviewed publication retains exactly the same nine weight files. A separate
Go driver replays saved offline/FP32 TDD selections through the actual native
compiler at source 68361e64def5457f0d0e6de972570a9885cceb96, using Go 1.27.1.
It checks compiler type checking, deterministic replay and compiler source ID,
then executes the exact emitted Go bytes in isolated packages.
Twenty native generations (five families × two languages × offline/FP32) pass 240/240 independent arithmetic cases. These executions repeat five unique functions; the result is not twenty independent code synthesis tasks. Both arms select the same final functions after finite TDD. Native replay makes zero additional model predictions and zero external calls. Model ranking already happened in the preceding 5,760-prediction probe; native execution does not call the model a second time. Structural inference is still a Go stage before native codegen, not the native compiler's operation provider.
The first replay incorrectly expanded the runner's short commit into a full
revision. Its raw metadata is preserved and excluded from primary source-bound
evidence. The driver now checks the declared full revision against its clean
tracked checkout before execution and captures its running binary digest. A
fresh run at 643ca6ead45cef35d85175864aa3b16556346bca provides the primary
native review. Across both local attempts there were 40 native calls and 480
executed finite cases, with no new model calls. Only the correctly bound fresh
run supplies the 20-call/240-case primary review; the earlier correction record
is included here, and full captures are retained in GitHub.
This validates an end-to-end compiler path for the measured closed structures. General intent completeness, arbitrary code generation and runtime speedup are still unmeasured. Future work can use failure/type/provenance receipts to improve bilingual ranking and reduce setup/candidate cost, while keeping explicit finite completion and deterministic disconnection.
Fresh native structural inference (2026-10-01 KST)
The subsequent direct integration uses native compiler source
d21ce275ec0832e38ad4962ad03f4546c1b46e9f and public Go SDK
v0.2.0-experimental at e67cb5934d892fbc81be18038be2d624f1a6b5e0.
The SDK's CI passed before tagging. Native implementation and fixtures are
public in PR 1112.
This supersedes the preceding review's limitation to replaying saved selections:
native body-codegen --path-plan [--path-model] now loads a structural bundle
and performs one in-process prediction per decision before finite tests.
The compiler first binds the declared fallback to the authoritative Gooo body with its existing canonical equivalence witness. It validates all individually offered options and combined scope/types, ranks finite alternatives, retains partial results, replaces only an in-memory computes literal, and emits marked Go with stable semantic identity, native type checking and deterministic replay. It calls no external provider, even if Laya environment settings are configured. Disconnection retains deterministic finite search from the declared fallback.
The fresh probe at runner source b4d4ebe08335ab96427e2add3b012126d4687a3f
uses one compound body with local reference, branch assignment/layout, and
declaration order: 8 combinations, 2 scope-invalid. Two languages × four arms
(offline/FP32/PTQ/QAT) × attempt budgets 4/8 × full/partial finite contracts give
32 native calls and 72 fresh native model predictions. These are views of
one compound experiment, not 32 distinct synthesis ideas. No weights are updated;
the nine published bundles retain their prior bytes.
All 32 exact native emitted Go functions independently execute 10 integer inputs absent from the search cases, including int64 overflow edges. The 32 Go tests pass selected arena/native semantic parity, while the separate arithmetic intent observations pass 226/320 cases. All budget-8 arms pass 160/160 of those arithmetic observations. At budget 4, offline misses all 40; the models pass all 60 English cases but only 6/60 Korean cases. Thus a small search budget can still miss the intended assembly, even with identical initial selected labels, because ranking weights change subsequent exploration. No claim of unseen natural-language accuracy or full-domain proof is made.
The deliberately wrong partial contract requests 999 for input 3. With budget 8 the correct arithmetic function still yields a recorded 2/3 = 66.67% on that declared suite. Native lowering remains complete. Neither the incorrect expectation nor the partial fraction is silently repaired or presented as 100%.
| Arm | Median prediction | Median native stages | Median lifetime native RSS |
|---|---|---|---|
| Offline | no prediction | 1.571 ms | 18,350,080 bytes |
| FP32 | 8.375 us | 1.681 ms | 18,554,880 bytes |
| PTQ ternary | 9.646 us | 1.623 ms | 18,415,616 bytes |
| QAT ternary | 9.833 us | 1.658 ms | 18,415,616 bytes |
Process CPU-time/wall ratios have arm medians 81.49–82.89% of one core. These are per-child measurements, not host CPU utilization increases. The native stage durations exclude process startup; per-cell process wall time is retained separately. This single fixed-order local probe has no matched repetitions and does not establish a speedup. Inference uses no GPU. The packed ternary weights remain 2,759 bytes; process RSS includes Go runtime, parsing, typing, arenas, search and decoded tensors and is not the packed-weight size.
Initial fixture preparation rejected unsupported activity-ID syntax and then
caught a manually miscomputed arithmetic expectation before this source-bound
probe. The corrected explicit contract is 5*input+2, minus 3 for negatives,
plus 3 otherwise. Native full local tests also report failures in unrelated
macOS namespace replacement, symlink diagnostics and source-splitter assertions;
no full local-suite pass is claimed. The new focused race tests passed. Actual
Linux CI results remain independently inspectable on the public PR.
The earlier ad-hoc mixed-language FP32 check made one additional native call and three predictions; its 2.744 ms stage observation is excluded from the 32-cell primary probe. Publication contains only synthetic fixtures, public model metadata/weights and selected evidence. No Laya weights, environment, credentials, private repositories or user paths are included.
Native integration is on main
PR 1113 merged
the direct structural model integration to native compiler main at
9b8b900e32384b5397fb3d08dc545e22cc7530da on 2026-09-30 22:51:48 UTC.
The protected promotion retained the exact dev tree with the previous main
parent. The six canonical checks and actual PR-bound promotion authorization
passed in CI 36785977424.
The downloaded proof/append-only receipt independently passed the Go verifier
before the ordinary squash merge; no protection or required-review change was
made. Required approving reviews remain zero and Guardian is absent.
A clean main build with Go 1.27.1 then loaded the published FP32 structural model directly, made three fresh local predictions, passed the fixture's three explicit cases, type checking and replay, and recorded the actual main source SHA. It made zero external calls. This main smoke is separate from the primary 32-cell probe and is not an additional independent arithmetic holdout study. The main and dev trees were checked equal after merging. This records deployed source availability; it does not claim a newly packaged native release binary. Post-main push CI was queued when this note was written; the successful authorization above comes from the actual pre-merge PR run.
Use native body-codegen --path-plan [--path-model], documented in the
Gooo compiler body-codegen guide.
For example, use positioned-random/models/fp32/model.json from this public
repository with the compound plan/fixture supplied by the compiler. Omitting
the model retains deterministic finite search. These are closed bilingual
structural choice models; they do not freely generate arbitrary program text.
All nine published weight bundles remain unchanged.
Prepared native Gooo paths (2026-10-01 KST)
The Go-only SDK v0.2.1-experimental owns one validated immutable plan snapshot and its compiled fallback. Native PR 1114 uses that snapshot for source binding and search, preserving every combined candidate's existing scope/type checks, final native verification and zero source writes. It introduces no global cache, retained test outcomes or additional model call.
The fixed new compound intent combines six decisions: comparison operand order,
Boolean local and assignment target, Integer assignment target, branch bodies
and Boolean-update order. English/Korean, four arms and budgets 8/64 form 16
cells, paired over five balanced-order repetitions. The clean baseline is native
main 9b8b900e32384b5397fb3d08dc545e22cc7530da; the clean prepared candidate is
e155148f5ce114ee4c55e720a0e90cc34cc458ce. These are views of one idea, not
160 independent tasks. No checkpoint is selected or trained on this study.
There are 160 actual native generations, 720 local predictions, zero external calls and zero optimizer steps. All 80 old/new pairs retain identical exact Go source and search semantics. The emitted Go independently executes in 160 isolated packages and matches the selected arena on all 1,440 observations. Independent intent arithmetic passes 645/720 per version, including 360/360 at budget 64 and 285/360 at budget 8. Fifty native calls retain a partial result.
Authored structural target-label agreement is separately 390/480 per version. Even at full budget, arithmetic succeeds on 360/360 inputs while target-label agreement is 220/240. Equivalent finite behavior does not establish every requested internal variable/path choice. The target labels and nine independent arithmetic inputs are inspected after selection, never passed into the model or finite search API. The seven stated cases alone guide bounded search.
| Local five-repetition comparison, median | Baseline | Prepared |
|---|---|---|
| Native processing stages, excludes startup | 3.202 ms | 1.562 ms |
| Whole native child wall time | 9.839 ms | 8.448 ms |
| Lifetime child peak RSS | 20,807,680 bytes | 19,308,544 bytes |
| Process CPU/wall, relative to one core | 89.96% | 86.04% |
The median paired stage saving is 1.590 ms. Balanced ordering limits systematic
first-run bias, but this remains one local workload on one host. RSS is an OS
child-lifetime high-water value, not packed weights or prepared-object resident
size; CPU/wall is not the host's utilization increase. New plan_prepare_ms is
separate, while older binding/search stages included preparation. Compare total
processing before attributing per-stage differences.
The compact comparison/cohort/independent-Go evidence accompanies this model
repository. Full 643-file raw captures, the initial missing-entity fixture
failure, preliminary smokes and the runner are in the
public raw capture.
Runner/workload source is frozen at research revision
61f12318262e69520ab4af6663afb95004c494dc.
The first authoring failure was refused before inference and corrected by
declaring the fixture's Integer entity. All nine model weight bundles stay
unchanged; the update improves compiler integration and measurement, not model
first-choice accuracy. Earlier model-card sections retain historical observations.
Incremental finite construction, 2026-10-01
This edition preserves the previous nine tiny weight bundles and all historical evidence. It adds a rank-once session experiment, not new training or an upstream Laya invocation. Korean and English intentions select bounded Gooo structures; type-safe candidate bodies can remain functionally partial.
The public native feature source is
f6323846f18602c2cf3cef5a40692340cae9fe3d, using public Go SDK
v0.2.2-experimental from a124721ab8538ebfbec9a8fddee6e7b88331a31c
(SDK CI 36797429511 passed). The measurement runner source is
4aa721d24997e76ca7aa1c95124de415bddc44c8. This records feature execution;
main merge receipts are published separately on GitHub.
Across 32 paired policies and 288 native calls, continued search preserves the final Go body and finite search in all pairs, with 166 matched intermediate prefixes. Repeated restart budgets 8,16,...64 cost 1,152 model judgments and 6,340 candidate attempts; one continued request costs 144 judgments and 1,268 attempts. The latter ranks once, remembers tried masks, and advances eight new candidates at a time. No test/CI-feedback reranking occurs.
Sixteen deliberately inconsistent contracts remain partial at 6/7 cases; sixteen normal contracts reach 7/7. Final finite observations are 208/224. These cases drive selection and are not new heldout language accuracy. The workload is one compound intent; language, model, contract and repeat views are not 32 separate ideas or another 100-experiment claim.
Median whole-policy wall time is 69.59 ms for eight independent requests and 10.22 ms for one continued request. Child CPU time is 62.58 ms versus 9.37 ms. The median maximum child RSS is 20,742,144 versus 21,004,288 bytes: the retained receipt costs memory. This is a policy comparison on one host, not equal-call kernel throughput or a host CPU utilization increase. The eight-byte bitset for 64 masks excludes frontier, body, model, runtime and receipt memory.
--path-step-attempts 8 enables this native experiment. Without a model,
declared-fallback ordering is deterministic. The result contains initial
ranking and batch progress; it emits final Go once. Progress digests link
observations without authorizing ontology mutations. Models are optional.
Source binding, stable identity, type checks, replay and native/arena parity
remain required. Generic output writers require a draining, bounded process
harness; arbitrary I/O deadlock freedom is not claimed.
The new publication allowlist contains only three additional synthetic files:
native-incremental/report.json, preexecution.json and summary.json.
The Go verifier checks exact frozen digests, 1,268 arena candidate observations,
batch hash links, cumulative denominators and raw native captures at assembly.
It makes zero additional model predictions. Old revisions stay reproducible.