- A small local judge for Gooo program construction
- Keep a compiled Gooo graph for current inputs
- Record field assembly — installed controls
- Earlier candidate and ternary observations
- Language and earlier observations
- 한국어 요약
- Connecting model files to Gooo — 2026-10-04
- What we want this to contribute
- Choose an entry point
- New native generation and execution observation
- New full-input research exports — 2026-10-03
- How the model is used
- Original V3 architecture, training and compatible files
- Decision quality
- Input sensitivity diagnosis
- Actual Gooo generation and execution
- Storage, speed and CPU observations
- Reproducibility and development
- Research acknowledgments
- Keep a compiled Gooo graph for current inputs
A small local judge for Gooo program construction
Keep a compiled Gooo graph for current inputs
Compiler main 7acccf564b3def31ddd368050c2288f05bd7076f is merged after all12
authoritative CI jobs and both local executables are installed. Ordered
current-input suites retain one compiled graph within an owned request. The optional model guides
construction once; later runtime suites make zero model calls. Keeping an
assembled tool on a workbench lets the next piece of material arrive immediately.
Current-input observations retain the original clean candidate and a separate fresh installed cohort. Installed controls keep the same program identities and actual finite results in all20workloads. Three paired saved FP32-origin workloads cost a median2,327.40ms across six stateless commands versus497.14ms in one retained command. CPU time is1.45s/0.36s and peak RSS82.88MiB/82.14MiB. Startup count differs. Each cohort retains50commands, 120suite frames and240native executions; original timings remain separate. Three changed-input suites repeated twice retain4/4,2/2,1/2,4/4,2/2,1/2 named expectations; the third label is deliberately different. All record fields match. The Go reader independently checks current inputs and recomputes finite counts.
Fresh installed FP32 prediction median is18.792µs across three constructions. One
PTQ and one QAT pilot use18,752 resident tensor bytes each and retain the same
program/current results. Their matrix encoding is five trits per byte, about1.6
stored matrix bits; each file is3,854bytes including FP32 biases. Decoded tensor
residency is separate. The earlier2-bit wording is corrected from actual files.
Weights and training are unchanged. Compiler PR1228
contains the merged release; research PR13 passed all23 CI jobs and merged.
Independent Linux readback
rechecks eight native captures. Followup PR14
also passed all23 CI jobs and merged into research main d799e9cd. Its Go text
reader shows current percentages, mismatches and separate build/native-run times.
Installed-source Linux readback
includes all eight downloaded native responses and independent recomputed reports.
한국어: 작은 모델이 한 번 조립 방향을 제안하고, 만든 프로그램은 다음 입력을 받아 바로 실행합니다. 같은 공구를 작업대에 두고 재료만 바꾸는 방식입니다. 틀린 기대값도 실제 값과 함께 남겨 다음 개선에 사용할 수 있습니다. 한 작성 예제의 반복 비용을 측정했으며 전체 호스트 CPU 사용률은 아직 측정하지 않았습니다.
Record field assembly — installed controls
The field assembly feature is merged into compiler main
f6da667e11951b6939fc5e30442e01a09ca06e85 after complete PR1226 CI and source-bound
result verification. Both local executables are installed. The
separate installed cohort
retains 24 fresh controls and six saved replays, matching the candidate's selected
source/Go hashes, rankings, masks and finite results in all 30 observations.
Fresh full-budget prediction median is31.625µs; graph generation is9.890ms/9.337ms
and whole command is323.462ms/324.194ms for model/deterministic controls. Peak
process RSS is82.94MiB/82.42MiB; CPU time is0.24s in each mode. These resources
include native Go builds and children; host-wide CPU changes remain unmeasured.
The Go completeness reader now shows selection case/field ratios and each remaining field value alongside actual compiled runtime scores. Saved composition replay made zero new predictions and built the native graph again in that earlier measurement. The subsequent candidate above retains the compiled graph. This update publishes compiler integration and observations with unchanged weights.
Earlier candidate and ternary observations
Gooo can now declare alternatives for the actual expressions inside a record constructor, together with typed JSON cases and a small attempt budget. The field assembly observation retains 24 fresh controls and six saved full-budget replays of one authored two-node graph. With 1/2/4/8 attempts, both modes complete 6/9/12/15 of 15 expected selection fields. Whole selection cases reach 5/5 and compiled execution reaches 14/14 named outputs and 21/21 record-output fields at full budget.
These are clean candidate measurements at compiler 48f70275, before main
installation. Compiler PR1225
contains the feature and subsequent standard-library-only syntax and formatted
JSON replay fixes. The existing frozen own FP32 weights are unchanged. A named
field-ordinal context transfers the integer-choice model to field alternatives;
case inputs and expected outputs are excluded. The model and deterministic
controls reached equal finite counts, with no attempt reduction on this source
shape. Full-budget prediction median was 29µs and command medians were about
349ms/342ms. Field-specific training remains subsequent work.
The separate ternary pilot
executes one PTQ and one QAT request on clean compiler cb316851. Both reach
15/15 selection fields and 14/14 compiled outputs with 18,752 resident tensor
bytes each, versus 74,624 bytes in the FP32 cohort. The actual matrices pack five
trits per byte, approximately1.6stored matrix bits. Each complete weight file is
3,854bytes including FP32 biases. The earlier2-bit description is corrected;
1.58bits describes ternary information content. The two pilot timings and decoded
tensor counts leave whole-process memory and
broad speed improvements unmeasured.
한국어: 기록의 제목·상태·사유 안에 후보식을 심고, 작은 예산에서 얼마나 채워지는지 측정했습니다. 시도 1·2·4·8회에 따라 기대 필드 완성률은 40·60·80·100%였습니다. 기존 자체 모델을 실제로 호출했지만 이 예제의 완성률과 시도 수는 결정론 방식과 같았습니다. 후보 관측과 설치 상태를 구분해 기록하며, 다음 학습은 다양한 필드 문맥에서 같은 예산으로 더 많이 채우는 것을 목표로 합니다. 새 가중치 학습은 이번 관측에 포함되지 않았습니다.
Language and earlier observations
Updated 2026-10-04. This model ranks eight permitted Gooo body paths built from three binary decisions. The project explores a language and a small model working together: Gooo provides the construction plan, the model suggests a route through it, and the compiler checks the assembled body and generates Go. A Go experiment runner then builds and executes the resulting program.
Like a workshop, the plan describes what the parts mean and how they may fit. The model helps choose the next assembly; observed test failures guide another attempt. A receipt connects the original intention to the choices and result.
The compiler now accepts an activity's assembling block for permitted source
choices, Korean/English intent, finite examples and search budget. The regular
commands expand it into typed construction and immediate compiled execution.
The latest two-task observation
retains weaker initial model ranks, bounded test recovery and measured costs.
Selected bodies now return as reusable Gooo checkpoints with the original
planning baseline and picked paths. The source reuse observation
chains three generations per activity/mode: 12 constructions, 48/48 selection
expectations and 108/108 independently executed expectations. Six real predictions
used these unchanged weights; twelve realization replays made zero predictions.
These are two repeated tasks, with no new training or measured accuracy gain.
The feature merged into compiler main f62eb5f0 (PR1218, CI37190298986) and both
executables were installed. The installed-source recheck
repeats those twelve controls and records eight parallel checkpoint requests,
runtime72/72, four actual model predictions and normal EOF completion.
Native composition now connects actual generated activity bodies through typed
Gooo bind edges. Compiler main 2d1ce2d0 is merged and installed after its own
complete CI. The fresh installed controls
retain588/588 expectations on one repeated graph, two concurrent requests98/98,
three extra input combinations42/42 and a deliberately changed expectation48/49.
The exact public FP32 model location
includes a two-file download command. Weights are unchanged; first candidates,
additional search and slower fresh timing observations remain visible.
The language now supports source-ordered scalar input joins. A body
such as Add(Integer, Integer) receives two separate inputs, and a mixed-type
body can combine bound producer values with values supplied by the caller.
The input-join observation
records 588/588 named outputs, 1,008 input-slot observations and 504 bound
deliveries across twelve repeated controls of one graph. Three fresh model
requests made three actual predictions with unchanged FP32 weights; saved
replays made zero. The model checked two candidates versus eight deterministic
candidates, while whole-generation medians were 9.711ms and 6.396ms respectively.
Development PR1221 and
main PR1222 passed their
own complete CI and independently verified proofs. Compiler main 1bef47fc
and both executables are installed. The separate installed-source controls
retain588/588 with matching source/Go/driver hashes, three actual predictions,
concurrency98/98, additional scenarios42/42 and the changed expectation48/49.
Installed fresh generation medians are9.449ms with the model and8.480ms
deterministically. Original candidate measurements retain exact source 5af78bcb.
Gooo bodies now carry source-declared record values such as a Candidate
with a title/state and a Review with a summary/reason. The
record observation retains
432/432 named expectations across twelve repeated controls of one six-activity
graph and 1,152 actual input/output field observations. The own model ranks the
scalar Score assembly; record bodies follow Gooo source. Its first candidate
met0/6, the next6/6. Median prediction was26.9µs, while graph generation took
12.80ms with the model versus11.77ms deterministically. Saved replay makes zero
new predictions. Separate concurrent requests met72/72 and eight further
runtime-only cases through saved compositions met96/96. One intentionally wrong
field expectation gives35/36 and retains the actual difference. Source candidate
fa3c8927 is retained separately from installation. Development
PR1223 and main
PR1224 passed their own
complete CI and independent proof verification. Compiler main b629a664 and both
executables are installed. Fresh installed observations
retain432/432 outputs and1,152 field observations with the same program hashes.
Actual prediction median was29.7µs; fresh generation was13.17ms versus12.01ms
deterministically. A Go result reader shows35/36 and Propose.state when the
state expectation differs. An actual replay omitting record expectations keeps
36 output fields unobserved with a null percentage. A second graph accepts
external records and replaces local values, meeting8/8 in fresh/saved controls.
Weights remain the frozen published FP32 files; this update publishes compiler
integration and finite observations.
| Component | Public home |
|---|---|
| Introductory Wiki in Korean | Gooo: language, models, metrics and research |
| Gooo language, compiler and execution | meta-ontology-go |
| Go inference and typed search | gooo-decision-runtime |
| Training, raw evidence and comparisons | gooo-neural-decision-experiments |
| Project direction in Korean | 언어·모델·실험 안내 |
한국어 요약
기록의 필드까지 실제 본문에서 전달하기
Candidate에 제목·상태를 담고 다음 본문이 Review를 만드는 경로를
구현했습니다. 서류를 다음 작업대에 전달하듯, 실제 값과 필드의 정체성을
연결마다 남깁니다. 한 그래프의 열두 반복 제어에서 출력 432/432와 필드 관측
1,152개를 확인했고, 상태 기대값 하나를 바꾸면 35/36으로 기록됩니다.
자체 모델은 정수 조립의 후보 순위를 제안합니다. 기록 구성과 읽기는 Gooo
선언을 따르며, 기존 가중치를 사용했습니다. 실제 추론 중앙값은 26.9μs지만
생성 전체는 모델 12.80ms·결정론 11.77ms로 모델 쪽이 조금 길었습니다.
사용 흐름·완전성의 분모·원본을
함께 공개합니다. 후보 관측 소스 fa3c8927을 유지하고, 개발·배포 검사를 마친
main b629a664로 두 실행기를 설치했습니다. 설치본의 별도 반복 제어도
432/432와 동일한 생성 지문을 유지했습니다. 실제 추론 중앙값은29.7μs이며,
새 기록 입력 예제와 필드별 완전성 요약을 추가했습니다. 기대값을 생략한
실행에서는 해당 출력 필드36개를 미관측으로 표시합니다.
같은 타입의 입력도 각 위치에 연결하기
새 입력 연결은 작업대에 번호가 붙은 소켓을 두는 방식입니다.
Add(Integer, Integer)의 input0과 input1은 각각 다른 생산자의 실제 값을
받습니다. Label(Integer, Boolean, Text)는 정수만 연결하고 나머지 입력은
호출자가 제공할 수 있습니다. 선언 순서가 의미 모델·IR·Go 인자 순서까지
이어져 뺄셈과 같은 순서에 민감한 본문을 구분합니다.
입력 연결 관측과 따라 할 명령을 공개했습니다. 같은 그래프를 반복한 열두 제어에서 출력 기대값 588/588, 입력 관측 1,008개·연결 전달 504개·실제 실행 24회를 확인했습니다. 자체 모델은 Left 본문을 조립할 때 세 번 호출했고 저장 기록은 추가 추론 없이 재실행했습니다. 모델의 첫 후보는 0/6, 다음 후보는 6/6 예시를 충족했습니다. 후보 확인은 모델 2개·결정론 8개였지만 전체 생성 시간 중앙값은 각각 9.711ms·6.396ms로 결정론 경로가 더 짧았습니다. 의도적으로 틀린 기대값은 48/49로 기록했습니다. 관측 수는 같은 그래프의 반복 횟수와 함께 읽습니다.
측정 소스는 5af78bcb이며 배포 검사에서 준비 함수의 분할을 더 작게 할 필요가
확인되어 6b6c86f2로 정리했습니다. 개발 #1221과
main #1222을 각 전체 CI와
독립 증거 확인 후 병합하고 main 1bef47fc로 CLI·실행기를 함께 설치했습니다.
설치본의 별도 관측에서도
같은588/588·입력1,008개·연결504개와 동일한 생성 지문을 확인했습니다.
모델은 세 번 실제 호출했고 저장 재실행은 추가 추론0번입니다. 동시 요청98/98,
추가 입력42/42, 틀린 기대값48/49도 기록했습니다. 설치본의 전체 생성 중앙값은
모델9.449ms·결정론8.480ms였습니다. 모델 가중치는 그대로 두었고 새 학습은
수행하지 않았습니다. 입력 순서와 실제 전달은 Gooo가 담당하고 작은 모델은
허용된 본문 선택의 순서를 제안합니다.
앞선 관측: 여러 활동을 연결한 프로그램으로 실행하기
컴파일러 개발 후보 c00bd714의 body-compose는 Gooo bind 연결을 따라 실제
본문을 함께 실행합니다. 정수 조립 두 개, 불리언 판단과 부정, 값의 분기 전달,
별도 텍스트 처리까지 일곱 활동을 한 번의 Go 빌드로 연결했습니다. 모델은
요청 안에서 한 번 로드하고 두 활동에 두 번 판단했습니다. 저장한 전체 조립
기록의 재실행은 추가 추론 0번입니다.
실제 사용한 파일은 공개 three-choice 모델의 고정 FP32 리비전에 있습니다.
hf download asketeddy/gooo-three-choice-feedback-tiny-v1 models/set-feedback/fp32/model.json models/set-feedback/fp32/weights.bin --revision 24238ca67048b36bc305731799985271bb752dd1 --local-dir gooo-models로 두 파일을 받고,
body-compose --model gooo-models/models/set-feedback/fp32/model.json에 연결합니다.
메타데이터 1,215바이트·가중치 74,624바이트이며 이번 관측의 실제 지문과 일치합니다.
새 생성 여섯 건과 별도 저장 기록 재실행 여섯 건에서 유한 실행 기대값 588/588, 중간 값 588개, 연결 전달 420개와 실제 실행 24회를 기록했습니다. 첫 모델 후보는 0/6·1/3으로 실패했고 유한 테스트 탐색으로 주어진 예시를 충족했습니다. 가중치는 유지했습니다. 조립 중앙값은 모델 13.527ms·결정론 13.011ms이며, 빌드와 두 번 실행을 포함한 명령 중앙값은 조건별 337.886–469.090ms입니다. 메모리와 CPU 기록에는 Go 빌드 자식도 포함됩니다. 세 번씩의 고정 순서 관측이며 속도 개선이나 전체 컴퓨터 CPU 증가량을 판단하지 않습니다. 원본·재현 도구·측정 범위, 컴파일러 PR1219, 연구 PR6. 이 기록은 main 배포 전의 위 개발 후보에서 수집했습니다.
같은 기능을 main 2d1ce2d07eaa29f00a412cbaa22c96008f0b7250으로 병합하고
두 실행기를 설치했습니다. 설치 후 기록은
같은 열두 제어의 588/588과 동일한 생성 지문, 동시 두 요청 98/98,
추가 세 입력의 두 경로 42/42, 틀린 기대값 사례의 48/49를 각각 남깁니다.
설치 후 조립 중앙값은 모델 36.933ms·결정론 32.216ms, 전체 명령은
조건별 721.523–778.166ms였습니다. 측정 시점과 동시 작업이 다르며 속도 개선을
결론내리지 않았습니다. 실행 후 CPU 스냅샷은 26.47%·17.70%이고, 모델에 따른
CPU 증가량을 구분할 실행 전·중 대조군은 없습니다.
선택한 본문을 다음 개발에 이어 쓰기
기준 본문과 선택 경로를 Gooo에 함께 기록하도록 발전시켰습니다. 모델이 고른
지역값을 원본에 넣었을 때 다음 생성의 기준이 어긋나던 문제를 다룹니다.
body-realize는 기록을 다시 구성해 Gooo 파일로 저장하며 추가 모델 호출은
0번입니다. 실제 자체 모델을 사용한 두 과제의 열두 번 생성·제어 관측에서
별도 실행 108/108을 확인했습니다. 생성 중앙값은 모델 경로에서 두 과제별
15.696ms·13.214ms, 저장은 15.738ms·13.984ms였습니다. 명령 시작과 자료를 읽는
비용이 포함되며 과제마다 세 번의 관측입니다. 원본과 범위를
함께 공개합니다. 가중치 학습·정확도 개선과 구분해 언어의 재사용성을 확인한 작업입니다.
Gooo 선언 안에 조립 의도를 함께 두기
computes 뒤에 assembling 블록을 두면 의도, 선택할 소스 위치, 허용된 대안,
유한 입출력 예시와 시도 예산을 Gooo에 함께 적을 수 있습니다. 선언은 의미 IR에
남고, 선택된 본문은 Go 생성과 실제 실행으로 이어집니다. 작은 조립 지시서를
본문 옆에 붙여 두는 방식입니다. 구문과 실행 예제.
gooo body-codegen --json --activity Qualified --path-model /path/to/model.json \
examples/body-codegen/source-assembly.gooo.fixture
모델 경로를 생략하면 같은 선언으로 결정론적 탐색을 합니다. 실행할 때에는 별도 입출력 예시를 제공해 생성 단계의 예시와 실제 Go 동작을 나누어 관찰합니다. 기존 모델의 입력 계약을 이용하며, 이 구문 때문에 추가 학습을 요구하지 않습니다.
이번 관측은 기존 자체 set-feedback FP32 joint 모델을 사용했습니다.
두 과제·두 작성 방식·모델/결정론적 경로에서 32회 조립을 수행했고, 선택 예시
128/128개와 별도 실행 예시 288/288개를 충족했습니다. 모델의 첫 후보는 각각
0/5와 1/3으로, 두 과제에서는 결정론적 순서보다 불리했습니다. 유한 테스트 탐색을
통해 완성됐고 이 실패와 추가 후보 수를 다음 학습의 재료로 남겼습니다.
선언 기반 모델 요청의 반복 응답 중앙값은 27.0–28.6ms, joint 판단 중앙값은 약 20µs였습니다. 원본 계약을 재구성하고 확인하는 비용 때문에 별도 계획 방식보다 응답 시간이 소폭 늘었습니다. 학습과 가중치는 유지했습니다. 원본 예제·측정 범위·자체 모델 출처.
별도 후속 실험은 첫 후보의 실패와 실제 개발 CI 상태를 다음 판단의 문맥으로 전달했습니다. 올바른 기대값으로 두 요청이 각각 선택 5/5와 실제 실행 11/11을 충족했습니다. 앞서 다른 예제의 기대값을 섞어 얻은 1/11도 보존했습니다. 위 32회 집계와 구분하고, CI 상태의 효과는 별도로 비교하지 않았습니다. 선언 기반 조립 관측과 후속 문맥 부록.
해당 언어 변경은 개발 #1215·main #1216의 자체 CI와 독립 증거 확인 후 병합했습니다.
main 8a08adbb에서 CLI·병렬 실행기를 함께 설치했고, 별도 16개 요청에서 제공한
실행 기대값 144/144와 정상 EOF 종료를 확인했습니다. 새 설치본 모델 재사용 응답은
두 예제에서 34.745·33.816ms의 단일 관측입니다.
설치 버전·후보 검사·시간 기록.
최근 재현 검사
과거 계산 방식의 macOS·Linux 기록을 각각 고정 원본과 대조하는 Go 검증기를 추가했습니다. 새 실행은 환경마다 18,432개 개발 기록과 24개 compact 형식 모델 파일을 모두 재현했습니다. 원래 두 환경에서 달랐던 요약 101개 항목과 후보 순위 272개는 양쪽 값과 입력 해시를 함께 보존합니다. 검증기는 모델을 호출하지 않으며 새 감사 실행의 Go 추론 호출은 환경마다 25,824회입니다. 원래 가중치와 기대값은 유지합니다.
Linux 검사 37126899513의
19개 작업이 통과한 뒤 연구 PR #1을 병합했습니다. 새 산술의 플랫폼 간 검사는
계속 별도로 실행합니다. 새 원본과 재현 방법을
공개했습니다. 당시 설치된 컴파일러 main ed2c2cac는 SDK v0.2.20-experimental을
사용했습니다. 최근 파일 연결과 설치 상태는 아래 갱신에, 이전 연결 실험의
버전·결과는 각 날짜의 기록에 남아 있습니다.
Gooo 선언을 설계도, 작은 모델을 조립 순서를 고르는 장치로 생각하면 됩니다. 현재 모델은 세 가지 이진 판단으로 이루어진 여덟 경로에 순위를 매깁니다. 컴파일러는 조건식·할당·분기 등을 조립하고, 실패한 입출력 예시를 다음 시도의 문맥으로 전달합니다. 생성된 Go의 빌드·실행 결과와 남은 의문도 기록합니다.
개발은 Laya의 구조화된 선택 실험에서 출발해, Gooo 자료로 새로 학습한 2,072개 파라미터 모델과 Go 실행기로 이어졌습니다. 현재 공개물에는 FP32와 삼진 가중치, 학습 기록, 실제 코드 생성·실행 관측이 함께 들어 있습니다. 목표는 적은 메모리로 의도와 생성 경로를 연결하고, 완전성을 여러 항목으로 관찰하며 다음 작업을 이어갈 수 있게 하는 것입니다.
문장 앞의 도입 표현에 민감한 문제를 확인한 뒤, 전체 입력을 유지하는 네 가지 새 모델을 학습했습니다. 기존 개발 입력에서 FP32의 첫 경로 충족 수는 대조군 113/512, 문장 조각의 빈도를 쓰는 조건에서 368/512였습니다. 표현을 다양하게 학습한 조건과 삼진 양자화에서는 후퇴한 결과도 있었습니다. 새 결과는 아래 연구 부록에 있으며, 기존 모델의 코드 생성·실행 관측은 해당 실험별로 남아 있습니다.
새 모델을 arm64와 Linux에서 대조한 결과, 첫 선택은 같았고 후순위까지 포함한 272개 후보 순위에 차이가 있었습니다. 이후 반올림 시점을 명시한 규칙으로 147,456회 호출을 대조했고, 18,432개 입력 쌍의 계산값과 전체 순위가 모두 일치했습니다. 새 계약은 SDK v0.2.15로 공개했고, 로컬과 Linux에서 각각 18,432개 입력·36,864회 판단으로 연구 결과를 재현했습니다. 이어서 SDK v0.2.15를 연결한 컴파일러로 400회 생성과 800회 실행을 완료했고, 9,600개 기대값이 모두 일치했습니다. 실제 모델 판단은 816회였으며, 모델을 끈 경로도 결정론적으로 완료했습니다. 아래에 새 관측과 기존 V3 모델 사용 예시를 각각 연결했습니다.
Connecting model files to Gooo — 2026-10-04
The released Go SDK v0.2.21
uses bounded model reads and checks the opened regular-file identity and extent.
On Unix, nonblocking/no-follow open prevents the observed replacement-FIFO
writer wait. The SDK's existing non-symlink model rule is retained. Use
hf download --local-dir to obtain regular metadata and adjacent weights.
The exact compact FP32 pair at Hub revision
7c2781cbbcf81edba89a50d0423276c042b082de was downloaded and checked:
metadata SHA256 6478334e1e181d0854865d1da54b9c2203eaa0e488f2b05617c92e27f08b755f,
weights SHA256 43c503b224a78f23725101e58cdef1ec4647831329eea4de80f0a259a0253d31.
Compiler dev PR1206
passed its exact-head CI and independent source proof. Candidate 2420ad19
completed 128 model-path swaps with zero writer releases/timeouts; 34 regular
completions kept 4,352/4,352 expectations. The standalone shared-model worker
kept 512/512 in concurrent Korean/English requests and joined both processes.
Original failures, counts and process CPU scope.
Main promotion PR1207
passed its own canonical CI37167954453, independently verified source proof and
live promotion authorization, then normally merged as 4f6c7566. Both CLI and
worker are freshly installed from that clean main with Go1.27.1 / SDK v0.2.21.
Fresh installed swaps completed 128 requests with zero writer releases/timeouts;
35 regular completions kept 4,480/4,480. Separate ordinary runs kept 1,024/1,024
and concurrent shared-model worker requests kept 512/512 with joined EOF.
The older d1bfd273 / SDK v0.2.20 wait remains recorded as the baseline.
The research comparator now binds each recorded source inventory to its own immutable Git revision. Complete local/Linux replays each made 73,728 actual predictions over 18,432 paired inputs. Explicit arithmetic kept all values, rankings and finite curves; fourteen source changes stay visible. Legacy arithmetic keeps 272 ranking and 38 finite-outcome differences between platforms. Research CI37168510833 completed SUCCESS, and its arithmetic archive was independently downloaded, hash-checked and fully consumed. Source deltas and original differences.
A direct file-based example
With Go1.27.1, a compatible compiler and the compiler repository as working directory, the existing compound source, plan and expectations can run directly:
hf download asketeddy/gooo-shared-judgment-tiny-v1 \
--revision 7c2781cbbcf81edba89a50d0423276c042b082de \
--include 'research/compact-runtime-20261003/models/fp32/*' \
--local-dir ./gooo-models
gooo body-path-run \
--source examples/body-codegen/typed-path-compound.gooo.fixture \
--activity Combined \
--path-plan examples/body-codegen/typed-path-compound-plan.json \
--cases examples/body-codegen/typed-path-runtime-cases.json \
--model gooo-models/research/compact-runtime-20261003/models/fp32/model.json \
--repeat 2 --timing --out compound-model-results
gooo body-path-run --verify-timing --out compound-model-results
Choose a fresh output path. Omitting --model uses deterministic construction.
Read assembled Go in run-N-generated.go, current case outputs in
run-N-runtime.json, and request status plus the finite numerator/denominator in
summary.json. Setup failures and unobserved cases have their own recorded state.
The clean candidate made two actual shared-model predictions and four native
runs, keeping 6/6 expectations in this existing example. First/repeated responses
were 884.413250/70.435292ms, covering construction and execution while excluding
initial model preparation and saving/output. Weights and training remain unchanged.
Fresh main 4f6c7566 also kept 6/6 in two constructions, two predictions and
four native runs. Its original first/repeated responses are
1,188.190292/75.380833ms, and the saved-file timing verifier passed. The earlier
candidate timing stays attached to its own run. Current installed concurrent
worker CPU averaged 72.69% of one core during a 1.796446292s process interval;
this includes setup, native children and EOF joining. Host CPU and model-only
RAM remain unobserved.
Commands, saved outputs and int64 readback
and Korean input/error guide
give the full scope.
What we want this to contribute
The 2026-10-03 research update
narrows the next work to useful bounded language features: choosing informative
execution inputs, preserving operation order, and reusing small typed assemblies.
An additive Go SDK probe-ranking API explores the first step with zero model
calls or training updates. Its compiler integration now has a
24-generation paired pilot:
the frozen compact bag-original FP32 judge made 24 actual predictions across the
model-enabled arms, and all configurations produced 48 compiled runs. A declared
Gooo reference activity supplied one new observation before construction. The
eight oracle-enabled generations matched 48/48 supplied expectations; controls
remain in the publication, with 84/144 matched across all 24 generations.
The original sparse selection case passed in every arm. This is one authored
task with two language views and two repetitions, with zero weight updates.
The Korean view needed extra search, and oracle-enabled generation added cost.
Compiler PR 1160 tracks
integration and deployment; the pilot pins compiler 2f02d244 and SDK v0.2.16.
Model artifacts and the existing task scores below remain attached to their
original observations.
An additional paired compiler experiment
uses the same frozen judge with SDK v0.2.17 and compiler 87d8afae:
96 generations, 120 actual model predictions and 192 compiled executions.
The compiler can retain candidate outputs and compare new reference observations
against them. Probe evaluations fall from 33 to 14 per oracle request, and all
24 fresh/reuse pairs preserve generated source, choices and runtime values.
Both oracle modes together meet 288/288 finite expectations; the full control
comparison is 396/576. This remains one authored task with Korean/English views,
six repetitions and zero training updates.
With the model enabled, observation-phase median falls from 0.603 to 0.333 ms, while complete codegen changes from 9.213 to 9.372 ms. Both oracle modes have 18.83 MiB median peak RSS. Process CPU changes from 92.35% to 93.02% of one core; whole-host utilization change is unmeasured. Reuse saves candidate evaluations; the end-to-end latency result is slightly slower in this collection. The compiler in that experiment still searches after a single candidate remains. PR 1162 tracks this opt-in integration. The existing model weights are preserved.
The next direct-projection experiment
completes 120 generations and 240 compiled executions, with 120 actual
model predictions in control arms. Complete observation can now select the
single surviving candidate before model loading or search. Resolved arms make
zero predictions, meet 144/144 finite expectations and match all 24 corresponding
cached-search bodies and runtime results. All oracle arms meet 432/432; the full
comparison including sparse-case controls is 540/720. Compiler
9158c3cd, PR 1164
adds this optional route with explicit source replay and skipped-work records.
For model-requested fresh processes, cached search vs direct projection has 9.491 vs 8.267 ms median codegen, 18.77 vs 17.29 MiB peak RSS, and 92.45% vs 88.01% CPU relative to one core. All samples are retained, including a slow initial control. This is one task, two language views and six repetitions with unchanged weights; host utilization change and broader speedup remain unmeasured. The model supplies a preference when choices remain, and a resolved source observation can complete this narrow construction without another prediction.
Smaller source-derived assembly recipes
The source-recipe pilot
uses the unchanged compact/bag-original/fp32 full-input export in actual native
construction. Gooo derives the typed base from its source body; a small recipe
names the permitted structural choices. Compiler
3a52232d, PR 1168
adds that input form and merged to dev as 79e7e20c. Its main promotion is
tracked in PR 1169.
Across 24 generations and 48 compiled runs, all 12 recipe/full-document pairs have equal expanded plans and emitted code. Compact authored JSON falls from 1,538 to 569 bytes (63.0%). Six real joint predictions each judge three choices; there are no training updates. Oracle/resolution arms satisfy 72/72 independent runtime expectations; sparse single-example controls satisfy 12/72, making the complete comparison 84/144. One authored subtraction task and repeated requests define this finite scope.
Own-model search has median whole-process generation time 8.143 ms for the full document and 8.598 ms for the recipe, including expansion. Peak RSS medians are 17.33 and 17.64 MiB; CPU relative to one core is 86.78% and 86.53%. Each cell has three samples. Whole-host utilization change and broader task accuracy are unmeasured. The pilot reduces authored representation while retaining the added processing cost. Initial collector accounting failure and all controls are public.
The language carries the plan, and the model supplies small local judgments.
The constant-body follow-up
uses the same frozen model with SDK v0.2.18 and compiler b742ba75 in
PR 1170. It repairs a
language gap: bodies that never read their declared input can now use typed
recipe construction. Six generations and twelve compiled runs met 36/36 finite
expectations, including both integer endpoints, with three actual predictions
and zero training updates. Three repeats per arm cover one authored body.
For that constant body, deterministic generation took a median 9.209 ms and 17.42 MiB peak RSS; model-connected generation took 9.731 ms and 19.31 MiB. Process CPU medians were 90.56% and 97.96% of one core; host utilization change was unmeasured. Isolated prediction median was 7.708 microseconds. Model ranking increased evaluated candidates from five to seven, so this task favors the deterministic route. The first proposed body failed in both arms; finite search reached the same successful source. All failed attempts and process records are retained. The result informs where a small model helps and where ordinary language machinery is already sufficient.
Condition chains using the same model
The condition-chain pilot
extends source recipes to if ... else if ... else. Compiler development source
bf51d65a lowers that form into existing typed nested branches, retaining
condition order, local scope and source binding. Its branch is public; main
deployment is tracked separately in the compiler wiki.
One authored clamp, one bilingual intention and three repeats per arm give six generations and twelve compiled runs. All 42 finite expectations match, including 24 selection-disjoint expectations and both int64 endpoints. The frozen own model makes three actual joint predictions; training updates are zero. Both arms choose the first candidate and emit equal source. Seven of eight declared combinations remain unattempted per request.
Deterministic/model generation medians are 7.958/9.227 ms, process peak RSS 17.45/17.91 MiB and one-core-relative CPU 90.64/91.58%. Prediction median is 7.917 microseconds. The first deterministic process takes 367.183 ms; all samples are public and its cause was not isolated. Cache state and whole-host CPU change were not instrumented. This already complete body favors deterministic assembly. The useful result is that a common source form now reaches both construction routes using the existing small runtime and unchanged weights.
An instruction-order feature preflight uses this frozen judge for 64 actual predictions. Four authored arithmetic pairs, two languages and four wrapper forms produce 32 pairs with identical V4 features and identical model predictions despite different requested operation orders. An experimental 128-byte directed-clause sketch distinguishes those 32 pairs; a repeated-clause counterexample still aliases and is retained. The feature kernel measures 372–619 ns/op with zero heap allocations on M4 / Go 1.27.1. It has no trained model head or new accuracy score, and this preflight includes zero optimizer updates or native Gooo executions. The current weights and input ABI remain attached to their original experiments.
The native order follow-up
uses main compiler 729482ed with four operation pairs, Korean/English views and
both requested orders. Across 96 generations, 192 compiled runs and 64 actual
predictions, first-attempt functional completion is 8/16 requests for deterministic
construction and 6/16 for each frozen positioned/bag model. With up to eight
candidate attempts, all three arms complete 16/16 requests and 128/128 finite
expectations each. All controls together meet 554/768 expectations. Training
updates are zero, and the two models also have different learned weights.
The positioned model distinguishes all eight reversed-instruction pairs in its features and probability distributions, while its first masks remain unchanged. Bounded-search generation medians are 8.129 ms deterministic, 9.352 ms positioned and 8.861 ms bag; total candidate evaluations are 24, 53 and 46. The first deterministic process took 574.741 ms and is retained. These are one sample per authored request, with process costs and limits in the raw publication.
A source-order counterexample reverses the source assignments while keeping the instruction fixed. The two correct source-relative choices are opposite, but the local root-order model inputs are identical. Both calls choose reverse: native outcomes are 8/8 and 0/8. The shared local judge lacks a distinction needed for this pair. These observations leave the released weights unchanged and retain bounded deterministic TDD as the working completion path.
The source-order preflight
now describes two source-derived assignment operations in 48 bytes per
alternative. Compiler ba00850f exposes its validated plan through the explicit
body-context --include-plan option in PR1174.
Across 48 exports from four authored families, the new description distinguishes
24/24 source-reversal pairs that have identical existing V3 local root inputs.
Renaming locals and commuting operands each preserve 16/16 description pairs.
Korean/English intent changes preserve 24/24 source-description pairs; this last
check measures the source channel, not language understanding.
Five M4/Go1.27.1 samples measure 28.02–29.31 ns/op and zero allocations for the descriptor alone. The complete context exports take a median 6.119 ms per fresh process; parsing, binding and validation remain separate costs. These observations add no model predictions, native runs or training updates. The new representation is isolated from released V3/V4 weights, and its trained-model quality is still unmeasured. It covers two simple integer updates; nested expressions and branches require further work. The initial collector digest-format failure is retained and covered by a regression. The next learning study must evaluate actual assembled bodies and additional attempts with the same deterministic-search budget.
Our aim is to make intent, construction and observed behavior travel together as a program evolves. We measure complete finite behavior, partial coverage, extra attempts, unresolved obligations, time and memory. These measurements help decide which language features and model changes to develop next.
For the full-input learning study, a first path is complete when it satisfies all 16 supplied examples for that input. A result of 368/512 therefore describes that specific authored task collection. Larger programs, broader types and reusable learned constructions are the next language questions.
Model probabilities describe which permitted path the judge favors. Finite
coverage counts the supplied examples that pass. An unresolved observation
stays UNKNOWN in its own receipt dimension. Keeping these meanings visible
helps the system identify the next useful experiment.
Choose an entry point
| What you want to do | Compatible starting point |
|---|---|
| Generate with SDK v0.2.14 or later | Pinned compact V3 bundle in the usage example below |
| Explore the new V3/V4 arithmetic contract in Go | SDK v0.2.15 and explicit-arithmetic models |
| Generate with the new V4 contract | Integration on main fc0e99c4 through merged PR 1159; native evidence |
| Follow the current language work | Language guide, experiment repository and the dated appendices here |
SDK release source 59c8d342da4475506b90954469aa201f85cadeb3 passed complete
replay on darwin/arm64 and linux/amd64: each made 36,864 predictions over 18,432
frozen inputs, matching intermediate values, probabilities and full rankings.
Both original SDK reports and
Linux CI
are public. This stage performed zero training updates. The native integration
observation below follows that library replay.
New native generation and execution observation
On sixteen known bilingual tasks, four training arms × three precisions × two layouts plus a disconnected arm completed 400 generations and 800 compiled runs. All 9,600 supplied expectations and 192 representation pairs matched. An independent reader checked the original source/result bindings, 1,760 progress records and 432 feedback records. Total actual predictions were 816; the disconnected requests made zero model calls.
Bag-original FP32 completed the first path for 16/16 requests in this integration slice. Its broader development result stays 368/512. Compact median prediction time was 8.33 µs, whole codegen 10.00 ms and compiler peak RSS 17.78 MiB. Process CPU was 78.84% of one core; host-wide utilization change was unmeasured. The disconnected arm completed after 96 extra candidates, with 9.80 ms median codegen. This task size shows reduced search attempts and similar process latency.
Full native results, every variant and original records
retain the timing conditions and known-task scope. All 400 receipts keep
permission_boundary unresolved. The observation used compiler e461c1d and
SDK v0.2.15, with zero training updates and no new default checkpoint. Compiler
integration merged to dev in PR 1158 and main in PR 1159 above.
New full-input research exports — 2026-10-03
The full-input appendix contains four freshly initialized shared judges, each exported as FP32, PTQ and QAT in expanded and compact layouts. Complete source-derived input and Korean/English instructions are retained. The comparison changes position-based features versus global fragment counts, and original versus varied training wording.
| FP32 arm | Original first-path complete /512 | Extra ranked attempts | Both KO/EN valid /256 |
|---|---|---|---|
| positioned-original | 113 | 1,469 | 1 |
| positioned-varied | 266 | 528 | 75 |
| bag-original | 368 | 186 | 180 |
| bag-varied | 320 | 326 | 132 |
Bag-original FP32 reaches 358/512 and 368/512 on the two preregistered new wording forms. Bag-varied falls to 173/512 on the new suffix. Bag-original PTQ and QAT reach 94/512 and 287/512 on original inputs. Its global fragment representation also maps two distinct operation orders to equal features; the appendix retains that explicit counterexample. These are previously observed authored source tasks.
The Go audit reconciled all 6,400 MPS updates and made 25,824 real predictions across export parity, compact parity, calibration and three development forms. All twelve exports and complete raw evidence are public. Optimization took 42.18 seconds; average process CPU was 49.83% of one core and peak RSS was 1,809,678,336 bytes. CPU usage describes the optimizer process. Model inference and native generation have their separate measurement above.
V4 compact artifacts use triple_semantic_context_bag_v4_joint_v1 and the
research Go runtime.
The matching SDK v0.2.15 is published with complete numerical replay and the
native integration observation above. Remaining split/form comparisons and new
intentions are subsequent steps. Later historical sections retain the measurements
of their pinned earlier models.
The subsequent Linux comparison
retained the same first selected mask on all 18,432 development rows, with 272
different complete candidate orders. Small score deltas changed partial-completion
curves, so the exact cross-platform comparison failed. Its complete observations
are public. The explicit-arithmetic follow-up
then made 73,728 predictions per platform with unchanged weight bytes and zero
training updates. All 18,432 paired inputs now have identical hidden/logit/
probability bits and complete rankings under float32_separate_v1. Original
first-path completeness and extra ranked attempts stay unchanged; partial
coverage shifts in both directions compared with legacy arm64 arithmetic.
Converted metadata and complete four-lane journals are public in that appendix.
The SDK replay reproduces these observations. The native stage above subsequently
completed generation and execution with this arithmetic contract.
How the model is used
The compiler derives context from an original Gooo body, its declared typed alternatives, and complete Korean/English intentions. The model ranks complete paths involving conditions, assignments, references, nested branches and arithmetic. Finite expectations evaluate a proposed body. Later predictions may include actual failed-case feedback and rank the remaining choices.
The model makes small decisions inside a supplied plan. The compiler supplies the source binding, allowed alternatives, attempt budget and deterministic continuation. Disconnected or unsupported inputs keep their complete text and continue with zero model predictions. Generated Go then has a separate replay, build and execution stage.
The current artifacts target authored Integer → Integer tasks. Broader source discovery, larger programs, reusable abstractions and Korean/English meaning alignment are active development areas.
Original V3 architecture, training and compatible files
The shared model reuses a 256 → 8 → 2 local network over three ordered decisions: 2,072 trainable parameters. The dense control has 18,656. Both were freshly initialized using the same Go initializer recipe and trained on Gooo-derived data. Each arm completed 800 FP32 and 800 QAT updates on local MPS, totaling 3,200 optimizer updates. PTQ exports come from the FP32 models. That total belongs to the original shared/dense comparison. The later four-arm full-input study above adds 6,400 updates and twelve exports in its own appendix.
The data contains 1,024 training, 256 calibration and 256 development program groups. The development set has 512 Korean/English views and had been observed in earlier studies. The source/teacher feature bank has 10,739 initial and feedback rows. Quality numbers below describe this known development cohort.
| Files in this Hub repository | Purpose | Go loader |
|---|---|---|
research/compact-runtime-20261003/models/{fp32,ptq_ternary,qat_ternary}/ |
Current compact shared representation | jointdecision.LoadSharedThree |
models/shared-local/{fp32,ptq_ternary,qat_ternary}/ |
Original expanded shared exports | jointdecision.LoadThree |
models/dense/{fp32,ptq_ternary,qat_ternary}/ |
Matched dense controls | jointdecision.LoadThree |
evidence.zip, go-audit.json, manifest.json |
Original training and native study evidence | Go audit tools in the research repository |
research/compact-runtime-20261003/ |
Exact representation parity and kernel measurements | SDK v0.2.14-experimental |
research/compact-native-20261003/ |
Actual paired Gooo generation and compiled execution | Compiler and independent receipt reader |
research/full-input-initial-20261003/ |
Four-arm study: full inputs, twelve exports and original observations | Research runtime and SDK v0.2.15 |
research/full-input-platform-20261003/ |
Retained Linux replay and complete ranking diagnosis | Offline Go diagnostic reader |
research/full-input-separate-20261003/ |
Explicit arithmetic, paired observations and converted metadata | Research runtime and SDK v0.2.15 |
research/full-input-native-20261003/ |
400 generations, 800 compiled runs, 192 matched pairs and independent receipt consumption | Compiler e461c1d with SDK v0.2.15 |
research/full-input-sdk-20261003/ |
Complete local/Linux SDK replay reports | SDK release source 59c8d34 |
Compact files use gooo/tiny-shared-three-choice-path-model/v1 metadata and
three tensors: 8×256 input weights, eight biases and 2×8 output weights.
The total feature input is 768 values; a caller workspace holds 24 hidden
values and eight mask scores. Go loads model.json and its adjacent weights.
With a compatible Gooo binary and an authored three-choice source/plan:
hf download asketeddy/gooo-shared-judgment-tiny-v1 \
--revision 985999a89caba6a31cc7147f66ba29a5ce76a1d9 \
--include 'research/compact-runtime-20261003/models/fp32/*' \
--local-dir ./gooo-models
gooo body-codegen --json --activity ChoosePath \
--path-plan plan.json \
--path-model gooo-models/research/compact-runtime-20261003/models/fp32/model.json \
--path-step-attempts 1 --path-feedback-rounds 7 source.gooo
source.gooo and plan.json above are caller inputs. Complete measured examples
and expectations are in the native evidence archive. The
compiler integration guide
describes the input contract. Offline PyTorch performs training/export; the Go
SDK performs local inference and path search.
Decision quality
Each development view has 16 finite expectations. This table ranks all eight masks from one initial prediction and evaluates their stored finite targets. Adaptive failure feedback is measured separately in actual codegen.
| Export | First choice complete /512 | Initial cases matched /8192 | Extra ranked attempts | Complete by budget 4 /512 | EN/KO mask disagreements /256 | Same-mask both wrong /256 |
|---|---|---|---|---|---|---|
| Dense FP32 | 95 | 2265 | 1572 | 289 | 256 | 0 |
| Shared FP32 | 113 | 2592 | 1469 | 319 | 255 | 0 |
| Dense PTQ | 74 | 1967 | 1595 | 285 | 45 | 180 |
| Shared PTQ | 80 | 2052 | 1439 | 336 | 136 | 98 |
| Dense QAT | 82 | 2050 | 1584 | 293 | 240 | 12 |
| Shared QAT | 89 | 2191 | 1512 | 327 | 255 | 1 |
Shared FP32 improves first-choice completion 18.55% → 22.07%, finite-case
matches 27.65% → 31.64%, and extra ranked attempts 1,572 → 1,469.
Its development passing-set negative log likelihood worsens 975.663 → 992.402.
Comparison/assignment, nested branches and successive assignments need more
extra attempts. The full per-family counts remain in go-audit.json.
Paired-language disagreement remains 255/256 for shared FP32. Agreement also needs a correctness check: dense PTQ gives the same wrong mask on 180 pairs. The subsequent full-input study above measures both language consistency and membership in the valid-path set. Full-budget success here depends on the authored finite search space containing a passing alternative.
Input sensitivity diagnosis
A subsequent fixed-weight study varies five authored instruction forms while keeping the original Gooo source, eight paths and finite expectations constant. The collector made 92,160 actual Go predictions across all 3,072 corpus views and six exports. An independent reader replayed all 92,160 predictions with zero numerical difference locally and recomputed every condition's finite and bilingual summary. These replay calls are recorded separately from collection.
The Linux CI replay also reproduced all chosen masks and 30 condition summaries. Its largest floating-point difference was 0.000001430511474609375, within the declared 0.00001 tolerance. The run's 13 jobs completed successfully.
For shared FP32, the 512 development views give:
| Complete input form | First-choice complete /512 | Extra ranked attempts | EN/KO disagreements /256 |
|---|---|---|---|
| Original development prefix | 113 | 1,469 | 255 |
| Bare authored instruction | 480 | 38 | 32 |
| Calibration prefix | 92 | 1,372 | 250 |
| Development phrase as suffix | 233 | 662 | 180 |
The bare form matches the format of the original training instructions. Its result describes a controlled intervention on an already observed cohort; original-input quality remains 113/512. Production Gooo keeps the complete caller text. The completed initial full-input comparison above varies training phrasing and positional features while preserving full input and source binding.
The source v3 intent encoder uses four relative-position buckets for byte
fragments. Added prefixes alter the fragments, their buckets and normalization.
The study identifies sensitivity to this combined change; further experiments
will separate those factors. All six models, all forms and regressions are in
the report
and the Hub appendix research/bilingual-wrapper-20261003/. Its 30-member archive
includes every full constructed input and prediction, the exact source dataset,
six small existing models and the independent reader. The appendix adds zero
training updates. The model weights at their existing paths remain unchanged.
Actual Gooo generation and execution
The latest compact study used 16 bilingual views, three model variants and two representations, for 96 generations and 298 real model predictions. Of those predictions, 202 used finite-failure feedback. Each generation was immediately followed by a native build and two executions: 192 compiled runs.
All 2,304 supplied finite expectations passed, producing 4,608 ordered values. Of these expectations, 768 use inputs absent from the current selection suite; their relationship to model training remains part of the dataset scope. All 48 unseeded expanded/compact pairs retain matching generated source, search/feedback semantics and ordered outputs.
An independent compiler reader consumed all 96 runtime receipts, 596 progress
records and 202 feedback records. The first unresolved completeness dimension
is permission_boundary throughout. Gooo keeps each observation axis and its
remaining questions alongside the finite functional results.
This compact study uses compiler 7db19b6bc9a2909aa059f1265d39a539c3573a57,
SDK v0.2.14-experimental and Go 1.27.1 on local darwin/arm64. The earlier
architecture study separately made 320 predictions in 96 generations.
Earlier quality/native report
and compact native report
retain both experiments with their own identities.
Storage, speed and CPU observations
| Compact variant | Weight bytes | Resident tensor bytes | Native-study prediction median | Fresh-process codegen median |
|---|---|---|---|---|
| FP32 | 8,288 | 8,288 | 23.709 µs | 28.927 ms |
| PTQ ternary | 446 | 2,096 + 8 scale bytes | 25.000 µs | 27.765 ms |
| QAT ternary | 446 | 2,096 + 8 scale bytes | 30.208 µs | 29.511 ms |
Caller workspace is 3,200 bytes. Five trits fit in one byte, giving 1.6 stored matrix bits per weight; biases and scales have their own storage. The runtime decodes matrix values into int8 arrays and computes with int8/float32. Valid warmed kernel calls allocate zero heap objects in the measured contract.
The separate kernel audit checks all 10,739 frozen states across three variants: 64,434 actual predictions, with bit-for-bit agreement in features, hidden values, logits, probabilities and chosen mask. Compact conversion preserves the existing shared model; this step performs zero optimizer updates.
Native timings have 16 generations per representation/variant. Prediction timers exclude loading, startup, code generation and builds. QAT's complete codegen median increased from 29.237 to 29.511 ms after compact conversion. Cold-cache outliers remain recorded, including the first 6.45-second native build.
Codegen process CPU/wall-time medians are 83.7–84.6% of one core. Generated program RSS medians are 4.20–4.23 MiB, covering the compiled arithmetic child. Inference-process RSS and whole-host CPU increase were unmeasured in this study. Retained serving and additional hardware need their own controlled measurements.
Reproducibility and development
Every public evidence bundle has a closed byte inventory. The compact native archive has 684 members, 3,040,850 compressed bytes and 12,474,521 decoded bytes. Original failed preflights, negative variants, complete inputs and measurement outliers are retained. Public packaging checks private paths and credential patterns, including decoded embedded parent receipts.
Training source: 0f3249596b4dacdb34241e3b5e8b4cafac58b0df.
Original independent Go audit: 8398f1994d65b75af0483db3fefd637eec99ced4.
Compact native collector: 7e9be4489d15663cbac8f3e71bc02d2e4649dfe5.
Immutable compact model edition: 985999a89caba6a31cc7147f66ba29a5ce76a1d9.
Immutable native appendix edition: 37db3a1f8a06669081a370ee1e7b23a64cbc00fb.
Next work connects the published SDK to native generation and execution while continuing full-input evaluations and operation-order representation work. Korean/English intent alignment, more expressive Gooo construction, per-axis before/after observations and serving costs guide the wider development. Repeated successful structures may later become reusable language abstractions. The dated research protocols define each experiment's budgets and acceptance criteria; the compiler retains deterministic continuation between model versions.
Research acknowledgments
- Solar-Lezama et al., SKETCH (2006): partial programs completed under a specification inform our view of bounded construction. Combinatorial Sketching for Finite Programs.
- Ellis et al., DreamCoder (2020/2021): neural program search and learned reusable abstractions inform the longer-term language/model direction. Paper.
- Balog et al., DeepCoder (2016/2017): learned program properties guide synthesis search. It informs our question of how much a small judgment can reduce construction attempts. Paper.
- Laya / ConvAI Innovations: its structured decision interface was used in our early Gooo experiments. The present weights use fresh initialization and Gooo-specific training. Project.
- Ma et al., BitNet b1.58 (2024): ternary-weight research motivates exploring low-storage models with explicit quality and runtime measurements. Paper.
- W3C PROV-O (2013): entity, activity and agent relations inform how we trace inputs, construction and observed results. Recommendation.