A small local judge for Gooo program construction

Keep a compiled Gooo graph for current inputs

Compiler main 7acccf564b3def31ddd368050c2288f05bd7076f is merged after all12 authoritative CI jobs and both local executables are installed. Ordered current-input suites retain one compiled graph within an owned request. The optional model guides construction once; later runtime suites make zero model calls. Keeping an assembled tool on a workbench lets the next piece of material arrive immediately.

Current-input observations retain the original clean candidate and a separate fresh installed cohort. Installed controls keep the same program identities and actual finite results in all20workloads. Three paired saved FP32-origin workloads cost a median2,327.40ms across six stateless commands versus497.14ms in one retained command. CPU time is1.45s/0.36s and peak RSS82.88MiB/82.14MiB. Startup count differs. Each cohort retains50commands, 120suite frames and240native executions; original timings remain separate. Three changed-input suites repeated twice retain4/4,2/2,1/2,4/4,2/2,1/2 named expectations; the third label is deliberately different. All record fields match. The Go reader independently checks current inputs and recomputes finite counts.

Fresh installed FP32 prediction median is18.792µs across three constructions. One PTQ and one QAT pilot use18,752 resident tensor bytes each and retain the same program/current results. Their matrix encoding is five trits per byte, about1.6 stored matrix bits; each file is3,854bytes including FP32 biases. Decoded tensor residency is separate. The earlier2-bit wording is corrected from actual files. Weights and training are unchanged. Compiler PR1228 contains the merged release; research PR13 passed all23 CI jobs and merged. Independent Linux readback rechecks eight native captures. Followup PR14 also passed all23 CI jobs and merged into research main d799e9cd. Its Go text reader shows current percentages, mismatches and separate build/native-run times. Installed-source Linux readback includes all eight downloaded native responses and independent recomputed reports.

한국어: 작은 모델이 한 번 조립 방향을 제안하고, 만든 프로그램은 다음 입력을 받아 바로 실행합니다. 같은 공구를 작업대에 두고 재료만 바꾸는 방식입니다. 틀린 기대값도 실제 값과 함께 남겨 다음 개선에 사용할 수 있습니다. 한 작성 예제의 반복 비용을 측정했으며 전체 호스트 CPU 사용률은 아직 측정하지 않았습니다.

Record field assembly — installed controls

The field assembly feature is merged into compiler main f6da667e11951b6939fc5e30442e01a09ca06e85 after complete PR1226 CI and source-bound result verification. Both local executables are installed. The separate installed cohort retains 24 fresh controls and six saved replays, matching the candidate's selected source/Go hashes, rankings, masks and finite results in all 30 observations. Fresh full-budget prediction median is31.625µs; graph generation is9.890ms/9.337ms and whole command is323.462ms/324.194ms for model/deterministic controls. Peak process RSS is82.94MiB/82.42MiB; CPU time is0.24s in each mode. These resources include native Go builds and children; host-wide CPU changes remain unmeasured.

The Go completeness reader now shows selection case/field ratios and each remaining field value alongside actual compiled runtime scores. Saved composition replay made zero new predictions and built the native graph again in that earlier measurement. The subsequent candidate above retains the compiled graph. This update publishes compiler integration and observations with unchanged weights.

Earlier candidate and ternary observations

Gooo can now declare alternatives for the actual expressions inside a record constructor, together with typed JSON cases and a small attempt budget. The field assembly observation retains 24 fresh controls and six saved full-budget replays of one authored two-node graph. With 1/2/4/8 attempts, both modes complete 6/9/12/15 of 15 expected selection fields. Whole selection cases reach 5/5 and compiled execution reaches 14/14 named outputs and 21/21 record-output fields at full budget.

These are clean candidate measurements at compiler 48f70275, before main installation. Compiler PR1225 contains the feature and subsequent standard-library-only syntax and formatted JSON replay fixes. The existing frozen own FP32 weights are unchanged. A named field-ordinal context transfers the integer-choice model to field alternatives; case inputs and expected outputs are excluded. The model and deterministic controls reached equal finite counts, with no attempt reduction on this source shape. Full-budget prediction median was 29µs and command medians were about 349ms/342ms. Field-specific training remains subsequent work.

The separate ternary pilot executes one PTQ and one QAT request on clean compiler cb316851. Both reach 15/15 selection fields and 14/14 compiled outputs with 18,752 resident tensor bytes each, versus 74,624 bytes in the FP32 cohort. The actual matrices pack five trits per byte, approximately1.6stored matrix bits. Each complete weight file is 3,854bytes including FP32 biases. The earlier2-bit description is corrected; 1.58bits describes ternary information content. The two pilot timings and decoded tensor counts leave whole-process memory and broad speed improvements unmeasured.

한국어: 기록의 제목·상태·사유 안에 후보식을 심고, 작은 예산에서 얼마나 채워지는지 측정했습니다. 시도 1·2·4·8회에 따라 기대 필드 완성률은 40·60·80·100%였습니다. 기존 자체 모델을 실제로 호출했지만 이 예제의 완성률과 시도 수는 결정론 방식과 같았습니다. 후보 관측과 설치 상태를 구분해 기록하며, 다음 학습은 다양한 필드 문맥에서 같은 예산으로 더 많이 채우는 것을 목표로 합니다. 새 가중치 학습은 이번 관측에 포함되지 않았습니다.

Language and earlier observations

Updated 2026-10-04. This model ranks eight permitted Gooo body paths built from three binary decisions. The project explores a language and a small model working together: Gooo provides the construction plan, the model suggests a route through it, and the compiler checks the assembled body and generates Go. A Go experiment runner then builds and executes the resulting program.

Like a workshop, the plan describes what the parts mean and how they may fit. The model helps choose the next assembly; observed test failures guide another attempt. A receipt connects the original intention to the choices and result.

The compiler now accepts an activity's assembling block for permitted source choices, Korean/English intent, finite examples and search budget. The regular commands expand it into typed construction and immediate compiled execution. The latest two-task observation retains weaker initial model ranks, bounded test recovery and measured costs.

Selected bodies now return as reusable Gooo checkpoints with the original planning baseline and picked paths. The source reuse observation chains three generations per activity/mode: 12 constructions, 48/48 selection expectations and 108/108 independently executed expectations. Six real predictions used these unchanged weights; twelve realization replays made zero predictions. These are two repeated tasks, with no new training or measured accuracy gain. The feature merged into compiler main f62eb5f0 (PR1218, CI37190298986) and both executables were installed. The installed-source recheck repeats those twelve controls and records eight parallel checkpoint requests, runtime72/72, four actual model predictions and normal EOF completion.

Native composition now connects actual generated activity bodies through typed Gooo bind edges. Compiler main 2d1ce2d0 is merged and installed after its own complete CI. The fresh installed controls retain588/588 expectations on one repeated graph, two concurrent requests98/98, three extra input combinations42/42 and a deliberately changed expectation48/49. The exact public FP32 model location includes a two-file download command. Weights are unchanged; first candidates, additional search and slower fresh timing observations remain visible.

The language now supports source-ordered scalar input joins. A body such as Add(Integer, Integer) receives two separate inputs, and a mixed-type body can combine bound producer values with values supplied by the caller. The input-join observation records 588/588 named outputs, 1,008 input-slot observations and 504 bound deliveries across twelve repeated controls of one graph. Three fresh model requests made three actual predictions with unchanged FP32 weights; saved replays made zero. The model checked two candidates versus eight deterministic candidates, while whole-generation medians were 9.711ms and 6.396ms respectively. Development PR1221 and main PR1222 passed their own complete CI and independently verified proofs. Compiler main 1bef47fc and both executables are installed. The separate installed-source controls retain588/588 with matching source/Go/driver hashes, three actual predictions, concurrency98/98, additional scenarios42/42 and the changed expectation48/49. Installed fresh generation medians are9.449ms with the model and8.480ms deterministically. Original candidate measurements retain exact source 5af78bcb.

Gooo bodies now carry source-declared record values such as a Candidate with a title/state and a Review with a summary/reason. The record observation retains 432/432 named expectations across twelve repeated controls of one six-activity graph and 1,152 actual input/output field observations. The own model ranks the scalar Score assembly; record bodies follow Gooo source. Its first candidate met0/6, the next6/6. Median prediction was26.9µs, while graph generation took 12.80ms with the model versus11.77ms deterministically. Saved replay makes zero new predictions. Separate concurrent requests met72/72 and eight further runtime-only cases through saved compositions met96/96. One intentionally wrong field expectation gives35/36 and retains the actual difference. Source candidate fa3c8927 is retained separately from installation. Development PR1223 and main PR1224 passed their own complete CI and independent proof verification. Compiler main b629a664 and both executables are installed. Fresh installed observations retain432/432 outputs and1,152 field observations with the same program hashes. Actual prediction median was29.7µs; fresh generation was13.17ms versus12.01ms deterministically. A Go result reader shows35/36 and Propose.state when the state expectation differs. An actual replay omitting record expectations keeps 36 output fields unobserved with a null percentage. A second graph accepts external records and replaces local values, meeting8/8 in fresh/saved controls. Weights remain the frozen published FP32 files; this update publishes compiler integration and finite observations.

Component Public home
Introductory Wiki in Korean Gooo: language, models, metrics and research
Gooo language, compiler and execution meta-ontology-go
Go inference and typed search gooo-decision-runtime
Training, raw evidence and comparisons gooo-neural-decision-experiments
Project direction in Korean 언어·모델·실험 안내

한국어 요약

기록의 필드까지 실제 본문에서 전달하기

Candidate에 제목·상태를 담고 다음 본문이 Review를 만드는 경로를 구현했습니다. 서류를 다음 작업대에 전달하듯, 실제 값과 필드의 정체성을 연결마다 남깁니다. 한 그래프의 열두 반복 제어에서 출력 432/432와 필드 관측 1,152개를 확인했고, 상태 기대값 하나를 바꾸면 35/36으로 기록됩니다. 자체 모델은 정수 조립의 후보 순위를 제안합니다. 기록 구성과 읽기는 Gooo 선언을 따르며, 기존 가중치를 사용했습니다. 실제 추론 중앙값은 26.9μs지만 생성 전체는 모델 12.80ms·결정론 11.77ms로 모델 쪽이 조금 길었습니다. 사용 흐름·완전성의 분모·원본을 함께 공개합니다. 후보 관측 소스 fa3c8927을 유지하고, 개발·배포 검사를 마친 main b629a664로 두 실행기를 설치했습니다. 설치본의 별도 반복 제어도 432/432와 동일한 생성 지문을 유지했습니다. 실제 추론 중앙값은29.7μs이며, 새 기록 입력 예제와 필드별 완전성 요약을 추가했습니다. 기대값을 생략한 실행에서는 해당 출력 필드36개를 미관측으로 표시합니다.

같은 타입의 입력도 각 위치에 연결하기

새 입력 연결은 작업대에 번호가 붙은 소켓을 두는 방식입니다. Add(Integer, Integer)의 input0과 input1은 각각 다른 생산자의 실제 값을 받습니다. Label(Integer, Boolean, Text)는 정수만 연결하고 나머지 입력은 호출자가 제공할 수 있습니다. 선언 순서가 의미 모델·IR·Go 인자 순서까지 이어져 뺄셈과 같은 순서에 민감한 본문을 구분합니다.

입력 연결 관측과 따라 할 명령을 공개했습니다. 같은 그래프를 반복한 열두 제어에서 출력 기대값 588/588, 입력 관측 1,008개·연결 전달 504개·실제 실행 24회를 확인했습니다. 자체 모델은 Left 본문을 조립할 때 세 번 호출했고 저장 기록은 추가 추론 없이 재실행했습니다. 모델의 첫 후보는 0/6, 다음 후보는 6/6 예시를 충족했습니다. 후보 확인은 모델 2개·결정론 8개였지만 전체 생성 시간 중앙값은 각각 9.711ms·6.396ms로 결정론 경로가 더 짧았습니다. 의도적으로 틀린 기대값은 48/49로 기록했습니다. 관측 수는 같은 그래프의 반복 횟수와 함께 읽습니다.

측정 소스는 5af78bcb이며 배포 검사에서 준비 함수의 분할을 더 작게 할 필요가 확인되어 6b6c86f2로 정리했습니다. 개발 #1221과 main #1222을 각 전체 CI와 독립 증거 확인 후 병합하고 main 1bef47fc로 CLI·실행기를 함께 설치했습니다. 설치본의 별도 관측에서도 같은588/588·입력1,008개·연결504개와 동일한 생성 지문을 확인했습니다. 모델은 세 번 실제 호출했고 저장 재실행은 추가 추론0번입니다. 동시 요청98/98, 추가 입력42/42, 틀린 기대값48/49도 기록했습니다. 설치본의 전체 생성 중앙값은 모델9.449ms·결정론8.480ms였습니다. 모델 가중치는 그대로 두었고 새 학습은 수행하지 않았습니다. 입력 순서와 실제 전달은 Gooo가 담당하고 작은 모델은 허용된 본문 선택의 순서를 제안합니다.

앞선 관측: 여러 활동을 연결한 프로그램으로 실행하기

컴파일러 개발 후보 c00bd714의 body-compose는 Gooo bind 연결을 따라 실제 본문을 함께 실행합니다. 정수 조립 두 개, 불리언 판단과 부정, 값의 분기 전달, 별도 텍스트 처리까지 일곱 활동을 한 번의 Go 빌드로 연결했습니다. 모델은 요청 안에서 한 번 로드하고 두 활동에 두 번 판단했습니다. 저장한 전체 조립 기록의 재실행은 추가 추론 0번입니다.

실제 사용한 파일은 공개 three-choice 모델의 고정 FP32 리비전에 있습니다. hf download asketeddy/gooo-three-choice-feedback-tiny-v1 models/set-feedback/fp32/model.json models/set-feedback/fp32/weights.bin --revision 24238ca67048b36bc305731799985271bb752dd1 --local-dir gooo-models로 두 파일을 받고, body-compose --model gooo-models/models/set-feedback/fp32/model.json에 연결합니다. 메타데이터 1,215바이트·가중치 74,624바이트이며 이번 관측의 실제 지문과 일치합니다.

새 생성 여섯 건과 별도 저장 기록 재실행 여섯 건에서 유한 실행 기대값 588/588, 중간 값 588개, 연결 전달 420개와 실제 실행 24회를 기록했습니다. 첫 모델 후보는 0/6·1/3으로 실패했고 유한 테스트 탐색으로 주어진 예시를 충족했습니다. 가중치는 유지했습니다. 조립 중앙값은 모델 13.527ms·결정론 13.011ms이며, 빌드와 두 번 실행을 포함한 명령 중앙값은 조건별 337.886–469.090ms입니다. 메모리와 CPU 기록에는 Go 빌드 자식도 포함됩니다. 세 번씩의 고정 순서 관측이며 속도 개선이나 전체 컴퓨터 CPU 증가량을 판단하지 않습니다. 원본·재현 도구·측정 범위, 컴파일러 PR1219, 연구 PR6. 이 기록은 main 배포 전의 위 개발 후보에서 수집했습니다.

같은 기능을 main 2d1ce2d07eaa29f00a412cbaa22c96008f0b7250으로 병합하고 두 실행기를 설치했습니다. 설치 후 기록은 같은 열두 제어의 588/588과 동일한 생성 지문, 동시 두 요청 98/98, 추가 세 입력의 두 경로 42/42, 틀린 기대값 사례의 48/49를 각각 남깁니다. 설치 후 조립 중앙값은 모델 36.933ms·결정론 32.216ms, 전체 명령은 조건별 721.523–778.166ms였습니다. 측정 시점과 동시 작업이 다르며 속도 개선을 결론내리지 않았습니다. 실행 후 CPU 스냅샷은 26.47%·17.70%이고, 모델에 따른 CPU 증가량을 구분할 실행 전·중 대조군은 없습니다.

선택한 본문을 다음 개발에 이어 쓰기

기준 본문과 선택 경로를 Gooo에 함께 기록하도록 발전시켰습니다. 모델이 고른 지역값을 원본에 넣었을 때 다음 생성의 기준이 어긋나던 문제를 다룹니다. body-realize는 기록을 다시 구성해 Gooo 파일로 저장하며 추가 모델 호출은 0번입니다. 실제 자체 모델을 사용한 두 과제의 열두 번 생성·제어 관측에서 별도 실행 108/108을 확인했습니다. 생성 중앙값은 모델 경로에서 두 과제별 15.696ms·13.214ms, 저장은 15.738ms·13.984ms였습니다. 명령 시작과 자료를 읽는 비용이 포함되며 과제마다 세 번의 관측입니다. 원본과 범위를 함께 공개합니다. 가중치 학습·정확도 개선과 구분해 언어의 재사용성을 확인한 작업입니다.

Gooo 선언 안에 조립 의도를 함께 두기

computes 뒤에 assembling 블록을 두면 의도, 선택할 소스 위치, 허용된 대안, 유한 입출력 예시와 시도 예산을 Gooo에 함께 적을 수 있습니다. 선언은 의미 IR에 남고, 선택된 본문은 Go 생성과 실제 실행으로 이어집니다. 작은 조립 지시서를 본문 옆에 붙여 두는 방식입니다. 구문과 실행 예제.

gooo body-codegen --json --activity Qualified --path-model /path/to/model.json \
  examples/body-codegen/source-assembly.gooo.fixture

모델 경로를 생략하면 같은 선언으로 결정론적 탐색을 합니다. 실행할 때에는 별도 입출력 예시를 제공해 생성 단계의 예시와 실제 Go 동작을 나누어 관찰합니다. 기존 모델의 입력 계약을 이용하며, 이 구문 때문에 추가 학습을 요구하지 않습니다.

이번 관측은 기존 자체 set-feedback FP32 joint 모델을 사용했습니다. 두 과제·두 작성 방식·모델/결정론적 경로에서 32회 조립을 수행했고, 선택 예시 128/128개와 별도 실행 예시 288/288개를 충족했습니다. 모델의 첫 후보는 각각 0/5와 1/3으로, 두 과제에서는 결정론적 순서보다 불리했습니다. 유한 테스트 탐색을 통해 완성됐고 이 실패와 추가 후보 수를 다음 학습의 재료로 남겼습니다.

선언 기반 모델 요청의 반복 응답 중앙값은 27.0–28.6ms, joint 판단 중앙값은 약 20µs였습니다. 원본 계약을 재구성하고 확인하는 비용 때문에 별도 계획 방식보다 응답 시간이 소폭 늘었습니다. 학습과 가중치는 유지했습니다. 원본 예제·측정 범위·자체 모델 출처.

별도 후속 실험은 첫 후보의 실패와 실제 개발 CI 상태를 다음 판단의 문맥으로 전달했습니다. 올바른 기대값으로 두 요청이 각각 선택 5/5와 실제 실행 11/11을 충족했습니다. 앞서 다른 예제의 기대값을 섞어 얻은 1/11도 보존했습니다. 위 32회 집계와 구분하고, CI 상태의 효과는 별도로 비교하지 않았습니다. 선언 기반 조립 관측과 후속 문맥 부록.

해당 언어 변경은 개발 #1215·main #1216의 자체 CI와 독립 증거 확인 후 병합했습니다. main 8a08adbb에서 CLI·병렬 실행기를 함께 설치했고, 별도 16개 요청에서 제공한 실행 기대값 144/144와 정상 EOF 종료를 확인했습니다. 새 설치본 모델 재사용 응답은 두 예제에서 34.745·33.816ms의 단일 관측입니다. 설치 버전·후보 검사·시간 기록.

최근 재현 검사

과거 계산 방식의 macOS·Linux 기록을 각각 고정 원본과 대조하는 Go 검증기를 추가했습니다. 새 실행은 환경마다 18,432개 개발 기록과 24개 compact 형식 모델 파일을 모두 재현했습니다. 원래 두 환경에서 달랐던 요약 101개 항목과 후보 순위 272개는 양쪽 값과 입력 해시를 함께 보존합니다. 검증기는 모델을 호출하지 않으며 새 감사 실행의 Go 추론 호출은 환경마다 25,824회입니다. 원래 가중치와 기대값은 유지합니다.

Linux 검사 37126899513의 19개 작업이 통과한 뒤 연구 PR #1을 병합했습니다. 새 산술의 플랫폼 간 검사는 계속 별도로 실행합니다. 새 원본과 재현 방법을 공개했습니다. 당시 설치된 컴파일러 main ed2c2cac는 SDK v0.2.20-experimental을 사용했습니다. 최근 파일 연결과 설치 상태는 아래 갱신에, 이전 연결 실험의 버전·결과는 각 날짜의 기록에 남아 있습니다.

Gooo 선언을 설계도, 작은 모델을 조립 순서를 고르는 장치로 생각하면 됩니다. 현재 모델은 세 가지 이진 판단으로 이루어진 여덟 경로에 순위를 매깁니다. 컴파일러는 조건식·할당·분기 등을 조립하고, 실패한 입출력 예시를 다음 시도의 문맥으로 전달합니다. 생성된 Go의 빌드·실행 결과와 남은 의문도 기록합니다.

개발은 Laya의 구조화된 선택 실험에서 출발해, Gooo 자료로 새로 학습한 2,072개 파라미터 모델과 Go 실행기로 이어졌습니다. 현재 공개물에는 FP32와 삼진 가중치, 학습 기록, 실제 코드 생성·실행 관측이 함께 들어 있습니다. 목표는 적은 메모리로 의도와 생성 경로를 연결하고, 완전성을 여러 항목으로 관찰하며 다음 작업을 이어갈 수 있게 하는 것입니다.

문장 앞의 도입 표현에 민감한 문제를 확인한 뒤, 전체 입력을 유지하는 네 가지 새 모델을 학습했습니다. 기존 개발 입력에서 FP32의 첫 경로 충족 수는 대조군 113/512, 문장 조각의 빈도를 쓰는 조건에서 368/512였습니다. 표현을 다양하게 학습한 조건과 삼진 양자화에서는 후퇴한 결과도 있었습니다. 새 결과는 아래 연구 부록에 있으며, 기존 모델의 코드 생성·실행 관측은 해당 실험별로 남아 있습니다.

새 모델을 arm64와 Linux에서 대조한 결과, 첫 선택은 같았고 후순위까지 포함한 272개 후보 순위에 차이가 있었습니다. 이후 반올림 시점을 명시한 규칙으로 147,456회 호출을 대조했고, 18,432개 입력 쌍의 계산값과 전체 순위가 모두 일치했습니다. 새 계약은 SDK v0.2.15로 공개했고, 로컬과 Linux에서 각각 18,432개 입력·36,864회 판단으로 연구 결과를 재현했습니다. 이어서 SDK v0.2.15를 연결한 컴파일러로 400회 생성과 800회 실행을 완료했고, 9,600개 기대값이 모두 일치했습니다. 실제 모델 판단은 816회였으며, 모델을 끈 경로도 결정론적으로 완료했습니다. 아래에 새 관측과 기존 V3 모델 사용 예시를 각각 연결했습니다.

Connecting model files to Gooo — 2026-10-04

The released Go SDK v0.2.21 uses bounded model reads and checks the opened regular-file identity and extent. On Unix, nonblocking/no-follow open prevents the observed replacement-FIFO writer wait. The SDK's existing non-symlink model rule is retained. Use hf download --local-dir to obtain regular metadata and adjacent weights. The exact compact FP32 pair at Hub revision 7c2781cbbcf81edba89a50d0423276c042b082de was downloaded and checked: metadata SHA256 6478334e1e181d0854865d1da54b9c2203eaa0e488f2b05617c92e27f08b755f, weights SHA256 43c503b224a78f23725101e58cdef1ec4647831329eea4de80f0a259a0253d31.

Compiler dev PR1206 passed its exact-head CI and independent source proof. Candidate 2420ad19 completed 128 model-path swaps with zero writer releases/timeouts; 34 regular completions kept 4,352/4,352 expectations. The standalone shared-model worker kept 512/512 in concurrent Korean/English requests and joined both processes. Original failures, counts and process CPU scope. Main promotion PR1207 passed its own canonical CI37167954453, independently verified source proof and live promotion authorization, then normally merged as 4f6c7566. Both CLI and worker are freshly installed from that clean main with Go1.27.1 / SDK v0.2.21. Fresh installed swaps completed 128 requests with zero writer releases/timeouts; 35 regular completions kept 4,480/4,480. Separate ordinary runs kept 1,024/1,024 and concurrent shared-model worker requests kept 512/512 with joined EOF. The older d1bfd273 / SDK v0.2.20 wait remains recorded as the baseline.

The research comparator now binds each recorded source inventory to its own immutable Git revision. Complete local/Linux replays each made 73,728 actual predictions over 18,432 paired inputs. Explicit arithmetic kept all values, rankings and finite curves; fourteen source changes stay visible. Legacy arithmetic keeps 272 ranking and 38 finite-outcome differences between platforms. Research CI37168510833 completed SUCCESS, and its arithmetic archive was independently downloaded, hash-checked and fully consumed. Source deltas and original differences.

A direct file-based example

With Go1.27.1, a compatible compiler and the compiler repository as working directory, the existing compound source, plan and expectations can run directly:

hf download asketeddy/gooo-shared-judgment-tiny-v1 \
  --revision 7c2781cbbcf81edba89a50d0423276c042b082de \
  --include 'research/compact-runtime-20261003/models/fp32/*' \
  --local-dir ./gooo-models

gooo body-path-run \
  --source examples/body-codegen/typed-path-compound.gooo.fixture \
  --activity Combined \
  --path-plan examples/body-codegen/typed-path-compound-plan.json \
  --cases examples/body-codegen/typed-path-runtime-cases.json \
  --model gooo-models/research/compact-runtime-20261003/models/fp32/model.json \
  --repeat 2 --timing --out compound-model-results

gooo body-path-run --verify-timing --out compound-model-results

Choose a fresh output path. Omitting --model uses deterministic construction. Read assembled Go in run-N-generated.go, current case outputs in run-N-runtime.json, and request status plus the finite numerator/denominator in summary.json. Setup failures and unobserved cases have their own recorded state. The clean candidate made two actual shared-model predictions and four native runs, keeping 6/6 expectations in this existing example. First/repeated responses were 884.413250/70.435292ms, covering construction and execution while excluding initial model preparation and saving/output. Weights and training remain unchanged. Fresh main 4f6c7566 also kept 6/6 in two constructions, two predictions and four native runs. Its original first/repeated responses are 1,188.190292/75.380833ms, and the saved-file timing verifier passed. The earlier candidate timing stays attached to its own run. Current installed concurrent worker CPU averaged 72.69% of one core during a 1.796446292s process interval; this includes setup, native children and EOF joining. Host CPU and model-only RAM remain unobserved. Commands, saved outputs and int64 readback and Korean input/error guide give the full scope.

What we want this to contribute

The 2026-10-03 research update narrows the next work to useful bounded language features: choosing informative execution inputs, preserving operation order, and reusing small typed assemblies. An additive Go SDK probe-ranking API explores the first step with zero model calls or training updates. Its compiler integration now has a 24-generation paired pilot: the frozen compact bag-original FP32 judge made 24 actual predictions across the model-enabled arms, and all configurations produced 48 compiled runs. A declared Gooo reference activity supplied one new observation before construction. The eight oracle-enabled generations matched 48/48 supplied expectations; controls remain in the publication, with 84/144 matched across all 24 generations. The original sparse selection case passed in every arm. This is one authored task with two language views and two repetitions, with zero weight updates. The Korean view needed extra search, and oracle-enabled generation added cost. Compiler PR 1160 tracks integration and deployment; the pilot pins compiler 2f02d244 and SDK v0.2.16. Model artifacts and the existing task scores below remain attached to their original observations.

An additional paired compiler experiment uses the same frozen judge with SDK v0.2.17 and compiler 87d8afae: 96 generations, 120 actual model predictions and 192 compiled executions. The compiler can retain candidate outputs and compare new reference observations against them. Probe evaluations fall from 33 to 14 per oracle request, and all 24 fresh/reuse pairs preserve generated source, choices and runtime values. Both oracle modes together meet 288/288 finite expectations; the full control comparison is 396/576. This remains one authored task with Korean/English views, six repetitions and zero training updates.

With the model enabled, observation-phase median falls from 0.603 to 0.333 ms, while complete codegen changes from 9.213 to 9.372 ms. Both oracle modes have 18.83 MiB median peak RSS. Process CPU changes from 92.35% to 93.02% of one core; whole-host utilization change is unmeasured. Reuse saves candidate evaluations; the end-to-end latency result is slightly slower in this collection. The compiler in that experiment still searches after a single candidate remains. PR 1162 tracks this opt-in integration. The existing model weights are preserved.

The next direct-projection experiment completes 120 generations and 240 compiled executions, with 120 actual model predictions in control arms. Complete observation can now select the single surviving candidate before model loading or search. Resolved arms make zero predictions, meet 144/144 finite expectations and match all 24 corresponding cached-search bodies and runtime results. All oracle arms meet 432/432; the full comparison including sparse-case controls is 540/720. Compiler 9158c3cd, PR 1164 adds this optional route with explicit source replay and skipped-work records.

For model-requested fresh processes, cached search vs direct projection has 9.491 vs 8.267 ms median codegen, 18.77 vs 17.29 MiB peak RSS, and 92.45% vs 88.01% CPU relative to one core. All samples are retained, including a slow initial control. This is one task, two language views and six repetitions with unchanged weights; host utilization change and broader speedup remain unmeasured. The model supplies a preference when choices remain, and a resolved source observation can complete this narrow construction without another prediction.

Smaller source-derived assembly recipes

The source-recipe pilot uses the unchanged compact/bag-original/fp32 full-input export in actual native construction. Gooo derives the typed base from its source body; a small recipe names the permitted structural choices. Compiler 3a52232d, PR 1168 adds that input form and merged to dev as 79e7e20c. Its main promotion is tracked in PR 1169.

Across 24 generations and 48 compiled runs, all 12 recipe/full-document pairs have equal expanded plans and emitted code. Compact authored JSON falls from 1,538 to 569 bytes (63.0%). Six real joint predictions each judge three choices; there are no training updates. Oracle/resolution arms satisfy 72/72 independent runtime expectations; sparse single-example controls satisfy 12/72, making the complete comparison 84/144. One authored subtraction task and repeated requests define this finite scope.

Own-model search has median whole-process generation time 8.143 ms for the full document and 8.598 ms for the recipe, including expansion. Peak RSS medians are 17.33 and 17.64 MiB; CPU relative to one core is 86.78% and 86.53%. Each cell has three samples. Whole-host utilization change and broader task accuracy are unmeasured. The pilot reduces authored representation while retaining the added processing cost. Initial collector accounting failure and all controls are public.

The language carries the plan, and the model supplies small local judgments.

The constant-body follow-up uses the same frozen model with SDK v0.2.18 and compiler b742ba75 in PR 1170. It repairs a language gap: bodies that never read their declared input can now use typed recipe construction. Six generations and twelve compiled runs met 36/36 finite expectations, including both integer endpoints, with three actual predictions and zero training updates. Three repeats per arm cover one authored body.

For that constant body, deterministic generation took a median 9.209 ms and 17.42 MiB peak RSS; model-connected generation took 9.731 ms and 19.31 MiB. Process CPU medians were 90.56% and 97.96% of one core; host utilization change was unmeasured. Isolated prediction median was 7.708 microseconds. Model ranking increased evaluated candidates from five to seven, so this task favors the deterministic route. The first proposed body failed in both arms; finite search reached the same successful source. All failed attempts and process records are retained. The result informs where a small model helps and where ordinary language machinery is already sufficient.

Condition chains using the same model

The condition-chain pilot extends source recipes to if ... else if ... else. Compiler development source bf51d65a lowers that form into existing typed nested branches, retaining condition order, local scope and source binding. Its branch is public; main deployment is tracked separately in the compiler wiki.

One authored clamp, one bilingual intention and three repeats per arm give six generations and twelve compiled runs. All 42 finite expectations match, including 24 selection-disjoint expectations and both int64 endpoints. The frozen own model makes three actual joint predictions; training updates are zero. Both arms choose the first candidate and emit equal source. Seven of eight declared combinations remain unattempted per request.

Deterministic/model generation medians are 7.958/9.227 ms, process peak RSS 17.45/17.91 MiB and one-core-relative CPU 90.64/91.58%. Prediction median is 7.917 microseconds. The first deterministic process takes 367.183 ms; all samples are public and its cause was not isolated. Cache state and whole-host CPU change were not instrumented. This already complete body favors deterministic assembly. The useful result is that a common source form now reaches both construction routes using the existing small runtime and unchanged weights.

An instruction-order feature preflight uses this frozen judge for 64 actual predictions. Four authored arithmetic pairs, two languages and four wrapper forms produce 32 pairs with identical V4 features and identical model predictions despite different requested operation orders. An experimental 128-byte directed-clause sketch distinguishes those 32 pairs; a repeated-clause counterexample still aliases and is retained. The feature kernel measures 372–619 ns/op with zero heap allocations on M4 / Go 1.27.1. It has no trained model head or new accuracy score, and this preflight includes zero optimizer updates or native Gooo executions. The current weights and input ABI remain attached to their original experiments.

The native order follow-up uses main compiler 729482ed with four operation pairs, Korean/English views and both requested orders. Across 96 generations, 192 compiled runs and 64 actual predictions, first-attempt functional completion is 8/16 requests for deterministic construction and 6/16 for each frozen positioned/bag model. With up to eight candidate attempts, all three arms complete 16/16 requests and 128/128 finite expectations each. All controls together meet 554/768 expectations. Training updates are zero, and the two models also have different learned weights.

The positioned model distinguishes all eight reversed-instruction pairs in its features and probability distributions, while its first masks remain unchanged. Bounded-search generation medians are 8.129 ms deterministic, 9.352 ms positioned and 8.861 ms bag; total candidate evaluations are 24, 53 and 46. The first deterministic process took 574.741 ms and is retained. These are one sample per authored request, with process costs and limits in the raw publication.

A source-order counterexample reverses the source assignments while keeping the instruction fixed. The two correct source-relative choices are opposite, but the local root-order model inputs are identical. Both calls choose reverse: native outcomes are 8/8 and 0/8. The shared local judge lacks a distinction needed for this pair. These observations leave the released weights unchanged and retain bounded deterministic TDD as the working completion path.

The source-order preflight now describes two source-derived assignment operations in 48 bytes per alternative. Compiler ba00850f exposes its validated plan through the explicit body-context --include-plan option in PR1174. Across 48 exports from four authored families, the new description distinguishes 24/24 source-reversal pairs that have identical existing V3 local root inputs. Renaming locals and commuting operands each preserve 16/16 description pairs. Korean/English intent changes preserve 24/24 source-description pairs; this last check measures the source channel, not language understanding.

Five M4/Go1.27.1 samples measure 28.02–29.31 ns/op and zero allocations for the descriptor alone. The complete context exports take a median 6.119 ms per fresh process; parsing, binding and validation remain separate costs. These observations add no model predictions, native runs or training updates. The new representation is isolated from released V3/V4 weights, and its trained-model quality is still unmeasured. It covers two simple integer updates; nested expressions and branches require further work. The initial collector digest-format failure is retained and covered by a regression. The next learning study must evaluate actual assembled bodies and additional attempts with the same deterministic-search budget.

Our aim is to make intent, construction and observed behavior travel together as a program evolves. We measure complete finite behavior, partial coverage, extra attempts, unresolved obligations, time and memory. These measurements help decide which language features and model changes to develop next.

For the full-input learning study, a first path is complete when it satisfies all 16 supplied examples for that input. A result of 368/512 therefore describes that specific authored task collection. Larger programs, broader types and reusable learned constructions are the next language questions.

Model probabilities describe which permitted path the judge favors. Finite coverage counts the supplied examples that pass. An unresolved observation stays UNKNOWN in its own receipt dimension. Keeping these meanings visible helps the system identify the next useful experiment.

Choose an entry point

What you want to do Compatible starting point
Generate with SDK v0.2.14 or later Pinned compact V3 bundle in the usage example below
Explore the new V3/V4 arithmetic contract in Go SDK v0.2.15 and explicit-arithmetic models
Generate with the new V4 contract Integration on main fc0e99c4 through merged PR 1159; native evidence
Follow the current language work Language guide, experiment repository and the dated appendices here

SDK release source 59c8d342da4475506b90954469aa201f85cadeb3 passed complete replay on darwin/arm64 and linux/amd64: each made 36,864 predictions over 18,432 frozen inputs, matching intermediate values, probabilities and full rankings. Both original SDK reports and Linux CI are public. This stage performed zero training updates. The native integration observation below follows that library replay.

New native generation and execution observation

On sixteen known bilingual tasks, four training arms × three precisions × two layouts plus a disconnected arm completed 400 generations and 800 compiled runs. All 9,600 supplied expectations and 192 representation pairs matched. An independent reader checked the original source/result bindings, 1,760 progress records and 432 feedback records. Total actual predictions were 816; the disconnected requests made zero model calls.

Bag-original FP32 completed the first path for 16/16 requests in this integration slice. Its broader development result stays 368/512. Compact median prediction time was 8.33 µs, whole codegen 10.00 ms and compiler peak RSS 17.78 MiB. Process CPU was 78.84% of one core; host-wide utilization change was unmeasured. The disconnected arm completed after 96 extra candidates, with 9.80 ms median codegen. This task size shows reduced search attempts and similar process latency.

Full native results, every variant and original records retain the timing conditions and known-task scope. All 400 receipts keep permission_boundary unresolved. The observation used compiler e461c1d and SDK v0.2.15, with zero training updates and no new default checkpoint. Compiler integration merged to dev in PR 1158 and main in PR 1159 above.

New full-input research exports — 2026-10-03

The full-input appendix contains four freshly initialized shared judges, each exported as FP32, PTQ and QAT in expanded and compact layouts. Complete source-derived input and Korean/English instructions are retained. The comparison changes position-based features versus global fragment counts, and original versus varied training wording.

FP32 arm Original first-path complete /512 Extra ranked attempts Both KO/EN valid /256
positioned-original 113 1,469 1
positioned-varied 266 528 75
bag-original 368 186 180
bag-varied 320 326 132

Bag-original FP32 reaches 358/512 and 368/512 on the two preregistered new wording forms. Bag-varied falls to 173/512 on the new suffix. Bag-original PTQ and QAT reach 94/512 and 287/512 on original inputs. Its global fragment representation also maps two distinct operation orders to equal features; the appendix retains that explicit counterexample. These are previously observed authored source tasks.

The Go audit reconciled all 6,400 MPS updates and made 25,824 real predictions across export parity, compact parity, calibration and three development forms. All twelve exports and complete raw evidence are public. Optimization took 42.18 seconds; average process CPU was 49.83% of one core and peak RSS was 1,809,678,336 bytes. CPU usage describes the optimizer process. Model inference and native generation have their separate measurement above.

V4 compact artifacts use triple_semantic_context_bag_v4_joint_v1 and the research Go runtime. The matching SDK v0.2.15 is published with complete numerical replay and the native integration observation above. Remaining split/form comparisons and new intentions are subsequent steps. Later historical sections retain the measurements of their pinned earlier models.

The subsequent Linux comparison retained the same first selected mask on all 18,432 development rows, with 272 different complete candidate orders. Small score deltas changed partial-completion curves, so the exact cross-platform comparison failed. Its complete observations are public. The explicit-arithmetic follow-up then made 73,728 predictions per platform with unchanged weight bytes and zero training updates. All 18,432 paired inputs now have identical hidden/logit/ probability bits and complete rankings under float32_separate_v1. Original first-path completeness and extra ranked attempts stay unchanged; partial coverage shifts in both directions compared with legacy arm64 arithmetic. Converted metadata and complete four-lane journals are public in that appendix. The SDK replay reproduces these observations. The native stage above subsequently completed generation and execution with this arithmetic contract.

How the model is used

The compiler derives context from an original Gooo body, its declared typed alternatives, and complete Korean/English intentions. The model ranks complete paths involving conditions, assignments, references, nested branches and arithmetic. Finite expectations evaluate a proposed body. Later predictions may include actual failed-case feedback and rank the remaining choices.

The model makes small decisions inside a supplied plan. The compiler supplies the source binding, allowed alternatives, attempt budget and deterministic continuation. Disconnected or unsupported inputs keep their complete text and continue with zero model predictions. Generated Go then has a separate replay, build and execution stage.

The current artifacts target authored Integer → Integer tasks. Broader source discovery, larger programs, reusable abstractions and Korean/English meaning alignment are active development areas.

Original V3 architecture, training and compatible files

The shared model reuses a 256 → 8 → 2 local network over three ordered decisions: 2,072 trainable parameters. The dense control has 18,656. Both were freshly initialized using the same Go initializer recipe and trained on Gooo-derived data. Each arm completed 800 FP32 and 800 QAT updates on local MPS, totaling 3,200 optimizer updates. PTQ exports come from the FP32 models. That total belongs to the original shared/dense comparison. The later four-arm full-input study above adds 6,400 updates and twelve exports in its own appendix.

The data contains 1,024 training, 256 calibration and 256 development program groups. The development set has 512 Korean/English views and had been observed in earlier studies. The source/teacher feature bank has 10,739 initial and feedback rows. Quality numbers below describe this known development cohort.

Files in this Hub repository Purpose Go loader
research/compact-runtime-20261003/models/{fp32,ptq_ternary,qat_ternary}/ Current compact shared representation jointdecision.LoadSharedThree
models/shared-local/{fp32,ptq_ternary,qat_ternary}/ Original expanded shared exports jointdecision.LoadThree
models/dense/{fp32,ptq_ternary,qat_ternary}/ Matched dense controls jointdecision.LoadThree
evidence.zip, go-audit.json, manifest.json Original training and native study evidence Go audit tools in the research repository
research/compact-runtime-20261003/ Exact representation parity and kernel measurements SDK v0.2.14-experimental
research/compact-native-20261003/ Actual paired Gooo generation and compiled execution Compiler and independent receipt reader
research/full-input-initial-20261003/ Four-arm study: full inputs, twelve exports and original observations Research runtime and SDK v0.2.15
research/full-input-platform-20261003/ Retained Linux replay and complete ranking diagnosis Offline Go diagnostic reader
research/full-input-separate-20261003/ Explicit arithmetic, paired observations and converted metadata Research runtime and SDK v0.2.15
research/full-input-native-20261003/ 400 generations, 800 compiled runs, 192 matched pairs and independent receipt consumption Compiler e461c1d with SDK v0.2.15
research/full-input-sdk-20261003/ Complete local/Linux SDK replay reports SDK release source 59c8d34

Compact files use gooo/tiny-shared-three-choice-path-model/v1 metadata and three tensors: 8×256 input weights, eight biases and 2×8 output weights. The total feature input is 768 values; a caller workspace holds 24 hidden values and eight mask scores. Go loads model.json and its adjacent weights.

With a compatible Gooo binary and an authored three-choice source/plan:

hf download asketeddy/gooo-shared-judgment-tiny-v1 \
  --revision 985999a89caba6a31cc7147f66ba29a5ce76a1d9 \
  --include 'research/compact-runtime-20261003/models/fp32/*' \
  --local-dir ./gooo-models

gooo body-codegen --json --activity ChoosePath \
  --path-plan plan.json \
  --path-model gooo-models/research/compact-runtime-20261003/models/fp32/model.json \
  --path-step-attempts 1 --path-feedback-rounds 7 source.gooo

source.gooo and plan.json above are caller inputs. Complete measured examples and expectations are in the native evidence archive. The compiler integration guide describes the input contract. Offline PyTorch performs training/export; the Go SDK performs local inference and path search.

Decision quality

Each development view has 16 finite expectations. This table ranks all eight masks from one initial prediction and evaluates their stored finite targets. Adaptive failure feedback is measured separately in actual codegen.

Export First choice complete /512 Initial cases matched /8192 Extra ranked attempts Complete by budget 4 /512 EN/KO mask disagreements /256 Same-mask both wrong /256
Dense FP32 95 2265 1572 289 256 0
Shared FP32 113 2592 1469 319 255 0
Dense PTQ 74 1967 1595 285 45 180
Shared PTQ 80 2052 1439 336 136 98
Dense QAT 82 2050 1584 293 240 12
Shared QAT 89 2191 1512 327 255 1

Shared FP32 improves first-choice completion 18.55% → 22.07%, finite-case matches 27.65% → 31.64%, and extra ranked attempts 1,572 → 1,469. Its development passing-set negative log likelihood worsens 975.663 → 992.402. Comparison/assignment, nested branches and successive assignments need more extra attempts. The full per-family counts remain in go-audit.json.

Paired-language disagreement remains 255/256 for shared FP32. Agreement also needs a correctness check: dense PTQ gives the same wrong mask on 180 pairs. The subsequent full-input study above measures both language consistency and membership in the valid-path set. Full-budget success here depends on the authored finite search space containing a passing alternative.

Input sensitivity diagnosis

A subsequent fixed-weight study varies five authored instruction forms while keeping the original Gooo source, eight paths and finite expectations constant. The collector made 92,160 actual Go predictions across all 3,072 corpus views and six exports. An independent reader replayed all 92,160 predictions with zero numerical difference locally and recomputed every condition's finite and bilingual summary. These replay calls are recorded separately from collection.

The Linux CI replay also reproduced all chosen masks and 30 condition summaries. Its largest floating-point difference was 0.000001430511474609375, within the declared 0.00001 tolerance. The run's 13 jobs completed successfully.

For shared FP32, the 512 development views give:

Complete input form First-choice complete /512 Extra ranked attempts EN/KO disagreements /256
Original development prefix 113 1,469 255
Bare authored instruction 480 38 32
Calibration prefix 92 1,372 250
Development phrase as suffix 233 662 180

The bare form matches the format of the original training instructions. Its result describes a controlled intervention on an already observed cohort; original-input quality remains 113/512. Production Gooo keeps the complete caller text. The completed initial full-input comparison above varies training phrasing and positional features while preserving full input and source binding.

The source v3 intent encoder uses four relative-position buckets for byte fragments. Added prefixes alter the fragments, their buckets and normalization. The study identifies sensitivity to this combined change; further experiments will separate those factors. All six models, all forms and regressions are in the report and the Hub appendix research/bilingual-wrapper-20261003/. Its 30-member archive includes every full constructed input and prediction, the exact source dataset, six small existing models and the independent reader. The appendix adds zero training updates. The model weights at their existing paths remain unchanged.

Actual Gooo generation and execution

The latest compact study used 16 bilingual views, three model variants and two representations, for 96 generations and 298 real model predictions. Of those predictions, 202 used finite-failure feedback. Each generation was immediately followed by a native build and two executions: 192 compiled runs.

All 2,304 supplied finite expectations passed, producing 4,608 ordered values. Of these expectations, 768 use inputs absent from the current selection suite; their relationship to model training remains part of the dataset scope. All 48 unseeded expanded/compact pairs retain matching generated source, search/feedback semantics and ordered outputs.

An independent compiler reader consumed all 96 runtime receipts, 596 progress records and 202 feedback records. The first unresolved completeness dimension is permission_boundary throughout. Gooo keeps each observation axis and its remaining questions alongside the finite functional results.

This compact study uses compiler 7db19b6bc9a2909aa059f1265d39a539c3573a57, SDK v0.2.14-experimental and Go 1.27.1 on local darwin/arm64. The earlier architecture study separately made 320 predictions in 96 generations. Earlier quality/native report and compact native report retain both experiments with their own identities.

Storage, speed and CPU observations

Compact variant Weight bytes Resident tensor bytes Native-study prediction median Fresh-process codegen median
FP32 8,288 8,288 23.709 µs 28.927 ms
PTQ ternary 446 2,096 + 8 scale bytes 25.000 µs 27.765 ms
QAT ternary 446 2,096 + 8 scale bytes 30.208 µs 29.511 ms

Caller workspace is 3,200 bytes. Five trits fit in one byte, giving 1.6 stored matrix bits per weight; biases and scales have their own storage. The runtime decodes matrix values into int8 arrays and computes with int8/float32. Valid warmed kernel calls allocate zero heap objects in the measured contract.

The separate kernel audit checks all 10,739 frozen states across three variants: 64,434 actual predictions, with bit-for-bit agreement in features, hidden values, logits, probabilities and chosen mask. Compact conversion preserves the existing shared model; this step performs zero optimizer updates.

Native timings have 16 generations per representation/variant. Prediction timers exclude loading, startup, code generation and builds. QAT's complete codegen median increased from 29.237 to 29.511 ms after compact conversion. Cold-cache outliers remain recorded, including the first 6.45-second native build.

Codegen process CPU/wall-time medians are 83.7–84.6% of one core. Generated program RSS medians are 4.20–4.23 MiB, covering the compiled arithmetic child. Inference-process RSS and whole-host CPU increase were unmeasured in this study. Retained serving and additional hardware need their own controlled measurements.

Reproducibility and development

Every public evidence bundle has a closed byte inventory. The compact native archive has 684 members, 3,040,850 compressed bytes and 12,474,521 decoded bytes. Original failed preflights, negative variants, complete inputs and measurement outliers are retained. Public packaging checks private paths and credential patterns, including decoded embedded parent receipts.

Training source: 0f3249596b4dacdb34241e3b5e8b4cafac58b0df. Original independent Go audit: 8398f1994d65b75af0483db3fefd637eec99ced4. Compact native collector: 7e9be4489d15663cbac8f3e71bc02d2e4649dfe5. Immutable compact model edition: 985999a89caba6a31cc7147f66ba29a5ce76a1d9. Immutable native appendix edition: 37db3a1f8a06669081a370ee1e7b23a64cbc00fb.

Next work connects the published SDK to native generation and execution while continuing full-input evaluations and operation-order representation work. Korean/English intent alignment, more expressive Gooo construction, per-axis before/after observations and serving costs guide the wider development. Repeated successful structures may later become reusable language abstractions. The dated research protocols define each experiment's budgets and acceptance criteria; the compiler retains deterministic continuation between model versions.

Research acknowledgments

  • Solar-Lezama et al., SKETCH (2006): partial programs completed under a specification inform our view of bounded construction. Combinatorial Sketching for Finite Programs.
  • Ellis et al., DreamCoder (2020/2021): neural program search and learned reusable abstractions inform the longer-term language/model direction. Paper.
  • Balog et al., DeepCoder (2016/2017): learned program properties guide synthesis search. It informs our question of how much a small judgment can reduce construction attempts. Paper.
  • Laya / ConvAI Innovations: its structured decision interface was used in our early Gooo experiments. The present weights use fresh initialization and Gooo-specific training. Project.
  • Ma et al., BitNet b1.58 (2024): ternary-weight research motivates exploring low-storage models with explicit quality and runtime measurements. Paper.
  • W3C PROV-O (2013): entity, activity and agent relations inform how we trace inputs, construction and observed results. Recommendation.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for asketeddy/gooo-shared-judgment-tiny-v1