aharrar commited on
Commit
8fa0427
Β·
verified Β·
1 Parent(s): fe4fe50

Add Runtime Operators graph research card

Browse files
Files changed (1) hide show
  1. README.md +239 -0
README.md ADDED
@@ -0,0 +1,239 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: neosyntropy
3
+ tags:
4
+ - neosyntropy
5
+ - state-machine
6
+ - software-engineering
7
+ - agentic-workflow
8
+ - tool-use
9
+ - structured-output
10
+ - reasoning
11
+ - lora
12
+ - swe-bench
13
+ ---
14
+
15
+ # NeoSyntropy Runtime Operators
16
+
17
+ ## Overfit the graph, not the benchmark.
18
+
19
+ Language models are good at proposing actions. Production software must decide
20
+ which actions are legal, what evidence is required, and when a task is complete.
21
+
22
+ NeoSyntropy Runtime Operators is an experimental model family for testing one
23
+ idea: small, specialized models can become reliable software-engineering workers
24
+ when they are trained for narrow state-machine operators and executed inside an
25
+ application-owned graph.
26
+
27
+ This repository is currently a research card and evaluation specification. It
28
+ does **not** contain trained weights or make benchmark-performance claims yet.
29
+
30
+ ## The model family
31
+
32
+ The graph composes several model roles instead of asking one unconstrained model
33
+ to own the entire workflow.
34
+
35
+ | Model | Responsibility |
36
+ | --- | --- |
37
+ | [Runtime Structure](https://huggingface.co/Neosyntropy/runtime-structure) | Convert observations into application-owned schemas |
38
+ | [Runtime Route](https://huggingface.co/Neosyntropy/runtime-route) | Propose a legal next node from declared candidates |
39
+ | [Runtime Deterministic Reasoning](https://huggingface.co/Neosyntropy/runtime-deterministic-reasoning) | Perform a bounded reasoning step and declare tool calls |
40
+ | [Runtime Stochastic Reasoning](https://huggingface.co/Neosyntropy/runtime-stochastic-reasoning) | Generate alternative plans or repairs inside an allowed search space |
41
+ | [Runtime Guard](https://huggingface.co/Neosyntropy/runtime-guard) | Validate claims against rules, tests, and supplied evidence |
42
+ | [Runtime Score](https://huggingface.co/Neosyntropy/runtime-score) | Produce rubric-grounded measurements for evaluation and selection |
43
+
44
+ The planned unified operator adapter conditions these roles with explicit tokens
45
+ such as `<OPERATOR:UNDERSTAND>`, `<OPERATOR:PROPOSE>`, and
46
+ `<OPERATOR:REPAIR>`. NeoSyntropy remains responsible for execution, transition
47
+ legality, state commits, and side effects.
48
+
49
+ ## The operator graph
50
+
51
+ ```text
52
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
53
+ β”‚ RETRIEVE │◄──────────────┐
54
+ β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
55
+ β”‚ evidence β”‚ need information
56
+ β–Ό β”‚
57
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” requirements β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” candidates β”Œβ”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
58
+ β”‚ UNDERSTAND │───────────────►│ DECOMPOSE │─────────────►│ PROPOSE β”‚
59
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
60
+ β”‚ selected plan
61
+ β–Ό
62
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” complete β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” evidence β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
63
+ β”‚ SUCCESS │◄────────────│ VERIFY │◄───────────────│ OBSERVE β”‚
64
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β–²β”€β”€β”€β”€β”€β”€β”˜
65
+ β”‚ failed test β”‚ result
66
+ β–Ό β”‚
67
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” corrected plan β”Œβ”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”
68
+ β”‚ REPAIR │─────────────────►│ EXECUTE β”‚
69
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
70
+ ```
71
+
72
+ The same graph expressed as NeoSyntropy code:
73
+
74
+ ```python
75
+ from neosyntropy import END, FSM, edge_deterministic
76
+
77
+ def signal_is(expected):
78
+ return lambda state: state.get("signal") == expected
79
+
80
+ def verified_complete(state):
81
+ verification = state.get("verification", {})
82
+ return state.get("signal") == "COMPLETE" and verification.get("passed") is True
83
+
84
+ graph = FSM(
85
+ entry=understand,
86
+ nodes=[
87
+ understand,
88
+ decompose,
89
+ retrieve,
90
+ propose,
91
+ execute,
92
+ observe,
93
+ verify,
94
+ repair,
95
+ compress_memory,
96
+ success,
97
+ ],
98
+ edges=[
99
+ edge_deterministic("Understand", "Decompose", guard=signal_is("CONTINUE")),
100
+ edge_deterministic("Understand", "Retrieve", guard=signal_is("NEED_INFORMATION")),
101
+ edge_deterministic("Decompose", "Propose", guard=signal_is("CONTINUE")),
102
+ edge_deterministic("Decompose", "Retrieve", guard=signal_is("NEED_INFORMATION")),
103
+ edge_deterministic("Retrieve", "Propose", guard=signal_is("CONTINUE")),
104
+ edge_deterministic("Propose", "Execute", guard=signal_is("CONTINUE")),
105
+ edge_deterministic("Execute", "Observe", guard=signal_is("EXECUTED")),
106
+ edge_deterministic("Observe", "Verify", guard=signal_is("CONTINUE")),
107
+ edge_deterministic("Observe", "Compress", guard=signal_is("MEMORY_PRESSURE")),
108
+ edge_deterministic("Verify", "Success", guard=verified_complete),
109
+ edge_deterministic("Verify", "Repair", guard=signal_is("FAILED_TEST")),
110
+ edge_deterministic("Verify", "Retrieve", guard=signal_is("NEED_INFORMATION")),
111
+ edge_deterministic("Verify", "Decompose", guard=signal_is("WRONG_PLAN")),
112
+ edge_deterministic("Repair", "Execute", guard=signal_is("CONTINUE")),
113
+ edge_deterministic("Repair", "Retrieve", guard=signal_is("NEED_INFORMATION")),
114
+ edge_deterministic("Compress", "Decompose", guard=signal_is("CONTINUE")),
115
+ edge_deterministic("Success", END),
116
+ ],
117
+ )
118
+ ```
119
+
120
+ `Edge` guards are authoritative. A model may propose a route, but it cannot
121
+ commit an undeclared transition or bypass a failed verification gate.
122
+
123
+ ## One shared state
124
+
125
+ Operators do not own isolated hidden memories. They read a projection of one
126
+ auditable workflow state and return a validated state patch.
127
+
128
+ ```json
129
+ {
130
+ "goal": "Fix the reported repository issue",
131
+ "requirements": [],
132
+ "task_tree": [],
133
+ "current_plan": "",
134
+ "retrieved": [],
135
+ "observations": [],
136
+ "errors": [],
137
+ "evidence": [],
138
+ "history": [],
139
+ "scratch": {},
140
+ "signal": "CONTINUE"
141
+ }
142
+ ```
143
+
144
+ Each operator receives only the fields it needs:
145
+
146
+ | Operator | Reads | Writes |
147
+ | --- | --- | --- |
148
+ | `UNDERSTAND` | goal, initial context | requirements, unknowns, signal |
149
+ | `DECOMPOSE` | goal, requirements | task tree, signal |
150
+ | `RETRIEVE` | goal, errors, repository context | retrieved evidence, signal |
151
+ | `PROPOSE` | requirements, task tree, evidence | candidate plan, signal |
152
+ | `EXECUTE` | selected plan | trusted execution result |
153
+ | `OBSERVE` | execution result | observations, errors, signal |
154
+ | `VERIFY` | requirements, observations, tests | evidence, verification result, signal |
155
+ | `REPAIR` | current plan, errors, evidence | corrected plan, signal |
156
+ | `COMPRESS` | accumulated state | compact state patch, signal |
157
+ | `SUCCESS` | verified evidence | terminal outcome |
158
+
159
+ `EXECUTE`, deterministic observation parsing, transition checks, and final
160
+ success gates should remain trusted runtime operations. Learned models propose;
161
+ the graph validates and commits.
162
+
163
+ ## What β€œoverfit the graph” means
164
+
165
+ The phrase describes deliberate specialization to a stable execution protocol:
166
+
167
+ - fixed operator vocabulary;
168
+ - explicit input projections and output schemas;
169
+ - declared tools and legal transitions;
170
+ - repair loops driven by real test evidence;
171
+ - consistent prompts across teacher generation, training, and inference.
172
+
173
+ It does **not** mean training on benchmark test patches, hidden tests, or expected
174
+ answers. Benchmark instances and repositories used for final evaluation must be
175
+ kept out of training and teacher-label generation. The hypothesis is that a model
176
+ can learn the reusable procedure while still generalizing to unseen issues.
177
+
178
+ ## SWE evaluation plan
179
+
180
+ The first experiment will compare the same base model and tool environment under
181
+ four scaffolds:
182
+
183
+ 1. **Direct agent** β€” one unconstrained model call loop.
184
+ 2. **Prompted operators** β€” operator prompts without fine-tuning.
185
+ 3. **Trained operators** β€” the operator adapter without graph enforcement.
186
+ 4. **NeoSyntropy graph** β€” trained operators with schemas, guards, state, and
187
+ verified transitions.
188
+
189
+ Evaluation will begin with the official
190
+ [SWE-bench](https://www.swebench.com/) harness for reproducibility, while newer
191
+ or contamination-resistant suites should be used for primary generalization
192
+ claims. Static benchmark scores will always be reported with the exact harness,
193
+ model, scaffold, context limit, tool budget, retry budget, and task exclusions.
194
+
195
+ Primary metrics:
196
+
197
+ - resolved instances (`pass@1`);
198
+ - legal-transition rate;
199
+ - schema-valid output rate;
200
+ - tool-call success rate;
201
+ - verification precision and false-success rate;
202
+ - repair-loop recovery rate;
203
+ - model tokens, wall-clock time, tool calls, and cost per resolved task;
204
+ - average state transitions and repeated-transition rate;
205
+ - outcome consistency across repeated seeds.
206
+
207
+ The most important comparison is not only whether a task was solved, but whether
208
+ the graph can reach the same or better result with smaller models, fewer wasted
209
+ actions, and no unverified success transition.
210
+
211
+ ## Data-generation protocol
212
+
213
+ Training traces are generated from complete graph executions rather than isolated
214
+ first-step prompts:
215
+
216
+ 1. Sample an unseen repository task and construct the initial state.
217
+ 2. Use a stronger teacher to label the current operator only.
218
+ 3. Validate the output schema and reject invented tools or evidence.
219
+ 4. Execute approved actions in the benchmark environment.
220
+ 5. Record observations and test results as new evidence.
221
+ 6. Route through the declared graph and continue until success or budget expiry.
222
+ 7. Store each operator transition with its state projection, output, signal,
223
+ provenance, and final task outcome.
224
+ 8. Freeze repository-level train, validation, and test splits before fine-tuning.
225
+
226
+ This produces training examples for the decision that was actually available at
227
+ each state, including failed attempts and evidence-grounded repairs.
228
+
229
+ ## Status
230
+
231
+ - Graph contract: specified.
232
+ - Role-model cards: published separately.
233
+ - Unified operator dataset: planned.
234
+ - Unified operator adapter: not trained yet.
235
+ - SWE baseline runs: not run yet.
236
+ - Benchmark claims: none.
237
+
238
+ Results, weights, datasets, and exact run manifests will be published only after
239
+ reproducible evaluation.