|
Download README.md from Neosyntropy/runtime-operators: direct link, hf CLI and curl.
- Browser
- Download file 18.9 kB
-
https://huggingface.co/Neosyntropy/runtime-operators/resolve/main/README.md
- Command line
-
hf download hf://Neosyntropy/runtime-operators/README.md
-
curl -L -o README.md https://huggingface.co/Neosyntropy/runtime-operators/resolve/main/README.md
18.9 kB
| library_name: neosyntropy | |
| tags: | |
| - neosyntropy | |
| - state-machine | |
| - software-engineering | |
| - agentic-workflow | |
| - tool-use | |
| - structured-output | |
| - reasoning | |
| - lora | |
| - swe-bench | |
| # NeoSyntropy Runtime Operators | |
| ## Overfit the graph, not the benchmark. | |
| Language models are good at proposing actions. Production software must decide | |
| which actions are legal, what evidence is required, and when a task is complete. | |
| NeoSyntropy Runtime Operators is an experimental model family for testing one | |
| idea: small, specialized models can become reliable software-engineering workers | |
| when they are trained for narrow state-machine operators and executed inside an | |
| application-owned graph. | |
| This repository is currently a research card and evaluation specification. It | |
| does **not** contain trained weights or make benchmark-performance claims yet. | |
| ## The model family | |
| The graph composes several model roles instead of asking one unconstrained model | |
| to own the entire workflow. | |
| | Model | Responsibility | | |
| | --- | --- | | |
| | [Runtime Structure](https://huggingface.co/Neosyntropy/runtime-structure) | Convert observations into application-owned schemas | | |
| | [Runtime Route](https://huggingface.co/Neosyntropy/runtime-route) | Propose a legal next node from declared candidates | | |
| | [Runtime Deterministic Reasoning](https://huggingface.co/Neosyntropy/runtime-deterministic-reasoning) | Perform a bounded reasoning step and declare tool calls | | |
| | [Runtime Stochastic Reasoning](https://huggingface.co/Neosyntropy/runtime-stochastic-reasoning) | Generate alternative plans or repairs inside an allowed search space | | |
| | [Runtime Guard](https://huggingface.co/Neosyntropy/runtime-guard) | Validate claims against rules, tests, and supplied evidence | | |
| | [Runtime Score](https://huggingface.co/Neosyntropy/runtime-score) | Produce rubric-grounded measurements for evaluation and selection | | |
| The planned unified operator adapter conditions these roles with explicit tokens | |
| such as `<OPERATOR:UNDERSTAND>`, `<OPERATOR:PROPOSE>`, and | |
| `<OPERATOR:REPAIR>`. NeoSyntropy remains responsible for execution, transition | |
| legality, state commits, and side effects. | |
| ## The operator graph | |
| ```text | |
| ┌───────────────┐ | |
| │ RETRIEVE │◄──────────────┐ | |
| └───────┬───────┘ │ | |
| │ evidence │ need information | |
| ▼ │ | |
| ┌────────────┐ requirements ┌─────────────┐ candidates ┌┴───────────┐ | |
| │ UNDERSTAND │───────────────►│ DECOMPOSE │─────────────►│ PROPOSE │ | |
| └────────────┘ └─────────────┘ └─────┬──────┘ | |
| │ selected plan | |
| ▼ | |
| ┌────────────┐ complete ┌─────────────┐ evidence ┌─────────────┐ | |
| │ SUCCESS │◄────────────│ VERIFY │◄───────────────│ OBSERVE │ | |
| └────────────┘ └──────┬──────┘ └──────▲──────┘ | |
| │ failed test │ result | |
| ▼ │ | |
| ┌─────────────┐ corrected plan ┌─────┴───────┐ | |
| │ REPAIR │─────────────────►│ EXECUTE │ | |
| └─────────────┘ └─────────────┘ | |
| ``` | |
| ### Complete node declarations | |
| The following definitions make every symbol in the graph explicit. The three | |
| tool names are application integrations: `search_repository`, `apply_plan`, and | |
| `run_tests`. Learned operators use schema-constrained model calls; execution and | |
| verification stay in trusted Python handlers. | |
| ```python | |
| from typing import Any, Literal | |
| from pydantic import BaseModel, ConfigDict | |
| from neosyntropy import NodeContext, OpenInput, SchemaNode, TextOutput, node | |
| Signal = Literal[ | |
| "CONTINUE", | |
| "NEED_INFORMATION", | |
| "EXECUTED", | |
| "FAILED_TEST", | |
| "WRONG_PLAN", | |
| "MEMORY_PRESSURE", | |
| "COMPLETE", | |
| ] | |
| class StrictModel(BaseModel): | |
| model_config = ConfigDict(extra="forbid") | |
| class TaskInput(StrictModel): | |
| repository: str | |
| issue: str | |
| class UnderstandOutput(StrictModel): | |
| requirements: list[str] | |
| unknowns: list[str] | |
| signal: Literal["CONTINUE", "NEED_INFORMATION"] | |
| class DecomposeOutput(StrictModel): | |
| task_tree: list[str] | |
| search_queries: list[str] | |
| signal: Literal["CONTINUE", "NEED_INFORMATION"] | |
| class RetrieveOutput(StrictModel): | |
| retrieved: list[str] | |
| signal: Literal["CONTINUE", "NEED_INFORMATION"] | |
| class PlanOutput(StrictModel): | |
| current_plan: str | |
| signal: Literal["CONTINUE", "NEED_INFORMATION"] | |
| class ExecuteOutput(StrictModel): | |
| executed: bool | |
| summary: str | |
| signal: Literal["EXECUTED"] | |
| class ObserveOutput(StrictModel): | |
| observations: list[str] | |
| errors: list[str] | |
| signal: Literal["CONTINUE", "MEMORY_PRESSURE"] | |
| class Verification(StrictModel): | |
| passed: bool | |
| summary: str | |
| class VerifyOutput(StrictModel): | |
| verification: Verification | |
| evidence: list[str] | |
| signal: Literal[ | |
| "COMPLETE", "FAILED_TEST", "NEED_INFORMATION", "WRONG_PLAN" | |
| ] | |
| class CompressOutput(StrictModel): | |
| observations: list[str] | |
| errors: list[str] | |
| history: list[str] | |
| signal: Literal["CONTINUE"] | |
| class SuccessOutput(StrictModel): | |
| outcome: Literal["satisfies_spec"] | |
| # Learned operator: infer explicit requirements without inventing facts. | |
| understand = SchemaNode( | |
| id="Understand", | |
| input_schema=TaskInput, | |
| output_schema=UnderstandOutput, | |
| prompt=( | |
| "<OPERATOR:UNDERSTAND> Read the repository issue and current state. " | |
| "Return explicit requirements, unresolved unknowns, and a legal signal." | |
| ), | |
| metadata={"operator": "UNDERSTAND", "model_role": "runtime-structure"}, | |
| ) | |
| # Learned operator: turn requirements into an ordered task tree and searches. | |
| decompose = SchemaNode( | |
| id="Decompose", | |
| input_schema=OpenInput, | |
| output_schema=DecomposeOutput, | |
| prompt=( | |
| "<OPERATOR:DECOMPOSE> Decompose the goal into ordered, testable subgoals. " | |
| "Generate only repository searches needed by those subgoals." | |
| ), | |
| prerequisites=("Understand",), | |
| metadata={"operator": "DECOMPOSE", "model_role": "runtime-deterministic-reasoning"}, | |
| ) | |
| # Trusted tool operator: execute only repository searches already in state. | |
| @node( | |
| id="Retrieve", | |
| input_schema=OpenInput, | |
| output_schema=RetrieveOutput, | |
| tools=("search_repository",), | |
| ) | |
| def retrieve(ctx: NodeContext): | |
| facts: list[str] = [] | |
| for query in ctx.state.get("search_queries", []): | |
| result = ctx.tools.invoke("search_repository", {"query": query}) | |
| facts.extend(result if isinstance(result, list) else [str(result)]) | |
| signal = "CONTINUE" if facts else "NEED_INFORMATION" | |
| output = {"retrieved": facts, "signal": signal} | |
| return ctx.result(output=output, state_updates=output) | |
| # Learned operator: propose one bounded implementation plan from evidence. | |
| propose = SchemaNode( | |
| id="Propose", | |
| input_schema=OpenInput, | |
| output_schema=PlanOutput, | |
| prompt=( | |
| "<OPERATOR:PROPOSE> Use requirements, task_tree, and retrieved evidence. " | |
| "Return one implementation plan grounded in repository symbols." | |
| ), | |
| metadata={"operator": "PROPOSE", "model_role": "runtime-stochastic-reasoning"}, | |
| ) | |
| # Trusted tool operator: apply the selected plan and run the real test harness. | |
| @node( | |
| id="Execute", | |
| input_schema=OpenInput, | |
| output_schema=ExecuteOutput, | |
| tools=("apply_plan", "run_tests"), | |
| ) | |
| def execute(ctx: NodeContext): | |
| plan = ctx.state["current_plan"] | |
| ctx.tools.invoke("apply_plan", {"plan": plan}) | |
| last_run = ctx.tools.invoke("run_tests", {}) | |
| output = { | |
| "executed": True, | |
| "summary": str(last_run), | |
| "signal": "EXECUTED", | |
| } | |
| return ctx.result(output=output, state_updates={**output, "last_run": last_run}) | |
| # Trusted operator: normalize tool output into observations and errors. | |
| @node(id="Observe", input_schema=OpenInput, output_schema=ObserveOutput) | |
| def observe(ctx: NodeContext): | |
| run = ctx.state.get("last_run", {}) | |
| errors = list(run.get("errors", [])) if isinstance(run, dict) else [] | |
| observations = [str(run)] | |
| signal = "MEMORY_PRESSURE" if len(observations) > 20 else "CONTINUE" | |
| output = {"observations": observations, "errors": errors, "signal": signal} | |
| return ctx.result(output=output, state_updates=output) | |
| # Trusted gate: only real test evidence may produce COMPLETE. | |
| @node(id="Verify", input_schema=OpenInput, output_schema=VerifyOutput) | |
| def verify(ctx: NodeContext): | |
| run = ctx.state.get("last_run", {}) | |
| passed = bool(isinstance(run, dict) and run.get("passed") is True) | |
| summary = "All required tests passed." if passed else "Required tests failed." | |
| output = { | |
| "verification": {"passed": passed, "summary": summary}, | |
| "evidence": [str(run)], | |
| "signal": "COMPLETE" if passed else "FAILED_TEST", | |
| } | |
| return ctx.result(output=output, state_updates=output) | |
| # Learned operator: repair the plan using concrete failure evidence. | |
| repair = SchemaNode( | |
| id="Repair", | |
| input_schema=OpenInput, | |
| output_schema=PlanOutput, | |
| prompt=( | |
| "<OPERATOR:REPAIR> Read current_plan, errors, and verification evidence. " | |
| "Return a corrected plan that addresses the observed failure only." | |
| ), | |
| metadata={"operator": "REPAIR", "model_role": "runtime-stochastic-reasoning"}, | |
| ) | |
| # Trusted operator: bound state growth without rewriting facts. | |
| @node(id="Compress", input_schema=OpenInput, output_schema=CompressOutput) | |
| def compress_memory(ctx: NodeContext): | |
| output = { | |
| "observations": list(ctx.state.get("observations", []))[-4:], | |
| "errors": list(ctx.state.get("errors", []))[-4:], | |
| "history": list(ctx.state.get("history", []))[-12:], | |
| "signal": "CONTINUE", | |
| } | |
| return ctx.result(output=output, state_updates=output) | |
| # Terminal node: reachable only through the verified_complete graph guard. | |
| @node(id="Success", input_schema=OpenInput, output_schema=SuccessOutput) | |
| def success(ctx: NodeContext): | |
| output = {"outcome": "satisfies_spec"} | |
| return ctx.result(output=output, state_updates=output) | |
| @node( | |
| id="OperatorFallback", | |
| input_schema=OpenInput, | |
| output_schema=TextOutput, | |
| is_fallback=True, | |
| ) | |
| def operator_fallback(ctx: NodeContext): | |
| return ctx.result(output={"message": "No legal operator transition."}) | |
| ``` | |
| Schema-node outputs are validated before their fields are proposed as state | |
| updates. Python handlers explicitly return `state_updates`; mutating `ctx.state` | |
| does not commit anything. | |
| ### Graph wiring | |
| The same nodes are connected as a NeoSyntropy FSM: | |
| ```python | |
| from neosyntropy import END, FSM, edge_deterministic, edge_fallback | |
| def signal_is(expected): | |
| return lambda state: state.get("signal") == expected | |
| def verified_complete(state): | |
| verification = state.get("verification", {}) | |
| return state.get("signal") == "COMPLETE" and verification.get("passed") is True | |
| graph = FSM( | |
| entry=understand, | |
| nodes=[ | |
| understand, | |
| decompose, | |
| retrieve, | |
| propose, | |
| execute, | |
| observe, | |
| verify, | |
| repair, | |
| compress_memory, | |
| success, | |
| operator_fallback, | |
| ], | |
| edges=[ | |
| edge_deterministic("Understand", "Decompose", guard=signal_is("CONTINUE")), | |
| edge_deterministic("Understand", "Retrieve", guard=signal_is("NEED_INFORMATION")), | |
| edge_deterministic("Decompose", "Propose", guard=signal_is("CONTINUE")), | |
| edge_deterministic("Decompose", "Retrieve", guard=signal_is("NEED_INFORMATION")), | |
| edge_deterministic("Retrieve", "Propose", guard=signal_is("CONTINUE")), | |
| edge_deterministic("Propose", "Execute", guard=signal_is("CONTINUE")), | |
| edge_deterministic("Execute", "Observe", guard=signal_is("EXECUTED")), | |
| edge_deterministic("Observe", "Verify", guard=signal_is("CONTINUE")), | |
| edge_deterministic("Observe", "Compress", guard=signal_is("MEMORY_PRESSURE")), | |
| edge_deterministic("Verify", "Success", guard=verified_complete), | |
| edge_deterministic("Verify", "Repair", guard=signal_is("FAILED_TEST")), | |
| edge_deterministic("Verify", "Retrieve", guard=signal_is("NEED_INFORMATION")), | |
| edge_deterministic("Verify", "Decompose", guard=signal_is("WRONG_PLAN")), | |
| edge_deterministic("Repair", "Execute", guard=signal_is("CONTINUE")), | |
| edge_deterministic("Repair", "Retrieve", guard=signal_is("NEED_INFORMATION")), | |
| edge_deterministic("Compress", "Decompose", guard=signal_is("CONTINUE")), | |
| edge_deterministic("Success", END), | |
| edge_fallback("Understand", "OperatorFallback"), | |
| edge_fallback("Decompose", "OperatorFallback"), | |
| edge_fallback("Retrieve", "OperatorFallback"), | |
| edge_fallback("Propose", "OperatorFallback"), | |
| edge_fallback("Execute", "OperatorFallback"), | |
| edge_fallback("Observe", "OperatorFallback"), | |
| edge_fallback("Verify", "OperatorFallback"), | |
| edge_fallback("Repair", "OperatorFallback"), | |
| edge_fallback("Compress", "OperatorFallback"), | |
| ], | |
| ) | |
| ``` | |
| `Edge` guards are authoritative. A model may propose a route, but it cannot | |
| commit an undeclared transition or bypass a failed verification gate. | |
| ## One shared state | |
| Operators do not own isolated hidden memories. They read a projection of one | |
| auditable workflow state and return a validated state patch. | |
| ```json | |
| { | |
| "goal": "Fix the reported repository issue", | |
| "requirements": [], | |
| "task_tree": [], | |
| "current_plan": "", | |
| "retrieved": [], | |
| "observations": [], | |
| "errors": [], | |
| "evidence": [], | |
| "history": [], | |
| "scratch": {}, | |
| "signal": "CONTINUE" | |
| } | |
| ``` | |
| Each operator receives only the fields it needs: | |
| | Operator | Reads | Writes | | |
| | --- | --- | --- | | |
| | `UNDERSTAND` | goal, initial context | requirements, unknowns, signal | | |
| | `DECOMPOSE` | goal, requirements | task tree, signal | | |
| | `RETRIEVE` | goal, errors, repository context | retrieved evidence, signal | | |
| | `PROPOSE` | requirements, task tree, evidence | candidate plan, signal | | |
| | `EXECUTE` | selected plan | trusted execution result | | |
| | `OBSERVE` | execution result | observations, errors, signal | | |
| | `VERIFY` | requirements, observations, tests | evidence, verification result, signal | | |
| | `REPAIR` | current plan, errors, evidence | corrected plan, signal | | |
| | `COMPRESS` | accumulated state | compact state patch, signal | | |
| | `SUCCESS` | verified evidence | terminal outcome | | |
| `EXECUTE`, deterministic observation parsing, transition checks, and final | |
| success gates should remain trusted runtime operations. Learned models propose; | |
| the graph validates and commits. | |
| ## What “overfit the graph” means | |
| The phrase describes deliberate specialization to a stable execution protocol: | |
| - fixed operator vocabulary; | |
| - explicit input projections and output schemas; | |
| - declared tools and legal transitions; | |
| - repair loops driven by real test evidence; | |
| - consistent prompts across teacher generation, training, and inference. | |
| It does **not** mean training on benchmark test patches, hidden tests, or expected | |
| answers. Benchmark instances and repositories used for final evaluation must be | |
| kept out of training and teacher-label generation. The hypothesis is that a model | |
| can learn the reusable procedure while still generalizing to unseen issues. | |
| ## SWE evaluation plan | |
| The first experiment will compare the same base model and tool environment under | |
| four scaffolds: | |
| 1. **Direct agent** — one unconstrained model call loop. | |
| 2. **Prompted operators** — operator prompts without fine-tuning. | |
| 3. **Trained operators** — the operator adapter without graph enforcement. | |
| 4. **NeoSyntropy graph** — trained operators with schemas, guards, state, and | |
| verified transitions. | |
| Evaluation will begin with the official | |
| [SWE-bench](https://www.swebench.com/) harness for reproducibility, while newer | |
| or contamination-resistant suites should be used for primary generalization | |
| claims. Static benchmark scores will always be reported with the exact harness, | |
| model, scaffold, context limit, tool budget, retry budget, and task exclusions. | |
| Primary metrics: | |
| - resolved instances (`pass@1`); | |
| - legal-transition rate; | |
| - schema-valid output rate; | |
| - tool-call success rate; | |
| - verification precision and false-success rate; | |
| - repair-loop recovery rate; | |
| - model tokens, wall-clock time, tool calls, and cost per resolved task; | |
| - average state transitions and repeated-transition rate; | |
| - outcome consistency across repeated seeds. | |
| The most important comparison is not only whether a task was solved, but whether | |
| the graph can reach the same or better result with smaller models, fewer wasted | |
| actions, and no unverified success transition. | |
| ## Data-generation protocol | |
| Training traces are generated from complete graph executions rather than isolated | |
| first-step prompts: | |
| 1. Sample an unseen repository task and construct the initial state. | |
| 2. Use a stronger teacher to label the current operator only. | |
| 3. Validate the output schema and reject invented tools or evidence. | |
| 4. Execute approved actions in the benchmark environment. | |
| 5. Record observations and test results as new evidence. | |
| 6. Route through the declared graph and continue until success or budget expiry. | |
| 7. Store each operator transition with its state projection, output, signal, | |
| provenance, and final task outcome. | |
| 8. Freeze repository-level train, validation, and test splits before fine-tuning. | |
| This produces training examples for the decision that was actually available at | |
| each state, including failed attempts and evidence-grounded repairs. | |
| ## Status | |
| - Graph contract: specified. | |
| - Role-model cards: published separately. | |
| - Unified operator dataset: planned. | |
| - Unified operator adapter: not trained yet. | |
| - SWE baseline runs: not run yet. | |
| - Benchmark claims: none. | |
| Results, weights, datasets, and exact run manifests will be published only after | |
| reproducible evaluation. | |