| # Full-scope acceptance tracking |
|
|
| The goal is still active. The following distinguishes partial artifacts from |
| proof of the requested end state. Nothing here narrows the original requirements. |
|
|
| | Requested deliverable | Current evidence | Remaining proof/work | |
| |---|---|---| |
| | 1. Chosen architecture and reasoning | Mean, GRU, Transformer implementations; experiment reports | Final selection on realistic tasks and target hardware | |
| | 2. Parameter-count target | First trained mean baseline: 128,900 | Evaluate requested 50M–1.5B range and smallest reliable option | |
| | 3. Model implementation | Learned classifier and pointer | History, planning, argument spans, all action types | |
| | 4. HF training pipeline | PyTorchModelHubMixin save/load, CPU trainer | Third-party data ingestion, pinned base checkpoints, reproducible larger runs | |
| | 5. Dataset schema | Synthetic v1 JSONL and manifest | Normalized real trajectories, license/privacy gates | |
| | 6. Synthetic generator | Five layouts, three action classes | Multi-step, dynamic, modal, scrolling, ordinal and adversarial tasks | |
| | 7. Trajectory recorder | HMAC metadata SQLite | Reviewed redacted training trajectories and outcome linkage | |
| | 8. Experience memory | Not implemented | Reusable retrieval with poisoned-data and privacy protection | |
| | 9. SHA-256 subsystem | Canonical state and signatures | Correct full-state cache identity and validated cache reuse | |
| | 10. Compact DOM | Semantic rows, frame/shadow handling | Delta integration, bounds, faithful ARIA, value/history representation | |
| | 11. Action DSL | Strict 18-action parser | Actual tokenizer comparisons where generative policies are evaluated | |
| | 12. Safe executor | Authority, stale node, role and host policy gates | Network/effect isolation, redirects/downloads, security audit | |
| | 13. Recovery controller | Loop rejection and observe action | Learned recovery from enumerated browser failures | |
| | 14. Completion verifier | Host callback bound to task | General evidence-based user-goal verification | |
| | 15. CPU runtime | Actual local CPU inference | Linux EPYC two-vCPU end-to-end validation | |
| | 16. Quantized model | Linear INT8 artifact exported, reloaded and benchmarked; slower than FP32 | Broader quantization/runtime comparison and deployment selection | |
| | 17. Benchmark harness | Browser fixtures and CPU microbenchmarks | Real sites, baselines, full metric set and held-out audit | |
| | 18. Regression suite | Parser, state, Chromium, model, registry tests | Complete action/security/generalization coverage | |
| | 19. Model registry | Immutable copies, integrity and promotion/rollback gates | Full evaluator evidence and real champion; none promoted | |
| | 20. Continual improvement | Offline trainer; promotion gates | Failure mining, curated dataset deltas, replay and automated cycles | |
| | 21. Reproducible commands | README and module CLIs | Clean Linux install verification and dependency locking | |
| | 22. Measured CPU/RAM/latency | Local reports with measurement scope | Target VPS and full service peak measurements | |
| | 23. Known limitations | README, DATASET, STATUS, experiment reports | Keep updated as scope expands | |
| | 24. Roadmap | STATUS experiment order | Evidence-driven next-generation plan after full evaluation | |
|
|
| Five meaningful iterations are required but are not a substitute for acceptance. |
| The production promotion gate requires target-hardware evidence and task, |
| generalization, recovery and security metrics that current microbenchmarks lack. |
| Registry unit tests use labelled fake fixtures to test transitions; those fixtures |
| are never registered in the actual project model registry. |
|
|