devils-agent / docs /ACCEPTANCE.md
devildasdf's picture
Upload experimental BAIM code, research checkpoints and measured evaluations
795f737 verified
|
Raw
History Blame Contribute Delete
3.64 kB

Full-scope acceptance tracking

The goal is still active. The following distinguishes partial artifacts from proof of the requested end state. Nothing here narrows the original requirements.

Requested deliverable Current evidence Remaining proof/work
1. Chosen architecture and reasoning Mean, GRU, Transformer implementations; experiment reports Final selection on realistic tasks and target hardware
2. Parameter-count target First trained mean baseline: 128,900 Evaluate requested 50M–1.5B range and smallest reliable option
3. Model implementation Learned classifier and pointer History, planning, argument spans, all action types
4. HF training pipeline PyTorchModelHubMixin save/load, CPU trainer Third-party data ingestion, pinned base checkpoints, reproducible larger runs
5. Dataset schema Synthetic v1 JSONL and manifest Normalized real trajectories, license/privacy gates
6. Synthetic generator Five layouts, three action classes Multi-step, dynamic, modal, scrolling, ordinal and adversarial tasks
7. Trajectory recorder HMAC metadata SQLite Reviewed redacted training trajectories and outcome linkage
8. Experience memory Not implemented Reusable retrieval with poisoned-data and privacy protection
9. SHA-256 subsystem Canonical state and signatures Correct full-state cache identity and validated cache reuse
10. Compact DOM Semantic rows, frame/shadow handling Delta integration, bounds, faithful ARIA, value/history representation
11. Action DSL Strict 18-action parser Actual tokenizer comparisons where generative policies are evaluated
12. Safe executor Authority, stale node, role and host policy gates Network/effect isolation, redirects/downloads, security audit
13. Recovery controller Loop rejection and observe action Learned recovery from enumerated browser failures
14. Completion verifier Host callback bound to task General evidence-based user-goal verification
15. CPU runtime Actual local CPU inference Linux EPYC two-vCPU end-to-end validation
16. Quantized model Linear INT8 artifact exported, reloaded and benchmarked; slower than FP32 Broader quantization/runtime comparison and deployment selection
17. Benchmark harness Browser fixtures and CPU microbenchmarks Real sites, baselines, full metric set and held-out audit
18. Regression suite Parser, state, Chromium, model, registry tests Complete action/security/generalization coverage
19. Model registry Immutable copies, integrity and promotion/rollback gates Full evaluator evidence and real champion; none promoted
20. Continual improvement Offline trainer; promotion gates Failure mining, curated dataset deltas, replay and automated cycles
21. Reproducible commands README and module CLIs Clean Linux install verification and dependency locking
22. Measured CPU/RAM/latency Local reports with measurement scope Target VPS and full service peak measurements
23. Known limitations README, DATASET, STATUS, experiment reports Keep updated as scope expands
24. Roadmap STATUS experiment order Evidence-driven next-generation plan after full evaluation

Five meaningful iterations are required but are not a substitute for acceptance. The production promotion gate requires target-hardware evidence and task, generalization, recovery and security metrics that current microbenchmarks lack. Registry unit tests use labelled fake fixtures to test transitions; those fixtures are never registered in the actual project model registry.