Action1.5 policy — frozen base plus simulator side policy
Upload the extracted contents of this ZIP directly to the private repository summerMC/action1.5-policy. There is no enclosing directory. No HF upload or account change has been performed.
model.safetensors is the original locally cached frozen Action1 base (48,927,232 unique parameters), converted losslessly without training. policy.safetensors is the separate evolved 30-parameter float64 side policy, ES seed1512. The base was not fine-tuned. This is a bounded simulator action selector, not a standalone conversational model, autonomous planner or demonstrated general intelligence.
Use
Python3.11 and the dependencies in requirements.txt are required. No downloads occur during inference:
python run.py --seed 777 --length 2
The entrypoint loads both safetensors files and the included tokenizer by relative paths. It runs an actual base forward check and simulator episode with live base encoder features. No disk feature bank or undeclared model cache is needed. CPU only; no real GUI executor. It is an explicit custom loader, not a promise that generic Transformers AutoModel/generate works.
The side policy uses a handcoded visible checklist, exact goal/value matching and cue-to-route support. Arguments copy visible goal values. A learned scalar memory retains an early cue;16 projected observation–candidate products come from the live frozen base. Supported demo field counts are1–8. This is substantial task scaffolding, not planning learned from scratch.
Provenance and license
Included base source: local cache summerMC/action1, revision d898d67af871b23fcc6a88c75762f64681990add. File hashes match the prior local verification record. Current read-only HF metadata for that name resolves to j-llm/action1 and declares Apache-2.0; the declared j-llm/j72-tokenizer repo also declares Apache-2.0. The exact local tokenizer.json/config are included, not guessed. Vocabulary72000 matches embeddings and local inference.
Exact remote snapshot hash equivalence and the historical training-tokenizer provenance remain unverified. Original repo LICENSE/NOTICE files were not available; LICENSE-2.0.txt supplies the standard declared Apache2.0 terms. See ATTRIBUTION.md and provenance.json for the exact source, attribution limits and conversion notices. No invented original copyright notice or upstream identity claim is made.
Existing pilot results
No new training was performed for this bundle. In the previous locked test, each horizon had32 episodes =16 cue seeds ×2 new workflow graphs. The handcoded reactive baseline had68.75% success because11/16 cues were amber, repeated on both graphs; scripted memory verifier100%, random policy0%. They are not32 independent cue samples.
| Encoder | ES seed | Short success | Long success |
|---|---|---|---|
| Local cached Action1 | 1511 | 0% | 0% |
| Local cached Action1 | 1512 | 100% | 84.375% |
| Local cached Action1 | 1513 | 50% | 50% |
| Same-architecture random | 1511 | 0% | 0% |
| Same-architecture random | 1512 | 0% | 0% |
| Same-architecture random | 1513 | 0% | 0% |
The included policy was strongest on development evaluation before test; all seeds are shown to avoid hiding instability. Only three optimizer seeds and one random backbone seed were tested; there is no established pretrained advantage. Memory decay causes longer-horizon errors. This ZIP omits comparison weights/logs/history but retains the honest result summary.
Contents and validation
Minimum use files only: frozen base safetensors; exact custom model code and tokenizer; inference-only base config; selected side-policy safetensors/config; simulator/policy and entrypoint; dependencies; README/license/attribution; compact provenance and checksum manifest. Original .bin, other seeds, random-control weights, .npy duplicates, feature banks, training code, logs, private paths/receipts and caches are excluded. Source project/history is untouched.
Base state-dict conversion restored every key with exactly equal values/dtypes, including tied embeddings. Side-policy serialization previously passed bitwise roundtrip and24 development-episode/265-action replay equivalence. Clean extracted offline load/smoke is reported below; it validates integration, not a new success benchmark.
Clean extracted offline smoke: PASSED. New isolated process loaded both weights and included tokenizer with offline flags/empty model cache locations, executed actual base forward, then completed demo seed777 with zero invalid actions. Exact outputs: 8 steps, 19 live encoder calls. No feature-bank file or source cache was used.
- Downloads last month
- -