Safetensors
benchmark-models / README.md
Caesarrr's picture
Index and verify 65 Standard-Pipeline model imports
ad37dff verified
|
Raw History Blame Contribute Delete
8.73 kB
---
license: other
---
# Latency-Sensitive Bench models
Inference-ready teachers and VLA policies for the supported benchmark tasks.
## Layout
- `zero-latency/<task>/small-policy/` and `zero-latency/<task>/vla/`: models trained without latency.
- `latency-aware/<task>/small-policy/` and `latency-aware/<task>/vla/`: models trained for latency. Profile and fixed-2 training conditions are identified by run ID and bundle provenance.
MIKASA H8 conditioned inference bundles use `latency-aware/mikasa-intercept-grab-fast/vla/starvla-<model>-h8/<run-id>/`, where `<model>` is `qwenoft`, `qwenpi_v3`, or `qwengr00t`. Each model has an SFT run `h8-conditioned-seed-reset-20260922-{profile|fixed-2}` and a DAgger run `h8-conditioned-dagger-20260923-{profile|fixed-2}`.
## Hist8 VLA release: 18 models, 27 formal evaluations
Flappy, Demon Attack and Deadly Corridor × QwenOFT, QwenGR00T and QwenPI v3 × zero/profile training. Each model is the evaluated step-5000 checkpoint, initialized from Qwen3-VL-4B revision `ebb281ec70b05090aa6165b016eac8ec08e71b17` with seed 42.
Input is one current RGB image plus eight causal decision histories, passed as raw 0/1 transport state. Action horizon is **1**; the model directory suffix `h1` refers to this action horizon, while `hist8` in each run ID refers to input history.
### Formal scores
Each value is the mean of 100 episodes, seeds 1,000,000–1,000,099; parallel 32 and capacity 1. Flappy/Demon/Deadly caps are 3,600/7,200/3,600 raw frames at FPS/decision Hz 10/10, 60/15, 35/8.75. Deadly uses the corrected sf-render-v4 view (160×120 RGB with HUD).
| Game / model | zero→zero | zero→profile | profile→profile |
|---|---:|---:|---:|
| flappy / qwenoft | 439.13 | 19.19 | 373.78 |
| flappy / qwengr00t | 428.48 | 6.62 | 414.52 |
| flappy / qwenpi_v3 | 408.72 | 5.58 | 344.48 |
| demon_attack / qwenoft | 2362.15 | 584.35 | 1452.80 |
| demon_attack / qwengr00t | 2370.90 | 231.90 | 1203.10 |
| demon_attack / qwenpi_v3 | 2350.95 | 112.30 | 907.30 |
| deadly_corridor / qwenoft | 2098.67 | 742.35 | 2091.23 |
| deadly_corridor / qwengr00t | 2100.46 | 341.76 | 2090.82 |
| deadly_corridor / qwenpi_v3 | 2104.09 | 140.94 | 1450.83 |
### Downloadable models
| Training | Game / model | Run |
|---|---|---|
| profile | flappy / qwenoft | [flappy-qwenoft-hist8-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/latency-aware/flappy/vla/starvla-qwenoft-h1/flappy-qwenoft-hist8-scratch-s42-v2) |
| profile | demon_attack / qwenoft | [demon_attack-qwenoft-hist8-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/latency-aware/demon-attack/vla/starvla-qwenoft-h1/demon_attack-qwenoft-hist8-scratch-s42-v2) |
| profile | flappy / qwengr00t | [flappy-qwengr00t-hist8-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/latency-aware/flappy/vla/starvla-qwengr00t-h1/flappy-qwengr00t-hist8-scratch-s42-v2) |
| profile | demon_attack / qwengr00t | [demon_attack-qwengr00t-hist8-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/latency-aware/demon-attack/vla/starvla-qwengr00t-h1/demon_attack-qwengr00t-hist8-scratch-s42-v2) |
| profile | demon_attack / qwenpi_v3 | [demon_attack-qwenpi_v3-hist8-scratch-s42-v2-vlm3e-6](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/latency-aware/demon-attack/vla/starvla-qwenpi_v3-h1/demon_attack-qwenpi_v3-hist8-scratch-s42-v2-vlm3e-6) |
| zero | flappy / qwenoft | [flappy-qwenoft-hist8-zero-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/zero-latency/flappy/vla/starvla-qwenoft-h1/flappy-qwenoft-hist8-zero-scratch-s42-v2) |
| profile | flappy / qwenpi_v3 | [flappy-qwenpi_v3-hist8-scratch-s42-v2-vlm3e-6](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/latency-aware/flappy/vla/starvla-qwenpi_v3-h1/flappy-qwenpi_v3-hist8-scratch-s42-v2-vlm3e-6) |
| zero | demon_attack / qwenoft | [demon_attack-qwenoft-hist8-zero-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/zero-latency/demon-attack/vla/starvla-qwenoft-h1/demon_attack-qwenoft-hist8-zero-scratch-s42-v2) |
| zero | flappy / qwengr00t | [flappy-qwengr00t-hist8-zero-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/zero-latency/flappy/vla/starvla-qwengr00t-h1/flappy-qwengr00t-hist8-zero-scratch-s42-v2) |
| zero | flappy / qwenpi_v3 | [flappy-qwenpi_v3-hist8-zero-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/zero-latency/flappy/vla/starvla-qwenpi_v3-h1/flappy-qwenpi_v3-hist8-zero-scratch-s42-v2) |
| zero | deadly_corridor / qwenoft | [deadly_corridor-qwenoft-hist8-zero-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/zero-latency/deadly-corridor/vla/starvla-qwenoft-h1/deadly_corridor-qwenoft-hist8-zero-scratch-s42-v2) |
| zero | demon_attack / qwengr00t | [demon_attack-qwengr00t-hist8-zero-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/zero-latency/demon-attack/vla/starvla-qwengr00t-h1/demon_attack-qwengr00t-hist8-zero-scratch-s42-v2) |
| profile | deadly_corridor / qwenpi_v3 | [deadly_corridor-qwenpi_v3-hist8-scratch-s42-v2-vlm3e-6](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/latency-aware/deadly-corridor/vla/starvla-qwenpi_v3-h1/deadly_corridor-qwenpi_v3-hist8-scratch-s42-v2-vlm3e-6) |
| profile | deadly_corridor / qwenoft | [deadly_corridor-qwenoft-hist8-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/latency-aware/deadly-corridor/vla/starvla-qwenoft-h1/deadly_corridor-qwenoft-hist8-scratch-s42-v2) |
| profile | deadly_corridor / qwengr00t | [deadly_corridor-qwengr00t-hist8-scratch-s42-v2-workers16](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/latency-aware/deadly-corridor/vla/starvla-qwengr00t-h1/deadly_corridor-qwengr00t-hist8-scratch-s42-v2-workers16) |
| zero | demon_attack / qwenpi_v3 | [demon_attack-qwenpi_v3-hist8-zero-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/zero-latency/demon-attack/vla/starvla-qwenpi_v3-h1/demon_attack-qwenpi_v3-hist8-zero-scratch-s42-v2) |
| zero | deadly_corridor / qwengr00t | [deadly_corridor-qwengr00t-hist8-zero-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/zero-latency/deadly-corridor/vla/starvla-qwengr00t-h1/deadly_corridor-qwengr00t-hist8-zero-scratch-s42-v2) |
| zero | deadly_corridor / qwenpi_v3 | [deadly_corridor-qwenpi_v3-hist8-zero-scratch-s42-v2](https://huggingface.co/latency-sensitive-bench/benchmark-models/tree/main/zero-latency/deadly-corridor/vla/starvla-qwenpi_v3-h1/deadly_corridor-qwenpi_v3-hist8-zero-scratch-s42-v2) |
Every model contains its README, checkpoint, inference/training configuration, statistics, task contract, collection/training provenance, exact latency profile and episode-level evaluation evidence. The original detailed JSON index remains in the benchmark publication receipt.
### Interpretation and reproduction
The same zero-trained checkpoint is used in both evaluation environments. Each architecture has its own latency profile, so this is not a same-latency cross-architecture ranking. Zero and profile training use different teachers and collected data; differences do not isolate learning rate or a single training factor. Results use one training seed.
Common training: two GPUs, bf16 ZeRO-2, batch 64/rank, accumulation 1, global batch 128, 16 workers/rank, prefetch 4, backbone LR 3e-6, action-head LR 1e-4, 100-step warmup, cosine min_lr_rate 1/30, 5,000 steps. Architecture-specific losses and other parameter groups are retained.
Hist8 corrects the earlier -1/1 history input mismatch. Deadly also corrects an evaluation rendering mismatch; invalid older results are not included in this release. Profiles are preserved byte-for-byte with their sampling sidecars. Use each model task contract and pinned base-model revision when replaying.
## Standard-Pipeline import: seven Gymnasium tasks
AirRaid, Ant, HalfCheetah, Hopper, Humanoid, InvertedPendulum and Walker2d: 37 H1 VLA bundles (QwenOFT, QwenGR00T, QwenPI v3) and 28 Sample Factory APPO bundles. Each bundle contains one selected checkpoint. These are source-preserving copies with recorded experiment status; this import does not imply new evaluation or acceptance.
[Model inventory and verification](reports/standard-pipeline-2208875f92b2/README.md). HalfCheetah lacks profile QwenGR00T/QwenPI v3 VLAs; Humanoid lacks all three profile VLAs.