EvSpark — lossless speculative decoding for Evo2
Drafter checkpoints for EvSpark: small distilled drafters that accelerate single-stream Evo2 (StripedHyena2) generation while staying distributionally lossless — greedy output is token-for-token identical to native decoding.
Paper (bioRxiv) · DOI · GitHub · ModelScope
| Target | Suite speedup (all-48 / real-43) | Setup |
|---|---|---|
Evo2 7B, flagship L27_g12_150M |
3.27× / 2.96× | RTX 4090, flash-attn, 3 training seeds |
Evo2 7B, cost-optimal L27_g12_30M |
3.15× / 2.82× | ~1.06 GPU-h distill on one 4090 |
Evo2 20B phase3_20b_L20_g12_b1/b2 |
2.53×/2.18× · 2.51×/2.22× | H20, SDPA, 1 seed |
Evo2 40B phase3_40b_L45_g12_b1/b2/b3 |
2.72×/2.44× · 2.78×/2.46× · 2.72×/2.42× | H20, SDPA, 1 seed |
Suite = 48 genomic prompts × 1024 tokens, single stream. Long context: E. coli 262k 1.97–2.43×, B. subtilis 262k 1.84–2.05×. Greedy losslessness: 0 non-tie divergences (48 prompts × 6 ckpts). One checkpoint serves any decode-time draft length γ′ ≤ training γ (exact causal prefix; no retraining). A predictor-guided regulatory-DNA design workflow sees a median 1.57× complete-design speedup over a calibrated batched native baseline (see paper).
Checkpoints
All drafters are single-layer injection (7B: blocks.27; 20B: blocks.20; 40B: blocks.45), d_model=1024, distilled offline from frozen Evo2 hidden states; the target model is never fine-tuned. Each .pt is self-contained (frozen embedding, Markov head, confidence head, metadata). 20B/40B drafters need their targets' official Transformer-Engine path (measured on one H20 96 GB; no official path on Ada).
| File | Target | Distill budget | Role |
|---|---|---|---|
L27_g12_150M_s1.pt / _s2.pt |
7B | 150M | flagship (default) |
L27_g12_30M_s1.pt / _s2.pt |
7B | 30M | ~1 GPU-hour cell |
phase3_20b_L20_g12_b1.pt / _b2.pt |
20B | 10M / 30M | larger-target transfer |
phase3_40b_L45_g12_b1.pt / _b2.pt / _b3.pt |
40B | 10M / 30M / 80M | larger-target transfer |
SHA-256
| File | sha256 |
|---|---|
L27_g12_150M_s1.pt |
be6ff9f8818a83533e39d112ee9ba0028a340fe3cead8dc342c2397baf32310a |
L27_g12_150M_s2.pt |
654576ae3713fb4a0bc767c70f9bdaddb2f9b468a1019d14c8fff659d693af40 |
L27_g12_30M_s1.pt |
18552e63f9facf853c7b8df31edaafaac4dec7276c9524cefa038f11b6cbc7c1 |
L27_g12_30M_s2.pt |
571eb2ccdbef5f11026157908e29cf06ce76c6587d437da1e6797c40cf2ab6df |
phase3_20b_L20_g12_b1.pt |
2f308c71c80d9b919aa2d11a2d7ee98488251ee64bd4505ceecd129820f630d3 |
phase3_20b_L20_g12_b2.pt |
3d9d4d9343935f18cfeec3731447566cff17e8a9d08df4cffaebcda771066c1c |
phase3_40b_L45_g12_b1.pt |
81241aaeb0cc49d5c8559085b8bb2d78d6a6f6e87a5ac099e2e60734446d1260 |
phase3_40b_L45_g12_b2.pt |
87507bc81edbde067392c8d8e35f397945af2cd6b1bd5ad666f3ecdc371f8930 |
phase3_40b_L45_g12_b3.pt |
95e75641ca48fba95d5500f339c91ef4f7ac2567538c6063f744afdff35c1b58 |
Quickstart
These files are drafters, not a standalone DNA LM. You need Evo2 + the EvSpark engine (block verify + Hyena/attention state-slice rollback):
git clone https://github.com/dhnihaoya/EvSpark && cd EvSpark # env setup: see GitHub README
python scripts/download_ckpt.py L27_g12_150M_s1 # HF first, ModelScope fallback
python scripts/demo.py --ckpt L27_g12_150M_s1 --n-tokens 1024
from evspark import EvSpark
with EvSpark.load("L27_g12_150M_s1") as es:
out = es.generate("ACGTACGT...", n_tokens=1024) # T=1.0, top_k=4
g = es.generate(prompt, greedy=True)
assert g.ids.tolist() == es.generate_native(prompt, greedy=True).ids.tolist()
If huggingface.co is unreachable: HF_ENDPOINT=https://hf-mirror.com, or pull the same files from ModelScope.
Citation
@article{ding2026evspark,
title = {EvSpark: Lossless Speculative Decoding for Hybrid DNA Foundation Models},
author = {Ding, Hao and Wu, Nannan and Qiu, Tianyi},
journal = {bioRxiv},
year = {2026},
doi = {10.64898/2026.09.02.749017},
url = {https://www.biorxiv.org/content/10.64898/2026.09.02.749017}
}
License
Checkpoints and code are MIT. Evo2 / Vortex weights and runtime follow their upstream licenses.
Model tree for dinghhhhhhhhhhhhhhh/EvSpark
Base model
arcinstitute/evo2_7b