- HeliosLM β A Hackable DeepSeek-V3/K3-Style LLM Stack in Pure PyTorch
- Who is this for?
- Harness Plugins (v1.0-v1.22, 23 releases)
- Paper Highlights (for HF Papers)
- Paper
- Citation
- Plugins (helios-harness)
- Deployment tiers
- Quick Start
- Capabilities at a glance
- Why HeliosLM vs. alternatives?
- Benchmarks
- Repository layout
- Roadmap
- Contributing
- Known Limitations
- Version history
- License
- Who is this for?
HeliosLM β A Hackable DeepSeek-V3/K3-Style LLM Stack in Pure PyTorch
A from-scratch PyTorch reference implementation of a modern LLM stack: MLA attention with weight absorption, sigmoid-gated MoE with auxiliary-loss-free load balancing, hybrid linear attention, speculative decoding, FP8 training, a DualPipe schedule simulation, a vLLM-style serving engine, a verifiable agent layer (strict tool schema, bitwise-replay oracle), DSA sparse attention, Mooncake-style prefill/decode disaggregation, and tool-tuned checkpoints. Built to be read, modified, and verified β every core path is unit-tested and many are checked with bitwise-equivalence tests. Everything runs on CPU.
One-liner: If you want to understand (or hack on) how DeepSeek-V3/K3-class models actually work β without needing a GPU cluster first β this repo is for you.
Also: the open-source reference for the RLCD "decision layer" paradigm β calibrated confidence on acting models, verified across three scales and on the 7,193-instance RLCDAlignBench. Working paper:
docs/paper_draft_2026-10-03.tex(repo | HF).
Who is this for?
| You are... | What HeliosLM gives you |
|---|---|
| A learner who wants to understand MLA, MoE routing, DualPipe, speculative decoding | Annotated, review-hardened PyTorch with 115+ tests (T11-T28) that act as executable documentation |
| A researcher who wants a stack to modify, ablate, and extend quickly | Single-process, CPU-iterable training + serving code β change one file, run one test |
| A practitioner evaluating serving/quantization techniques | vLLM-style paged engine, GPTQ/AWQ/FP8/MXFP4 quantization, MTP speculative decoding β all inspectable |
Honest positioning: this is a correctness-focused reference implementation, not a throughput-optimized production engine (see Known Limitations).
Harness Plugins (v1.0-v1.22, 23 releases)
All plugins register as ctx.<name>.<op> with per-call audit effects.
Install extras as needed: pip install openpyxl python-docx python-pptx pytesseract pyserial cryptography psycopg2 boto3 google-cloud-storage azure-storage-blob pyttsx3 speech_recognition ultralytics transformers
Core (always available)
| Plugin | Namespace | Key ops |
|---|---|---|
| ModelPlugin | ctx.model.* |
mid (360M HeliosLM), qwen (Qwen3-0.6B) |
| ToolPlugin | ctx.tools.* |
registry, impls (calc/str/file_read/file_write/finish) |
| SessionPlugin | ctx.session.* |
new (append-only, replay-verified) |
| DecisionPlugin | ctx.decision.* |
grounding, trust, prefetcher |
| LoopPlugin | ctx.loop.* |
make (AgentLoop factory) |
| StatePlugin | ctx.state |
dict-backed, cross-plugin persistence |
| PresetPlugin | ctx.preset |
minimal / full wiring |
Agent (v1.8)
| Plugin | Namespace | Key ops |
|---|---|---|
| MemoryPlugin | ctx.memory.* |
remember, recall (ExperienceStore) |
| TerminalPlugin | ctx.terminal.* |
run (TrustGate-gated) |
| FetchPlugin | ctx.fetch.* |
get (HTTP) |
| FilesystemPlugin | ctx.fs.* |
read, write (path-escape blocked) |
| TimePlugin | ctx.time.* |
now, utc |
Common (v1.9)
| Plugin | Namespace | Key ops |
|---|---|---|
| SearchPlugin | ctx.search.* |
query (local corpus) |
| WebSearchPlugin | ctx.web.* |
search (Tavily/Exa stub) |
| PDFPlugin | ctx.pdf.* |
extract (pypdf) |
| SQLitePlugin | ctx.sqlite.* |
query (constrained db) |
| TemplatePlugin | ctx.plugin.* |
scaffold (new-plugin template) |
Services (v1.10)
| Plugin | Namespace | Key ops |
|---|---|---|
| GitHubPlugin | ctx.github.* |
get_repo, list_issues, create_issue (TrustGate-gated) |
| PostgresPlugin | ctx.postgres.* |
query (psycopg2/pg8000) |
Office (v1.11)
| Plugin | Namespace | Key ops |
|---|---|---|
| ExcelPlugin | ctx.excel.* |
read, write (openpyxl) |
| DocxPlugin | ctx.docx.* |
read, write (python-docx) |
| PptxPlugin | ctx.pptx.* |
write (python-pptx) |
| YouTubePlugin | ctx.youtube.* |
transcript (youtube-transcript-api) |
Niche (v1.12)
| Plugin | Namespace | Key ops |
|---|---|---|
| MCPWizardPlugin | ctx.mcp.* |
wrap (MCP server) |
| TmuxPlugin | ctx.tmux.* |
new, list, send |
| EverythingPlugin | ctx.everything.* |
search (local file content) |
Media (v1.13)
| Plugin | Namespace | Key ops |
|---|---|---|
| VideoPlugin | ctx.video.* |
generate (Runway/Pika stub) |
| AudioPlugin | ctx.audio.* |
tts (pyttsx3/external), stt |
| NanoVideoPlugin | ctx.nanovideo.* |
from_images (ffmpeg slideshow) |
Hardware (v1.14, v1.16, v1.18, v1.21)
| Plugin | Namespace | Key ops |
|---|---|---|
| BlenderPlugin | ctx.blender.* |
gen, run (headless blender) |
| OmiPlugin | ctx.omi.* |
connect, transcribe (BLE wearable) |
| RobotControlPlugin | ctx.robot.* |
move_base, move_arm, gripper (mock/ROS2, TrustGate-gated) |
| MultiRobotPlugin | ctx.fleet.* |
register, allocate, formation |
| HardwarePlugin | ctx.hw.* |
arduino_write, gpio_write, i2c_write |
Science (v1.15, v1.20-v1.22)
| Plugin | Namespace | Key ops |
|---|---|---|
| BenchmarkPlugin | ctx.benchmark.* |
alignbench, api_probe |
| ReportPlugin | ctx.report.* |
gen (markdown from benchmarks) |
| BenchmarkV2Plugin | ctx.bench.* |
run_suite, trend |
| DNAPlugin | ctx.dna.* |
validate, reverse_complement, gc, transcribe, translate, align, pcr_primers, restriction_sites |
| ProteinPlugin | ctx.protein.* |
validate, mol_weight, hydrophobicity, fold_toy |
| ChemPlugin | ctx.chem.* |
formula_weight, ph, bond_energy |
| MathV2Plugin | ctx.linalg.*, ctx.stats.*, ctx.signal.* |
matmul, summary, fft_magnitudes |
Collaboration (v1.17, v1.19)
| Plugin | Namespace | Key ops |
|---|---|---|
| AgentSwarmPlugin | ctx.swarm.* |
spawn, delegate, aggregate, blackboard |
| VoiceDialogPlugin | ctx.voice.* |
turn (STT->LLM->TTS with barge-in) |
| TerminalUIPlugin | ctx.ui.* |
progress, menu |
Autonomy (v1.18, v1.19)
| Plugin | Namespace | Key ops |
|---|---|---|
| SLAMPlugin | ctx.slam.* |
create_map, scan, integrate, frontier |
| VisionPlugin | ctx.vision.* |
detect, ocr, caption |
| AutoDrivePlugin | ctx.autodrive.* |
simulate (lane keep + obstacle avoid) |
Security / Cloud / RL (v1.21-v1.22)
| Plugin | Namespace | Key ops |
|---|---|---|
| CryptoPlugin | ctx.crypto.* |
hash, gen_key, xor_encrypt/decrypt, hmac_sign/verify, fernet |
| RLPlugin | ctx.rl.* |
train_grpo, rlcd_reward |
| CloudPlugin | ctx.cloud.* |
aws_s3_list, gcp_storage_list, azure_blob_list |
Paper Highlights (for HF Papers)
One-liner: The open-source reference for the RLCD "decision layer" -- calibrated confidence on acting models, measured across three scales (8.5M / 360M / Qwen3-0.6B) and on the 7,193-instance RLCDAlignBench.
Why this matters for the HF community:
- Open-weight models now span the full price spectrum ($0 -> $0.14/M -> $0.20/M frontier tiers) but NO vendor publishes calibration-on-wrong- answers curves. This repo is the open, replay-verifiable measurement.
- The API probe (GPT-5.6 Luna + DeepSeek V4 Flash) found THREE distinct calibration failure modes across the price spectrum: stable hallucination (360M/Qwen), unstable hallucination (Luna), refusal (V4 Flash) -- none healthy. Probe cost: $0.02.
- On RLCDAlignBench our supervised TF-IDF readout reaches 0.726 median AUROC and beats the commercial Jev detector's zero-shot numbers on 11/41 benchmarks. The open audit is 15x cheaper than the commercial detector it audits.
- HeliosLM's policy-in-code route (GroundingGate, CalibratedPrefetcher) is the engineering embodiment of the vendor's own finding: prompt engineering adds only +0.006 AUROC (p=0.055) -- retaining full probability and routing decisions in code is what works.
Artifacts: paper PDF + LaTeX in docs/; all benchmark JSONs in
benchmarks/; three-tier deployment (lite 8.5M / mid 360M / full V3-class)
with per-task replay verification and tokenizer-fingerprint pairing.
cc @osanseviero @philschmid for HF Papers consideration
Paper
Calibrated Agency: An Open-Source RLCD Stack, from Toy Scale to 360M, with a Canonical-Benchmark Comparison β Chien-Hsin Lin, working draft 2026-10-03.
π Read the PDF (7pp, figures included) Β·
LaTeX source (arXiv-ready) Β·
markdown Β·
HF copy.
Headline numbers: overconfidence on wrong answers grows with scale
(0.94 -> 0.9999 -> ~1.0 at 8.5M/360M/Qwen3-0.6B); deterministic grounding
cures a 360M model's copy-shaped agentic failure (0/12 -> 9/12 + 3
abstained); trust is f(state, intervention) with an 8.4x measured
contrast; on RLCDAlignBench our open readout reaches 0.726 median AUROC
and beats the commercial Jev detector's zero-shot numbers on 11/41
benchmarks. Every claim traces to a versioned artifact in benchmarks/.
Citation
@article{lin2026calibrated,
title={Calibrated Agency: An Open-Source {RLCD} Stack,
from Toy Scale to 360M, with a Canonical-Benchmark Comparison},
author={Lin, Chien-Hsin},
year={2026},
note={Working draft. arXiv ID: [pending endorsement]},
url={https://github.com/tonythetiger168/helioslm}
}
Every numeric claim traces to a versioned artifact: code and tests in
this repository (CHANGELOG v5.30.2--v5.38j), weights/tokenizer/
fingerprints and benchmark data in
huggingface.co/chienhsinlin/helioslm,
and the paper source + PDF in docs/.
Plugins (helios-harness)
The harness ships with 35+ plugins covering DeepSeek-Harness scenarios and beyond.
All plugins register as ctx.<name>.<op> with per-call audit effects.
| Category | Plugins | Version |
|---|---|---|
| Core | Model, Tool, Session, Decision (grounding/trust/prefetcher), Loop, State, Preset | v1.1βv1.3 |
| Agent | Memory, Terminal (TrustGate-gated), Fetch, Filesystem, Time | v1.8 |
| Common | Search (local), WebSearch (stub), PDF, SQLite, Template | v1.9 |
| Services | GitHub (gated), PostgreSQL | v1.10 |
| Office | Excel, Docx, Pptx, YouTube | v1.11 |
| Niche | MCP-wizard, Tmux, Everything | v1.12 |
| Media | Video, Audio (TTS/STT), NanoVideo | v1.13 |
| Hardware | Blender, Omi (BLE) | v1.14 |
| Robot | RobotControl (mock/ROS2), MultiRobot, SLAM, Vision | v1.16βv1.18 |
| Agentic | AgentWorkflow (ReAct), AgentSwarm | v1.2, v1.17 |
| Science | DNA, Protein, Chem, Math-v2 | v1.20βv1.22 |
| RL | GRPO training, RLCD reward shaping | v1.21 |
| Crypto | Hash/XOR/HMAC/Fernet | v1.22 |
| Cloud | AWS S3, GCP Storage, Azure Blob | v1.22 |
| Benchmark | AlignBench summary, Report gen, Trend | v1.15, v1.22 |
| Voice/UI | VoiceDialog, TerminalUI, AutoDrive | v1.19 |
| IC | RTL gen, UVM scaffold, Coverage | v1.7 |
Usage:
from harness.harness_core import Context
from harness.harness_plugins import *
ctx = Context()
ctx.use(PresetPlugin("full", model_fn=my_fn))
ctx.use(RobotControlPlugin(mode="mock"))
ctx.robot.move_base(1.0, 0.0, 0.0) # TrustGate-gated, effect-audited
Deployment tiers
Three rungs, each with one recorded reason to exist (no spectrum theater):
| Tier | Params | Vocab | Weights (bf16) | Deployment RAM | Why it exists |
|---|---|---|---|---|---|
| lite | 8.5M | 1,024 char-level | ~17 MB | <500 MB, CPU, millisecond latency | Protocol/audit research at $0: agent, chat, decision layer, replay verification |
| mid (v5.33) | 360M | 32,768 BPE (HeliosBPE, T31) | ~720 MB | ~1.5 GB, CPU-runnable inference | Scale validation: separate toy artifacts from scale-invariant findings (acceptance oracles T32; GPU training run pending) |
| full | DeepSeek-V3/Kimi-K3-class spec | 160,000 | fp8 still needs hundreds of GB HBM | H200/B200-class servers | Architecture decision reference (MLA + MoE + FP8 + MTP), not a local target |
Any intermediate size is constructible by explicit field overrides --
caller-provided non-None values always win over presets and are validated
in __post_init__ (no silent truncation).
Quick Start
pip install torch
python -m helioslm_v5.tests.test_v5 # 48 unit tests
python integration_test_v51.py # 9 end-to-end integration tests
python -m helioslm_v5.tests.test_agent # 9 agent-layer oracles (v5.23)
python -m helioslm_v5.tests.test_v5_stage_a # T15 real-model oracles (v5.26)
python -m helioslm_v5.tests.test_v5_stage_b # T17 tool-tuned end-to-end (v5.27)
python -m helioslm_v5.tests.test_disagg_pareto # T18 Pareto sweep oracles (v5.28)
python helioslm_v5/tests/test_chat.py # T20 chat capability oracles (v5.30)
python helioslm_v5/tests/test_mhc.py # T22 mHC oracles (v5.31)
python helioslm_v5/tests/test_kv_compress.py # T23 compressed-attention causality (v5.31)
python helioslm_v5/tests/test_file_env.py # T24 long-horizon env (v5.31)
python helioslm_v5/tests/test_think.py # T25 think + experience reuse (v5.31)
python helioslm_v5/tests/test_async_grpo.py # T26 async GRPO sync-parity (v5.31)
python helioslm_v5/tests/test_decision.py # T27 typed decision primitives (v5.32)
python helioslm_v5/tests/test_decision_head.py # T28 DecisionHead calibration (v5.32)
from helioslm_v5.configs.config_v5 import HeliosLMv5Config
from helioslm_v5.src.model_v5 import HeliosLMv5
model = HeliosLMv5(HeliosLMv5Config(size="lite")) # CPU-friendly
out = model.generate([[1, 2, 3]], max_new_tokens=20, temperature=0)
print(out)
Capabilities at a glance
| Area | Implementation |
|---|---|
| Attention | MLA with weight absorption β latent-only KV cache, β97.7% memory vs MHA (full config), verified equivalent to the expanded path (<1e-4). Hybrid linear attention: Gated Delta Rule layers interleaved with MLA (3:1 default), fixed-size recurrent state cache, decode β‘ one-shot (<2e-7), optional per-channel (KDA-style) decay gate so state rows forget at independent rates (v5.21), NoPE option to drop RoPE entirely (v5.22). RoPE scaling: linear / NTK / YaRN. DSA-style sparse top-k decode over the latent cache (kβ₯L exactly dense), with a learned lightning indexer option (ReLU-scored low-dim heads reading the cached latents, DSA-style distill training helper; default head-mean "free" indexer bit-identical) and opt-in sparse prefill. Sliding-window attention with StreamingLLM sinks (O(W) decode), per-head QK-norm, Gemma-style logit soft-capping. |
| MoE | Sigmoid-gated fine-grained experts with auxiliary-loss-free load balancing (selection-only bias, heuristic or quantile updates). LatentMoE: routed experts in a shared latent space. SiTU-GLU tanh soft-capped activation. |
| Cross-layer | Attention Residuals β per-layer gated injection of accumulated lower-layer attention outputs, threaded through DualPipe (gradient-exact, bitwise-verified). |
| Agent | v5.23 agent layer: strict tool-call schema/parser (16 error classes), deterministic sandboxed tools, trajectory bitwise-replay oracle, ground-truth-by-construction toy envs, routing gates (v5.22 decision-audit discipline); v5.26 real-model oracles (T15); v5.27 tool-tuned checkpoint trained on agent-loop replays (data format == inference by construction) β T17 end-to-end baseline parse 0.15 / finish 1/9, sparse_top_k=4 ~= dense; hardened by real-model findings (TOOL_ERROR recovery, ASCII-safe docs) ; v5.29 three-modes benchmark with strict-monotone tau routing curves and a recorded overconfidence finding (max conf 0.925-0.944 on wrong answers), artifact oracles T19 |
| Chat | v5.30: dual-mode protocol (plain text OR @@tool@@ block; text bypasses the gate, gate governs tools only), multi-turn ChatSession with transcript replay, chat SFT data with inference-identical prompts; v5.30.2 fixed a dataset filter bias that had hidden the text channel (mode-choice: text 11/20) β T20/T21 |
| Frontier references | v5.31: five readable toy-scale references distilled from the 2026 frontier β manifold-constrained hyper-connections (DS-V4), HCA/CSA-style compressed attention with exact causality, long-horizon file env (8-14 steps), think-mode + experience reuse (Qwen3-Max direction), async GRPO with bitwise sync-parity (GLM-5 direction) β T22-T26 |
| Decision layer | v5.32: typed decision primitives Choice/Score/Noul (System One / Jev direction) with schema enforcement, gate ask/ask_batch extension point, non-autoregressive DecisionHead answering K questions in one pass, outcome-targeted Brier calibration (RLCD direction; the label-targeted variant is provably redundant with CE β recorded) β T27/T28 ; v5.36 continuous Noul (P(yes), floor retired); v5.34 deterministic grounding (0/12 -> 9/12 + 3 abstained on a real 360M checkpoint, zero hallucination leakage); v5.37 TrustGate calibrated abstention + TherapyPair composition (trust is f(state, intervention) β 8.4x measured contrast); v5.38 RLCDAlignBench: our supervised TF-IDF readout median 0.726 AUROC, beats the commercial Jev detector's zero-shot numbers on 11/41 benchmarks (charts in benchmarks/charts/) β T27-T39 |
| Benchmarks | Top-10 LLM position paper with verified leaderboard data and the calibration/replay axes no vendor publishes (helioslm_v5/docs/benchmark_top5_2026-09-27.md); probe suite (toy + file-env + chat) with scripted oracles, runnable against any API |
| Disaggregation | v5.25 Mooncake-style prefill/decode module behind a monotonicity gate; v5.28 three-axis Pareto sweep (makespan / workers / worker-seconds) with latency-cost curves per workload β cache-aware anti-monotonicity recorded as a structural finding, not hidden |
| Speculative decoding | DeepSeek-style MTP with strict verification (residual (pβq)β resampling), batch support, O(1) cache-truncation rollback; hybrid recurrent-state rollback via restore+replay. |
| Serving | vLLM-style engine: paged KV accounting, copy-on-write forks, watermark-aligned continuous batching. |
| Training | FP8 trainer (native float8 + STE, E5M2 gradient hooks, AdamW master weights), DualPipe schedule simulation (recompute-based, gradient-exact), GRPO (real sampling, k3 KL, answer-extraction rewards), Muon optimizer (NewtonβSchulz orthogonalized momentum, optional per-head blocks), QAT straight-through fake-quant training. |
| Adaptation | AdaptationLoop: transferred harnesses must be re-accepted under the target workload's gate or dropped; cold/prefix-free targets shed draft+pool, matching the v5.16 break-even data |
| Harness evolution | inference/harness_evolver.py: ModularRSI-style module-wise search over draft/pool/tier configs; the temp-0 bitwise gate makes 'latency evolves, answers never change' an enforced invariant (1.41x modeled speedup, 0 gate rejects) |
| Prefix pool | inference/prefix_pool.py: blake2b(token-block + config fingerprint) keyed KV snapshots, LRU, exact-length past; pooled greedy == from-scratch bitwise |
| Stream scoring | eval/score_stream.py: score any engine's (prompt, output) JSONL under a reference model; A/B compare with bootstrap CI β the audit-side complement to serving engines |
| Spec telemetry | bench_spec_breakeven.py: draft x cache-state sweep in the colibri-P3 schema; acceptance + expert hit-rate per decode context, best_draft_per_cache_state() picker |
| Expert streaming | expert_store.py: routed experts tiered to a memory-mapped file, LRU residency with hit/miss/eviction telemetry; streaming forward is bitwise-identical to dense (oracle-verified, roadmap #6) |
| Toy checkpoints | checkpoints/toy_v5.13.pt β 8.5M char-level model trained on the repo's own source in ~10 CPU-minutes (examples/train_toy_checkpoint.py); checkpoints/tool_tuned_v5.27.pt β tool-tuned on agent-loop replays (examples/train_tool_tuned.py, periodic save + resume); generate() / harness / MTP / agent loop run against trained weights |
| Quantization | True GPTQ (Hessian OBS with error compensation, optional act-order), AWQ with activation-aware grid search, native FP8, MXFP4 β all with from_linear real-weight packing. QAT fake-quant training (STE) targets: MXFP4, AWQ, and NVFP4 (E2M1Γ16 + E4M3 block scales, local-line v5.22-L). |
| Eval | Log-likelihood harness (helioslm_v5/eval/harness.py): loglikelihood / multiple_choice / run_harness + built-in synthetic tasks (v5.11), token-id based, lm-eval-harness spirit |
| Multimodal | NaViT vision encoder (row/col position decomposition, mixed-resolution packing), streaming audio encoder (causal, sliding-window memory, bit-equivalent to one-shot). |
Why HeliosLM vs. alternatives?
| HeliosLM | transformers |
vLLM |
nanoGPT-style | |
|---|---|---|---|---|
| Purpose | Understand + hack the full stack | Run pre-trained models | Max serving throughput | Learn basics |
| Runs on CPU end-to-end | β | partial | β | β |
| Training + serving + quantization in one repo | β | β | β | β |
| Bitwise/strict correctness checks on core paths | β | β | β | β |
| Production throughput | β (by design) | β | β | β |
Benchmarks
The full config holds 99.2% less KV-cache memory than MHA at 128k context
(2.0 GB vs 257.7 GB): 36 of 48 layers are Gated-Delta linear attention with a
fixed ~1 MB recurrent state, and the 12 MLA layers store only the compressed
latent (512 + 64 values/token). Full numbers and methodology:
docs/BENCHMARKS.md β reproducible on CPU via
python benchmarks/bench_cpu.py.
Sparse Γ MTP acceptance sweep (benchmarks/bench_sparse_mtp.py):
extends the v5.16 break-even sweep with the missing sparse_top_k axis β
MTP acceptance, CPU tok/s, and output overlap vs dense under DSA-style
top-k decode; k β₯ L row is bitwise-equal to dense (test-enforced oracle).
indexer="learned" rows need a checkpoint with indexer weights β
examples/train_indexer_distill.py distills one from the toy checkpoint
(DSA recipe: teacher = the head-mean free indexer, trunk frozen
bit-exact).
v5.28 serving Pareto (benchmarks/disagg_pareto_2026-09-25.json, regenerate via
python examples/disagg_pareto.py): latency-cost curves for cache-heavy / cold / mixed
workloads β the cost-axis alignment artifact, see
docs/benchmark_alignment.md.
Repository layout
helioslm_v5/ # source (configs, src/{attention,moe,inference,training,vision,audio,quantization}, agent/, tests)
examples/ # train_toy_checkpoint.py, train_tool_tuned.py (v5.27), disagg_pareto.py (v5.28)
benchmarks/ # CPU bench results + disagg_pareto_2026-09-25.json (v5.28 artifact)
docs/ # code review reports, benchmark_alignment.md, k3_alignment_targets(.md/.csv),
# competitive_intel_2026-09-25.md (+ raw claims CSV), helioslm_handoff.md
integration_test_v51.py
CHANGELOG.md # full version history (v5.0 β v5.28)
Roadmap
See the GitHub Project board for the live plan. Highlights:
- Pre-trained toy checkpoints (v5.13 text, v5.27 tool-tuned) β
loadandgenerate()/ agent-loop immediately - Epoch-2 + scaled tool-tuning (T17 parse 0.15 β target 0.4+; trainer resume-ready)
- Fused quantization kernels
- CUDA end-to-end verification (paths are currently static-checked; CPU-verified)
- Hugging Face Hub:
chienhsinlin/helioslm-agenthosts the agent layer + tool-tuned artifacts - Example notebooks: "Train a tiny HeliosLM on your laptop" / "Add a new attention variant in 30 lines"
Contributing
Contributions are very welcome β see CONTRIBUTING.md. Issues labeled good first issue are the best entry points.
Known Limitations
All five previously known limitations are resolved as of v5.14. Remaining hardware-dependent item: CUDA end-to-end verification (CPU-verified paths are static-checked) β tracked as a community issue.
Version history
Headlines (full details in CHANGELOG.md):
- v5.22-L (local line) β NVFP4-format QAT fake-quant target (E2M1Γ16 + E4M3 block scales) + NoPE option for MLA (Kimi-K3 direction; default off, bit-identical)
- v5.21 (local line) β KDA-style per-channel decay gate: linear-attention state rows forget at independent rates (Kimi Linear / GLM-5.3-Flash direction), bit-identical default, rollback-safe
- v5.29 β Three-mode benchmark on the real checkpoint (direct/routed/oracle): tau-routing curve strictly monotone (v5.22 gate PASS on real confidence); headline finding = systematic overconfidence on wrong answers (conf 0.944) β the exact failure class the v5.22 audit toolkit measures
- v5.28 β Cost-axis alignment: disagg three-axis Pareto sweep (latency-cost curves per workload; cache-aware anti-monotonicity recorded as structural finding)
- v5.27 β Tool-tuned checkpoint: agent-loop-replay training, T17 end-to-end (parse 0.15/finish 1/9 baseline, regression-guard floors), sparse_top_k=4 ~= dense in agent inference; agent hardened (TOOL_ERROR recovery, ASCII-safe docs)
- v5.26 β Stage A real-model oracles (T15): zero-gate attention residuals bitwise-verified on HeliosLMv5; sparse top-k decode oracle (kβ₯L bit-identical, selection validity + determinism); agent-loop smoke on the real toy checkpoint
- v5.25 β Disagg evolver module: Mooncake-style prefill/decode separation as a HarnessEvolver search module, monotonicity gate, three-axis Pareto (makespan / workers / worker-seconds)
- v5.24 β Attention variants with two-layer oracles: DSA sparse decode (fp32 certificate β fp64 gate), AttnRes mixing (zero-init β bitwise migration gate)
- v5.23 β Agent layer: strict tool-call schema + parser, deterministic sandboxed tools, trajectory bitwise-replay oracle, ground-truth-by-construction envs, routing gates in the agent loop
- v5.22 β Decision-layer audit toolkit: calibration metrics (ECE / Brier) against constructed ground truth β no reference LLM required
- v5.20 β Pareto-aware integration (memory axis) + cross-workload adaptation loop β ModularRSI gaps 2/3 closed at the inference layer
- v5.19 β Evolvable serving harness: ModularRSI-style module-wise search (draft/pool/tier) behind a deterministic bitwise oracle gate
- v5.18 β Content-addressed KV prefix pool: cross-session prefix reuse, fingerprint-guarded, bit-exact oracle (roadmap #6 complete)
- v5.17 β Standalone token-stream scorer: quality-gate any engine's output (A/B compare + bootstrap CI, colibri-container-style)
- v5.16 β Speculation break-even sweep: colibri-P3-compatible JSONL schema, first honest data point (draft pays only when warm)
- v5.15 β Disk-tier expert store: bit-exact streaming oracle, mmap + LRU, router/shared stay resident (roadmap #6)
- v5.14 β Multi-process DualPipe (one stage per process, phased queue protocol, gradient-exact vs single-process)
- v5.13 β CPU-trained toy char-level checkpoint (MTP aux loss, acceptance 1.00 on greedy) + train_toy_checkpoint.py
- v5.12 β Batched equal-length prefill in the engine (hybrid recurrent-state models included)
- v5.11 β Fused dequantΓmatmul kernels (AWQ/GPTQ/MXFP4), MXFP4 decode fix, eval loglikelihood harness
- v5.10 β CPU benchmark suite: analytic KV-cache accounting + wall-clock generation, BENCHMARKS.md
- v5.9 β Gemma-style attention + final logit soft-capping, per-head QK-norm, sliding-window attention with StreamingLLM sinks
- v5.8 β YaRN RoPE scaling, DSA-style sparse top-k decode, per-head Muon, GPTQ act-order
- v5.7 β RoPE scaling (linear/NTK), FP8 latent KV cache, Hyper-Connections, QAT training
- v5.6 β Hybrid packed-sequence training, MXFP4, Muon optimizer
- v5.5 β K3-aligned feature wave: hybrid Gated-Delta attention, LatentMoE, quantile balancing, attention residuals, SiTU-GLU
- v5.0βv5.4 β Review-hardened core: MLA weight absorption, true GPTQ, strict speculative sampling, NaViT + audio encoders
License
Apache License 2.0 β see LICENSE.

