Haan v1 โ full-training lineage
Haan is a Korean-first full-duplex speech model: a Qwen3-8B backbone reading Moshi-style multi-stream frames (Mimi codec, 12.5 Hz, 8 codebooks per audio stream) across four channels โ native (inner) text, surface text, self audio, user audio. This repository preserves the phase-wise weights of the full-training lineage, in which every backbone parameter learns.
All checkpoints are bf16, from_pretrained-loadable
(HaanForConditionalGeneration), 36 backbone layers, and descend from the
same scale-matched assembly (Qwen3-8B + kyutai/moshiko grafts, 2026-08-09;
see init_report.json inside each checkpoint).
Layout
| path | training | what it is |
|---|---|---|
phase0/ |
text-anchor corpus only | frame format and the native channel's block grammar |
phase1/ |
English TTS only | depth-adapter initialization gate |
phase2/ |
TTS + ASR + dialogue mix, step 1200 | speaking, listening, timing; the lineage's reference checkpoint |
Phase-2 selection
Step 1200 is the checkpoint later work builds on (the frozen-backbone lineage's compose base). The run continued to step 1700, but scored worse on free-running TTS (WER 2.48 vs 1.79) with duplex metrics flat, so 1200 is the one preserved.
Known state at step 1200, measured by self-play milestones: teacher-forced perplexities converge; free-running ASR does not transcribe (WER ~1.0) and turn overlap stays ~0.99. These are open problems of the lineage, documented here so the checkpoint is read for what it is.
Sibling repository
RetentionLabs/haan-v1-frozen holds the frozen-backbone lineage: pristine
Qwen3/Moshi spans restored and held, adapter blocks added at the ends,
resumed from this repository's phase2 fitted parts.