Haan v1 โ€” full-training lineage

Haan is a Korean-first full-duplex speech model: a Qwen3-8B backbone reading Moshi-style multi-stream frames (Mimi codec, 12.5 Hz, 8 codebooks per audio stream) across four channels โ€” native (inner) text, surface text, self audio, user audio. This repository preserves the phase-wise weights of the full-training lineage, in which every backbone parameter learns.

All checkpoints are bf16, from_pretrained-loadable (HaanForConditionalGeneration), 36 backbone layers, and descend from the same scale-matched assembly (Qwen3-8B + kyutai/moshiko grafts, 2026-08-09; see init_report.json inside each checkpoint).

Layout

path training what it is
phase0/ text-anchor corpus only frame format and the native channel's block grammar
phase1/ English TTS only depth-adapter initialization gate
phase2/ TTS + ASR + dialogue mix, step 1200 speaking, listening, timing; the lineage's reference checkpoint

Phase-2 selection

Step 1200 is the checkpoint later work builds on (the frozen-backbone lineage's compose base). The run continued to step 1700, but scored worse on free-running TTS (WER 2.48 vs 1.79) with duplex metrics flat, so 1200 is the one preserved.

Known state at step 1200, measured by self-play milestones: teacher-forced perplexities converge; free-running ASR does not transcribe (WER ~1.0) and turn overlap stays ~0.99. These are open problems of the lineage, documented here so the checkpoint is read for what it is.

Sibling repository

RetentionLabs/haan-v1-frozen holds the frozen-backbone lineage: pristine Qwen3/Moshi spans restored and held, adapter blocks added at the ends, resumed from this repository's phase2 fitted parts.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for RetentionLabs/haan-v1-full

Finetuned
Qwen/Qwen3-8B
Finetuned
(2145)
this model