Text-to-Speech
English
German
voice-acting
qwen3
moss-audio-tokenizer-v2
audio-generation
ChristophSchuhmann's picture
Document architecture, prompts, code, and full run statistics
5de8603 verified
|
Raw History Blame Contribute Delete
2.48 kB

Code map

This directory includes the exact historical Python training implementation and the source-level components needed to load the released model. Historical scripts retain absolute JUPITER paths and strict manifest/code hash guards; those paths are provenance, not runnable defaults on another machine. For a new dataset, regenerate manifests and plans and rehash changed source rather than editing the original plan in place.

  • continuous_train.py, train.sbatch, plan_caption_production_32nodes.json: one 128-GPU, single-warmup/cosine S1→S10 run. large_talker.py specifies the architecture. state_io.py writes resumable full states and BF16 model-only exports.
  • caption_curriculum_dataset.py, cascade_dataset.py, build_caption_assets.py, build_caption_stage.py, make_caption_training_plans.py, submit_caption_pipeline.py: data ingestion and multi-format presentation construction. Dataset audio and annotations are external, container-backed sources, not bundled into the model repo.
  • packing.py, conditioning.py, pack_moss.py, moss_pack.py, prompt_fmt.py, timed_prompt.py, anno_timed.py: prompt/reference packing and loss masking. The score-conditioned modules are carried for checkpoint compatibility; score-token prompts were not used in this caption run.
  • infer.py: example for one BF16 checkpoint plus the separately downloaded codec. The model construction and strict state-dictionary load were verified on CPU; generation requires CUDA and should be canaried in the target environment. It is not an HF pipeline() implementation.
  • eval_validation_loss.py, validation_sweep.sbatch, acting_challenge_eval.py, acting_challenge_eval.sbatch: historical evaluation implementation. Public benchmark summary and audio live in the linked dataset/Space.
  • assets/sft3 (one directory above): tokenizer/config and MOSS architecture modules used to instantiate this derived architecture. Original 4B SFT-3 weights are not bundled or needed. assets/qwen3/config.json is the backbone configuration; the trained weights are contained in each published model state.

Requirements for inference: a recent compatible Python/PyTorch/Transformers stack, NumPy, SoundFile, Torchaudio, Safetensors and access to OpenMOSS-Team/MOSS-Audio-Tokenizer-v2. The historical JUPITER environment used PyTorch 2.9.1 and its matching Torchaudio build. Pin a tested software environment for production; package APIs may change.