Text-to-Speech
English
German
voice-acting
qwen3
moss-audio-tokenizer-v2
audio-generation
File size: 2,477 Bytes
5de8603
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
# Code map

This directory includes the exact historical Python training implementation and the source-level components needed to load the released model. Historical scripts retain absolute JUPITER paths and strict manifest/code hash guards; those paths are provenance, **not** runnable defaults on another machine. For a new dataset, regenerate manifests and plans and rehash changed source rather than editing the original plan in place.

- `continuous_train.py`, `train.sbatch`, `plan_caption_production_32nodes.json`: one 128-GPU, single-warmup/cosine S1→S10 run. `large_talker.py` specifies the architecture. `state_io.py` writes resumable full states and BF16 model-only exports.
- `caption_curriculum_dataset.py`, `cascade_dataset.py`, `build_caption_assets.py`, `build_caption_stage.py`, `make_caption_training_plans.py`, `submit_caption_pipeline.py`: data ingestion and multi-format presentation construction. Dataset audio and annotations are external, container-backed sources, not bundled into the model repo.
- `packing.py`, `conditioning.py`, `pack_moss.py`, `moss_pack.py`, `prompt_fmt.py`, `timed_prompt.py`, `anno_timed.py`: prompt/reference packing and loss masking. The score-conditioned modules are carried for checkpoint compatibility; score-token prompts were not used in this caption run.
- `infer.py`: example for one BF16 checkpoint plus the separately downloaded codec. The model construction and strict state-dictionary load were verified on CPU; generation requires CUDA and should be canaried in the target environment. It is not an HF `pipeline()` implementation.
- `eval_validation_loss.py`, `validation_sweep.sbatch`, `acting_challenge_eval.py`, `acting_challenge_eval.sbatch`: historical evaluation implementation. Public benchmark summary and audio live in the linked dataset/Space.
- `assets/sft3` (one directory above): tokenizer/config and MOSS architecture modules used to instantiate this derived architecture. Original 4B SFT-3 weights are **not** bundled or needed. `assets/qwen3/config.json` is the backbone configuration; the trained weights are contained in each published model state.

Requirements for inference: a recent compatible Python/PyTorch/Transformers stack, NumPy, SoundFile, Torchaudio, Safetensors and access to `OpenMOSS-Team/MOSS-Audio-Tokenizer-v2`. The historical JUPITER environment used PyTorch 2.9.1 and its matching Torchaudio build. Pin a tested software environment for production; package APIs may change.