Download code/README.md from laion/Humaneness-Voice-Small: direct link, hf CLI and curl.
- Browser
- Download file 2.48 kB
-
https://huggingface.co/laion/Humaneness-Voice-Small/resolve/main/code/README.md
- Command line
-
hf download hf://laion/Humaneness-Voice-Small/code/README.md
-
curl -L -o README.md https://huggingface.co/laion/Humaneness-Voice-Small/resolve/main/code/README.md
Code map
This directory includes the exact historical Python training implementation and the source-level components needed to load the released model. Historical scripts retain absolute JUPITER paths and strict manifest/code hash guards; those paths are provenance, not runnable defaults on another machine. For a new dataset, regenerate manifests and plans and rehash changed source rather than editing the original plan in place.
continuous_train.py,train.sbatch,plan_caption_production_32nodes.json: one 128-GPU, single-warmup/cosine S1→S10 run.large_talker.pyspecifies the architecture.state_io.pywrites resumable full states and BF16 model-only exports.caption_curriculum_dataset.py,cascade_dataset.py,build_caption_assets.py,build_caption_stage.py,make_caption_training_plans.py,submit_caption_pipeline.py: data ingestion and multi-format presentation construction. Dataset audio and annotations are external, container-backed sources, not bundled into the model repo.packing.py,conditioning.py,pack_moss.py,moss_pack.py,prompt_fmt.py,timed_prompt.py,anno_timed.py: prompt/reference packing and loss masking. The score-conditioned modules are carried for checkpoint compatibility; score-token prompts were not used in this caption run.infer.py: example for one BF16 checkpoint plus the separately downloaded codec. The model construction and strict state-dictionary load were verified on CPU; generation requires CUDA and should be canaried in the target environment. It is not an HFpipeline()implementation.eval_validation_loss.py,validation_sweep.sbatch,acting_challenge_eval.py,acting_challenge_eval.sbatch: historical evaluation implementation. Public benchmark summary and audio live in the linked dataset/Space.assets/sft3(one directory above): tokenizer/config and MOSS architecture modules used to instantiate this derived architecture. Original 4B SFT-3 weights are not bundled or needed.assets/qwen3/config.jsonis the backbone configuration; the trained weights are contained in each published model state.
Requirements for inference: a recent compatible Python/PyTorch/Transformers stack, NumPy, SoundFile, Torchaudio, Safetensors and access to OpenMOSS-Team/MOSS-Audio-Tokenizer-v2. The historical JUPITER environment used PyTorch 2.9.1 and its matching Torchaudio build. Pin a tested software environment for production; package APIs may change.