WorldSonus: Bringing Sound to Worlds
Inference weights for the final C3 no-H0, 150k EMA model.
Code and usage: NoizAI/WorldSonus.
Contents
worldsonus_150k.pt: 638 inference tensors. No optimizer, training arguments, inactive history adapters, or training-only alignment projector.audio_codec.pt: causal stereo 48 kHz decoder, 64 latent channels at 30 Hz.z_stats.pt: matching latent normalization statistics. Do not substitute statistics from another model or recompute them for these weights.
All files load with torch.load(..., weights_only=True). DINOv3 and T5Gemma 2
encoders are downloaded separately under their upstream licenses.
Usage
Install the inference repository, then run:
python scripts/download_model.py --repo FF2416/WorldSonus --output assets
python scripts/download_encoders.py --output assets
Generate audio directly from a video and prompt:
python -m worldsonus.infer --video input.mp4 \
--prompt "A train passes beside a river." --output output.wav
Add --stream for incremental video encoding, audio generation, and causal
decoding in 100 ms chunks. See the code repository for playback instructions.
Limitations and responsible use
Generated audio may not match every visual event and may contain artifacts. Do not present generated audio as authentic recorded evidence. Respect source media rights and avoid recreating identifiable voices without consent.
This is a research inference release, not a fully qualified live-streaming deployment. Training code and datasets are not included.
License
WorldSonus contributions and these released model weights are provided under CC BY-NC 4.0 for non-commercial use with attribution. Upstream third-party components retain their own terms; see LICENSE and THIRD_PARTY_NOTICES.md.