WorldSonus: Bringing Sound to Worlds

Inference weights for the final C3 no-H0, 150k EMA model.

Code and usage: NoizAI/WorldSonus.

Contents

  • worldsonus_150k.pt: 638 inference tensors. No optimizer, training arguments, inactive history adapters, or training-only alignment projector.
  • audio_codec.pt: causal stereo 48 kHz decoder, 64 latent channels at 30 Hz.
  • z_stats.pt: matching latent normalization statistics. Do not substitute statistics from another model or recompute them for these weights.

All files load with torch.load(..., weights_only=True). DINOv3 and T5Gemma 2 encoders are downloaded separately under their upstream licenses.

Usage

Install the inference repository, then run:

python scripts/download_model.py --repo FF2416/WorldSonus --output assets
python scripts/download_encoders.py --output assets

Generate audio directly from a video and prompt:

python -m worldsonus.infer --video input.mp4 \
  --prompt "A train passes beside a river." --output output.wav

Add --stream for incremental video encoding, audio generation, and causal decoding in 100 ms chunks. See the code repository for playback instructions.

Limitations and responsible use

Generated audio may not match every visual event and may contain artifacts. Do not present generated audio as authentic recorded evidence. Respect source media rights and avoid recreating identifiable voices without consent.

This is a research inference release, not a fully qualified live-streaming deployment. Training code and datasets are not included.

License

WorldSonus contributions and these released model weights are provided under CC BY-NC 4.0 for non-commercial use with attribution. Upstream third-party components retain their own terms; see LICENSE and THIRD_PARTY_NOTICES.md.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support