Request access for non-commercial academic use

Access requires acceptance of the Matterport3D Terms of Use.

By requesting access, you confirm that your use is limited to non-commercial academic purposes and that you have read and agree to the Matterport3D Terms of Use: https://kaldir.vc.in.tum.de/matterport/MP_TOS.pdf

Log in or Sign Up to review the conditions and access this model content.

ST-AudioLM

This repository contains both released components from Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources:

  • ST-Audio Encoder, a standalone FOA encoder that produces 40 temporal tokens, one semantic token, and time-resolved direction, distance, and activity predictions; and
  • ST-AudioLM, which connects those 41 tokens to OLMo-2 for semantic, spatial, and temporal audio question answering.
ST-AudioLM/
β”œβ”€β”€ encoder/
β”‚   β”œβ”€β”€ config.json
β”‚   └── encoder.safetensors
└── model/
    β”œβ”€β”€ config.json
    β”œβ”€β”€ connector.safetensors
    β”œβ”€β”€ adapter/
    └── tokenizer/

Install the ST-AudioLM code before loading either component. Request access on this page and run hf auth login before downloading the weights.

ST-Audio Encoder

from staudiolm import STAudioEncoder, load_foa

audio = load_foa("example.wav")
encoder = STAudioEncoder.from_pretrained("HBoh/ST-AudioLM", device="cuda")
outputs = encoder(audio)
tokens = outputs["audio_tokens"]  # [1, 41, 768]

Loading the Encoder does not load an LLM.

ST-AudioLM

from staudiolm import STAudioLM, load_foa

audio = load_foa("example.wav")
model = STAudioLM.from_pretrained("HBoh/ST-AudioLM", device="cuda")
answer = model.generate(audio, "How does the source move?")[0]

The base language model is OLMo-2-1124-7B-Instruct and is downloaded separately. Its weights are not duplicated here.

Input and limitations

Input must be a 10-second, 32-kHz, four-channel AmbiX ACN/SN3D recording in [W, Y, Z, X] order. The models are intended for research on controlled spatial-audio understanding. Performance may degrade for other microphone formats, real recordings that differ from the simulated training environments, overlapping sources outside the training distribution, or clips with different duration and sample rate.

License

The released weights and metadata are provided under CC BY-NC-SA 3.0 US and are also subject to the Matterport3D Terms of Use where applicable. OLMo-2 and all other third-party components retain their original licenses.

Acknowledgements

ST-AudioLM builds on BAT and Spatial-AST, as well as OLMo-2, Transformers, and PEFT.

Citation

@article{hyun2026spatio,
  title={Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources},
  author={Hyun-Bin, Oh and Shimada, Kazuki and Takida, Yuhta and Sung-Bin, Kim and Uesaka, Toshimitsu and Shibuya, Takashi and Lee, Kyeongyoon and Oh, Tae-Hyun and Mitsufuji, Yuki},
  journal={arXiv preprint arXiv:2606.14141},
  year={2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for HBoh/ST-AudioLM

Papers for HBoh/ST-AudioLM