Safetensors
PEFT
st-audiolm
audio-language-model
audio-feature-extraction
spatial-audio
ambisonics
sound-event-localization
Instructions to use HBoh/ST-AudioLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use HBoh/ST-AudioLM with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download README.md from HBoh/ST-AudioLM: direct link, hf CLI and curl.
- Browser
- Download file 4.07 kB
-
https://huggingface.co/HBoh/ST-AudioLM/resolve/main/README.md
- Command line
-
hf download hf://HBoh/ST-AudioLM/README.md
-
curl -L -H "Authorization: Bearer $HF_TOKEN" -o README.md https://huggingface.co/HBoh/ST-AudioLM/resolve/main/README.md
4.07 kB
| license: cc-by-nc-sa-3.0 | |
| library_name: st-audiolm | |
| base_model: allenai/OLMo-2-1124-7B-Instruct | |
| tags: | |
| - audio-language-model | |
| - audio-feature-extraction | |
| - spatial-audio | |
| - ambisonics | |
| - sound-event-localization | |
| - peft | |
| extra_gated_heading: Request access for non-commercial academic use | |
| extra_gated_description: Access requires acceptance of the Matterport3D Terms of Use. | |
| extra_gated_prompt: "By requesting access, you confirm that your use is limited to non-commercial academic purposes and that you have read and agree to the Matterport3D Terms of Use: https://kaldir.vc.in.tum.de/matterport/MP_TOS.pdf" | |
| extra_gated_fields: | |
| Academic institution: text | |
| Principal investigator name: text | |
| Principal investigator email: text | |
| I agree to use this repository for non-commercial academic purposes only: checkbox | |
| I have read and agree to the Matterport3D Terms of Use: checkbox | |
| I consent to the collection of my name, email address, and academic institution, and to sharing this information with Matterport as described in the Matterport3D Terms of Use: checkbox | |
| # ST-AudioLM | |
| This repository contains both released components from [Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources](https://arxiv.org/abs/2606.14141): | |
| - **ST-Audio Encoder**, a standalone FOA encoder that produces 40 temporal tokens, one semantic token, and time-resolved direction, distance, and activity predictions; and | |
| - **ST-AudioLM**, which connects those 41 tokens to OLMo-2 for semantic, spatial, and temporal audio question answering. | |
| ```text | |
| ST-AudioLM/ | |
| βββ encoder/ | |
| β βββ config.json | |
| β βββ encoder.safetensors | |
| βββ model/ | |
| βββ config.json | |
| βββ connector.safetensors | |
| βββ adapter/ | |
| βββ tokenizer/ | |
| ``` | |
| Install the [ST-AudioLM code](https://github.com/SonyResearch/ST-AudioLM) before loading either component. | |
| Request access on this page and run `hf auth login` before downloading the | |
| weights. | |
| ## ST-Audio Encoder | |
| ```python | |
| from staudiolm import STAudioEncoder, load_foa | |
| audio = load_foa("example.wav") | |
| encoder = STAudioEncoder.from_pretrained("HBoh/ST-AudioLM", device="cuda") | |
| outputs = encoder(audio) | |
| tokens = outputs["audio_tokens"] # [1, 41, 768] | |
| ``` | |
| Loading the Encoder does not load an LLM. | |
| ## ST-AudioLM | |
| ```python | |
| from staudiolm import STAudioLM, load_foa | |
| audio = load_foa("example.wav") | |
| model = STAudioLM.from_pretrained("HBoh/ST-AudioLM", device="cuda") | |
| answer = model.generate(audio, "How does the source move?")[0] | |
| ``` | |
| The base language model is [OLMo-2-1124-7B-Instruct](https://huggingface.co/allenai/OLMo-2-1124-7B-Instruct) and is downloaded separately. Its weights are not duplicated here. | |
| ## Input and limitations | |
| Input must be a 10-second, 32-kHz, four-channel AmbiX ACN/SN3D recording in `[W, Y, Z, X]` order. The models are intended for research on controlled spatial-audio understanding. Performance may degrade for other microphone formats, real recordings that differ from the simulated training environments, overlapping sources outside the training distribution, or clips with different duration and sample rate. | |
| ## License | |
| The released weights and metadata are provided under [CC BY-NC-SA 3.0 US](LICENSE.md) and are also subject to the Matterport3D Terms of Use where applicable. OLMo-2 and all other third-party components retain their original licenses. | |
| ## Acknowledgements | |
| ST-AudioLM builds on [BAT](https://arxiv.org/abs/2402.01591) and [Spatial-AST](https://github.com/zszheng147/Spatial-AST), as well as [OLMo-2](https://huggingface.co/allenai/OLMo-2-1124-7B-Instruct), [Transformers](https://github.com/huggingface/transformers), and [PEFT](https://github.com/huggingface/peft). | |
| ## Citation | |
| ```bibtex | |
| @article{hyun2026spatio, | |
| title={Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources}, | |
| author={Hyun-Bin, Oh and Shimada, Kazuki and Takida, Yuhta and Sung-Bin, Kim and Uesaka, Toshimitsu and Shibuya, Takashi and Lee, Kyeongyoon and Oh, Tae-Hyun and Mitsufuji, Yuki}, | |
| journal={arXiv preprint arXiv:2606.14141}, | |
| year={2026} | |
| } | |
| ``` | |