Skip to content

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

Spatio-Temporal Audio Language Modeling
for Dynamic Sound Sources

EMNLP 2026 Main

Oh Hyun-Bin1,*, Kazuki Shimada2, Yuhta Takida2, Kim Sung-Bin1, Toshimitsu Uesaka2, Takashi Shibuya2,
Kyeongyoon Lee4, Tae-Hyun Oh5,†, Yuki Mitsufuji2,3,†

1POSTECH   2Sony AI   3Sony Group Corporation
4Sungkyunkwan University   5KAIST
*Work done during an internship at Sony AI.   †Co-corresponding authors.

Project Page arXiv ST-Audio Encoder and ST-AudioLM weights ST-AudioQA dataset

Overview of ST-Audio Encoder, ST-AudioLM, and ST-AudioQA

ST-Audio Encoder represents what is sounding and how it moves with one semantic token and 40 time-resolved tokens.
ST-AudioLM connects these tokens to a language model for semantic, spatial, and temporal reasoning over dynamic sound sources.

Overview

This repository contains two independently usable inference paths:

  • ST-Audio Encoder takes a 10-second first-order ambisonic recording and returns 40 temporal embeddings, one semantic embedding, and time-resolved direction, distance, and activity predictions. It can be used without loading an LLM.
  • ST-AudioLM loads the bundled Encoder, projects its 41 audio tokens into OLMo-2-1124-7B-Instruct, and answers natural-language questions about sound identity, position, and motion.

Both components are released in one Hugging Face model repository. The Encoder remains independently loadable without the LLM.

ST-AudioQA is the accompanying benchmark. It provides QA annotations, AudioSet segment references, acoustic-scene recipes, and the exact RIRs used for the test split. Original AudioSet waveforms and Matterport3D scene assets are not redistributed.

Getting Started

1. Install

Use Python 3.10+ and install a PyTorch build suitable for your hardware. Then run:

git clone https://github.com/SonyResearch/ST-AudioLM.git
cd ST-AudioLM
python -m pip install -e .

The model and dataset repositories are gated. Accept their access terms on Hugging Face, then authenticate once with hf auth login before downloading weights or data.

Verify the standalone Encoder without preparing an audio file:

python scripts/smoke_test.py --encoder-model HBoh/ST-AudioLM --device cpu

The full model has been tested with bfloat16 inference on an NVIDIA H100. The Encoder can also run on CPU, although it is slower.

2. Prepare FOA Audio

Pass your own 10-second, 32-kHz, four-channel AmbiX ACN/SN3D WAV, FLAC, or NPY recording in standard [W, Y, Z, X] channel order.

The loader validates the format but deliberately does not resample, crop, or peak-normalize audio. NPY arrays may have shape [4, 320000] or [320000, 4].

3. Run ST-Audio Encoder by Itself

Command line:

python scripts/encode.py example.wav --output example_features.npz

Python:

from staudiolm import STAudioEncoder, load_foa

audio = load_foa("example.wav")
encoder = STAudioEncoder.from_pretrained(
    "HBoh/ST-AudioLM",
    device="cuda",
)

outputs = encoder(audio)
tokens = outputs["audio_tokens"]  # [1, 41, 768]: 40 temporal + 1 semantic

direction = outputs["direction"]          # [1, 40, 3]
distance = outputs["distance_log"].exp()  # [1, 40]
activity = outputs["activity_logits"].sigmoid()

The CLI saves these outputs and the encoder tokens in a NumPy .npz file, making the Encoder usable as a standalone feature extractor or trajectory estimator.

4. Run ST-AudioLM

Command line:

python scripts/infer.py example.wav \
  "Where is the siren at the end of the clip?"

Python:

from staudiolm import STAudioLM, load_foa

audio = load_foa("example.wav")
model = STAudioLM.from_pretrained("HBoh/ST-AudioLM", device="cuda")
answer = model.generate(audio, "Where is the siren at the end of the clip?")[0]
print(answer)

The model repository contains both ST-Audio Encoder and ST-AudioLM. OLMo-2 is downloaded separately and is not duplicated in the release. Local directories are supported for offline use:

python scripts/infer.py example.wav "How does the source move?" \
  --model ./weights/ST-AudioLM \
  --base-model ./weights/olmo2 \
  --local-files-only

ST-AudioQA

Install the data dependencies first:

python -m pip install -e ".[data]"

Load each public table independently:

from datasets import load_dataset

qa = load_dataset("HBoh/ST-AudioQA", "qa")
recordings = load_dataset("HBoh/ST-AudioQA", "recordings")
sources = load_dataset("HBoh/ST-AudioQA", "sources")
rir_recipes = load_dataset("HBoh/ST-AudioQA", "rir_recipes")
rir_test_index = load_dataset("HBoh/ST-AudioQA", "rir_test_index")

The release contains:

  • paper-aligned train, validation, and test QA splits;
  • stable recording and source IDs with AudioSet YouTube IDs and time segments;
  • the complete Matterport3D/Habitat-Sim geometry and trajectory recipes;
  • exact, deduplicated test RIRs for benchmark reproduction; and
  • checksums and provenance metadata.

Users acquire the referenced source audio and Matterport3D assets from their original providers. Train and validation RIRs can then be rendered from the recipes; test evaluation uses the released exact RIRs because Habitat-Sim re-rendering is not guaranteed to be bit-exact across versions. Reconstructed model input is 10-second, 32-kHz AmbiX ACN/SN3D audio in [W, Y, Z, X] order. See ST-AudioQA data and reconstruction.

Download only the shard containing one exact test RIR and extract it with:

python scripts/extract_test_rir.py RIR_ID --output rirs/RIR_ID.npz

Documentation

Acknowledgements

ST-AudioLM builds on BAT and its Spatial-AST encoder, as well as OLMo-2, Transformers, and PEFT.

License

Original source code and documentation are released under the MIT License. Portions adapted from third-party projects remain subject to their original terms; copies are collected in third_party_licenses/. Model weights and ST-AudioQA assets have separate terms described in their Hugging Face cards.

Citation

@article{hyun2026spatio,
  title={Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources},
  author={Hyun-Bin, Oh and Shimada, Kazuki and Takida, Yuhta and Sung-Bin, Kim and Uesaka, Toshimitsu and Shibuya, Takashi and Lee, Kyeongyoon and Oh, Tae-Hyun and Mitsufuji, Yuki},
  journal={arXiv preprint arXiv:2606.14141},
  year={2026}
}

Contact

hyunbinoh@postech.ac.kr or hyunbin70@gmail.com

About

[EMNLP 2026] Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages