Loie/SpotSound-Bench
SpotSound-Bench: A 'Needle-in-a-Haystack' Evaluation for Audio Temporal Grounding Benchmark Summary SpotSound-Bench is a challenging temporal grounding benchmark designed to evaluate Large Audio-Language Models (ALMs). Existing benchmarks for audio temporal grounding often feature high ratios of target-window duration to full audio clip duration, which fail to simulate real-world scenarios where short events are obscured by dense background sounds. To bridge… See the full description on the dataset page: https://huggingface.co/datasets/Loie/SpotSound-Bench.
SpotSound-Bench: A 'Needle-in-a-Haystack' Evaluation for Audio Temporal Grounding
  
Benchmark Summary
SpotSound-Bench is a challenging temporal grounding benchmark designed to evaluate Large Audio-Language Models (ALMs).
Existing benchmarks for audio temporal grounding often feature high ratios of target-window duration to full audio clip duration, which fail to simulate real-world scenarios where short events are obscured by dense background sounds. To bridge this gap, we introduce SpotSound-Bench, featuring short acoustic events embedded within long, unstructured recordings.
This benchmark creates a rigorous ‘needle-in-a-haystack’ evaluation, demanding high temporal precision and robust resistance against hallucinations from audio-language models.
Benchmark Characteristics
- Average Clip Length: 54.2 seconds
- Average Target Event Length: 3.9 seconds
- Temporal Density: 7.2% (Target event duration / Full audio clip duration)
- Challenge: A large search space dominated by background content, requiring models to pinpoint exact timestamps of short events while ignoring complex background ambiance and avoiding hallucinated predictions for non-existent events.
Data Structure
<pre> { "audiopath": "Uro9suV3xU130187.wav", "caption": "hair dryer drying", "annotations": [[8.1, 10.9]] }, </pre>
Citation
If you use this code and data for your research or project, please cite:
@inproceedings{sun2026spotsound, title={SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding}, author={Sun, Luoyi and Zhou, Xiao and Li, Zeqian and Zhang, Ya and Wang, Yanfeng and Xie, Weidi}, journal={arXiv preprint arXiv:2604.13023}, year={2026} }
Contact
For questions, please contact: loiesun411@gmail.com.
