CoolFace
Datasetpublic

Loie/SpotSound-Bench

SpotSound-Bench: A 'Needle-in-a-Haystack' Evaluation for Audio Temporal Grounding Benchmark Summary SpotSound-Bench is a challenging temporal grounding benchmark designed to evaluate Large Audio-Language Models (ALMs). Existing benchmarks for audio temporal grounding often feature high ratios of target-window duration to full audio clip duration, which fail to simulate real-world scenarios where short events are obscured by dense background sounds. To bridge… See the full description on the dataset page: https://huggingface.co/datasets/Loie/SpotSound-Bench.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
1likes341downloads
Dataset Card

SpotSound-Bench: A 'Needle-in-a-Haystack' Evaluation for Audio Temporal Grounding

![Project Page](https://loiesun.github.io/spotsound/) ![GitHub](https://github.com/LoieSun/spotsound) ![Paper](https://arxiv.org/abs/2604.13023)

Benchmark Summary

SpotSound-Bench is a challenging temporal grounding benchmark designed to evaluate Large Audio-Language Models (ALMs).

Existing benchmarks for audio temporal grounding often feature high ratios of target-window duration to full audio clip duration, which fail to simulate real-world scenarios where short events are obscured by dense background sounds. To bridge this gap, we introduce SpotSound-Bench, featuring short acoustic events embedded within long, unstructured recordings.

This benchmark creates a rigorous ‘needle-in-a-haystack’ evaluation, demanding high temporal precision and robust resistance against hallucinations from audio-language models.

Benchmark Characteristics

  • Average Clip Length: 54.2 seconds
  • Average Target Event Length: 3.9 seconds
  • Temporal Density: 7.2% (Target event duration / Full audio clip duration)
  • Challenge: A large search space dominated by background content, requiring models to pinpoint exact timestamps of short events while ignoring complex background ambiance and avoiding hallucinated predictions for non-existent events.

Data Structure

<pre> { "audiopath": "Uro9suV3xU130187.wav", "caption": "hair dryer drying", "annotations": [[8.1, 10.9]] }, </pre>

Citation

If you use this code and data for your research or project, please cite:

@inproceedings{sun2026spotsound, title={SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding}, author={Sun, Luoyi and Zhou, Xiao and Li, Zeqian and Zhang, Ya and Wang, Yanfeng and Xie, Weidi}, journal={arXiv preprint arXiv:2604.13023}, year={2026} }

Contact

For questions, please contact: loiesun411@gmail.com.