datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AnimalQA
Dataset Card for SAKURA-AnimalQA
This dataset contains the audio and the single/multi-hop questions/answers of the animal track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information".
The fields of the dataset are:
file: The filename of the audio files.
audio: The audio recordings.
attribute_label: The attribute labels (i.e., the kinds of animal making the sounds) of the audio… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/AnimalQA.LanguageQA
Dataset Card for SAKURA-LanguageQA
This dataset contains the audio and the single/multi-hop questions/answers of the language track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information".
The fields of the dataset are:
file: The filename of the audio files.
audio: The audio recordings.
attribute_label: The attribute labels (i.e., the language spoken in the speech) of the audio… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/LanguageQA.GenderQA
Dataset Card for SAKURA-GenderQA
This dataset contains the audio and the single/multi-hop questions/answers of the gender track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information".
The fields of the dataset are:
file: The filename of the audio files.
audio: The audio recordings.
attribute_label: The attribute labels (i.e., gender of the speakers) of the audio files.… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/GenderQA.EmotionQA
Dataset Card for SAKURA-EmotionQA
This dataset contains the audio and the single/multi-hop questions/answers of the emotion track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information".
The fields of the dataset are:
file: The filename of the audio files.
audio: The audio recordings.
attribute_label: The attribute labels (i.e., the emotion of the speakers) of the audio files.… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/EmotionQA.
