datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LanguageQA
Dataset Card for SAKURA-LanguageQA
This dataset contains the audio and the single/multi-hop questions/answers of the language track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information".
The fields of the dataset are:
file: The filename of the audio files.
audio: The audio recordings.
attribute_label: The attribute labels (i.e., the language spoken in the speech) of the audio… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/LanguageQA.Multi_Language_Audio2TextThis dataset by Mozilla Common Voice (https://commonvoice.mozilla.org/en/datasets) is crafted by Udyan Sachdev
Voice datasets play a pivotal role in training and evaluating speech-to-text models, influencing advancements in natural language processing. This dataset outlines the creation of a comprehensive text dataset from 40,571 MP3 audio files sourced from the Common Voice project. The dataset aims to serve as a benchmark for training and evaluating speech-to-text models in English, French… See the full description on the dataset page: https://huggingface.co/datasets/UdyanSachdev/Multi_Language_Audio2Text.
