datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EgoIT-99KCheckout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information.
mmaulibrispeechvoicebench
License
The dataset is available under the Apache 2.0 license.
Citation
If you use the VoiceBench dataset in your research, please cite the following paper:
@article{chen2024voicebench,
title={VoiceBench: Benchmarking LLM-Based Voice Assistants},
author={Chen, Yiming and Yue, Xianghu and Zhang, Chen and Gao, Xiaoxue and Tan, Robby T. and Li, Haizhou},
journal={arXiv preprint arXiv:2410.17196},
year={2024}
}
WenetSpeech_tempWorldSensePerceptionTest_ValgigaspeechClothoAQAAIR_benchcommon_voice_15Omni_Bench_fixPerceptionTestcovost2_en-zhmuchomusicDataset Summary
MuChoMusic is a benchmark designed to evaluate music understanding in multimodal audio-language models (Audio LLMs). The dataset comprises 1,187 multiple-choice questions created from 644 music tracks, sourced from two publicly available music datasets: MusicCaps and the Song Describer Dataset (SDD). The questions test knowledge and reasoning abilities across dimensions such as music theory, cultural context, and functional applications. All questions and answers have been… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-audio/muchomusic.vocalsoundpeoples_speechfleursOmni_BenchAV_Odyssey_Bench_LMMs_EvalLibrispeech-concattedliumsong-describerMixEval-X-audio2textcovost2europal-asralpaca_audiococo-caption2017-lmms-qa-ttsopenhermes_instruction
