datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DitingBench
Diting Benchmark
Our paperGithub
Our benchmark is designed to evaluate the speech comprehension capabilities of Speech LLMs. We tested both humans and Speech LLMs in terms of speech understanding and provided further analysis of the results, along with a comparative study between the two. This offers insights for the future development of Speech LLMs. For more details, please refer to our paper.
Result
Level
Task
Human Baseline
GPT-4o
MuLLaMA
GAMA
SALMONN… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/DitingBench.vox2_3D_ditto_combined_shard_03vox2_3D_ditto_shard_20vox2_3D_ditto_shard_21vox2_3D_ditto_shard_07vox2_3D_ditto_combined_shard_16vox2_3D_ditto_shard_50vox2_3D_ditto_combined_shard_04vox2_3D_ditto_combined_shard_06vox2_3D_ditto_combined_shard_07vox2_3D_ditto_combined_shard_08vox2_3D_ditto_combined_shard_09vox2_3D_ditto_combined_shard_10vox2_3D_ditto_combined_shard_11vox2_3D_ditto_combined_shard_12vox2_3D_ditto_combined_shard_13vox2_3D_ditto_combined_shard_14vox2_3D_ditto_combined_shard_15vox2_3D_ditto_shard_14vox2_3D_ditto_shard_03vox2_3D_ditto_shard_19vox2_3D_ditto_shard_10vox2_3D_ditto_shard_02vox2_3D_ditto_shard_25vox2_3D_ditto_shard_01vox2_3D_ditto_shard_12vox2_3D_ditto_shard_00vox2_3D_ditto_shard_05vox2_3D_ditto_shard_24vox2_3D_ditto_shard_18
