CoolFace
Datasetpublic

sarvamai/audiollm-evals

This evaluation set contains ~100 questions in both text and audio format in Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, and Telugu. We use this dataset internally at Sarvam to evaluate the performance of our audio models. We open-source this data to enable the research community to replicate the results mentioned in our Shuka blog. By deisgn, the questions are sometimes vague, and the audio has noise and other inconsistencies, to measure the robustness of… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/audiollm-evals.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
3likes37downloads
Dataset Card

This evaluation set contains ~100 questions in both text and audio format in Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, and Telugu. We use this dataset internally at Sarvam to evaluate the performance of our audio models. We open-source this data to enable the research community to replicate the results mentioned in our Shuka blog.

By deisgn, the questions are sometimes vague, and the audio has noise and other inconsistencies, to measure the robustness of models.