datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VoiceAssistant-Eval
🔥 VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
[🌐 Homepage]
[🔮 Visualization]
[💻 Github]
[📖 Paper]
[📊 Leaderboard ]
[📊 Detailed Leaderboard ]
[📊 Roleplay Leaderboard ]
🚀 Data Usage
from datasets import load_dataset
for split in ['listening_general', 'listening_music', 'listening_sound', 'listening_speech',
'speaking_assistant', 'speaking_emotion', 'speaking_instruction_following'… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/VoiceAssistant-Eval.AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.LongSpeech-Eval
The proposed Long-Speech understanding evaluation dataset for the paper 'FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing'
Usage
First download the model from Model.
Then please refer to Github Page.
Requirements
We suggest to run with Python 3.10.
Examples of usage:
git clone https://github.com/ictnlp/FastLongSpeech.git
cd transformers-main
pip install -e .
pip install deepspeed sentencepiece librosa… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/LongSpeech-Eval.bba-storm-eval
BIG-Bench Audio — Thunderphone Storm Evaluation
Results
994 / 1000 correct (99.4%)
Category
Accuracy
formal_fallacies
250/250 (100%)
navigate
248/250 (99.2%)
object_counting
250/250 (100%)
web_of_lies
246/250 (98.4%)
What is this?
1,000 spoken reasoning questions from the BIG-Bench Audio benchmark, processed through Thunderphone's Storm voice pipeline (with intelligence boost enabled). Each example includes the model's spoken response… See the full description on the dataset page: https://huggingface.co/datasets/kolchinski/bba-storm-eval.VoiceAssistant-Eval
🔥 VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
[🌐 Homepage]
[🔮 Visualization]
[💻 Github]
[📖 Paper]
[📊 Leaderboard ]
[📊 Detailed Leaderboard ]
[📊 Roleplay Leaderboard ]
🚀 Data Usage
from datasets import load_dataset
for split in ['listening_general', 'listening_music', 'listening_sound', 'listening_speech',
'speaking_assistant', 'speaking_emotion', 'speaking_instruction_following'… See the full description on the dataset page: https://huggingface.co/datasets/SamSoko83/VoiceAssistant-Eval.BELLE-eval-S2S
BELLE-eval-S2S
💡 Dataset Description
BELLE-eval-S2S is a Chinese evaluation dataset for speech-to-speech conversational tasks. It contains 250 Chinese audio samples with corresponding text annotations and is intended for model evaluation rather than training.
🔗 Source
Original text source: the test set from LianjiaTech/BELLE
This dataset is built by filtering 250 samples from the original test set and synthesizing them into speech audio
📖 Data… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/BELLE-eval-S2S.
