datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MM-ContextASR-Bench
MM-ContextASR Bench
Metadata and evaluation splits for Multimodal Conversational Context for
LLM-Based ASR: Data Construction, Training, and Benchmark.
Dataset summary
Config
Examples
Audio
Context
Primary metric
mm_contextasr
1,250 (250 current utterances × 5 histories)
1,439 WAV files included
Controlled user-assistant dialogue
entity Recall
kespeech
19,212
Source ID only
Same-speaker speech and transcript
CER, SER, entity Recall
cv_yue
3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.MMOU
MMOU
Massive Multi-Task Omni Understanding and Reasoning
Benchmark for Long and Complex Real-World Videos
Project Page
·
HuggingFace
·
Videos (Community Hosted)
·
Paper
·
Evaluator
MMOU evaluates joint audio-visual understanding and reasoning in long and complex real-world videos.
Dataset Summary
MMOU is a benchmark for evaluating whether multimodal models can jointly reason over video, speech, sound, music, and long-range temporal context in… See the full description on the dataset page: https://huggingface.co/datasets/lililiy/MMOU.
