datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AV-SpeakerBench
AV-SpeakerBench
Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning.
Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/
Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench
Paper: https://arxiv.org/abs/2512.02231
Files
test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.aiera-speaker-assign
Transcript Speaker Identification Dataset
Description
This dataset is designed to facilitate the development and evaluation of models that identify and assign speakers and speaker changes within event transcripts. It consists of segments from various transcripts where the primary task is to determine who the speaker is, based on the given textual context and a list of possible speakers.
The dataset was assembled from three earnings events:
Q4 2023 Amazon.Com Inc Earnings… See the full description on the dataset page: https://huggingface.co/datasets/Aiera/aiera-speaker-assign.
