datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AV-SpeakerBench
AV-SpeakerBench
Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning.
Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/
Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench
Paper: https://arxiv.org/abs/2512.02231
Files
test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.PES-2018-2022
Dataset Card for PES-2018-2022
Update
The original dataset https://huggingface.co/datasets/amu-cai/PES-2018-2022 has been corrected and filtered.
Dataset Description
This is a dataset used and described in:
@misc{pokrywka2024gpt4,
title={GPT-4 passes most of the 297 written Polish Board Certification Examinations},
author={Jakub Pokrywka and Jeremi Kaczmarek and Edward Gorzelańczyk},
year={2024},
eprint={2405.01589}… See the full description on the dataset page: https://huggingface.co/datasets/speakleash/PES-2018-2022.arabic-tts-saudi-multi-speaker-xtts
Arabic Saudi TTS Dataset (LJSpeech Format) 🇸🇦
This dataset is designed for training Text-to-Speech (TTS) models such as XTTS_v2 using the LJSpeech format.
📌 Overview
Language: Arabic (Saudi Dialect)
Format: LJSpeech
Use Case: TTS training (XTTS_v2, YourTTS, Tacotron, etc.)
Speakers: Multi-speaker (Male & Female)
Audio Format: WAV (mono recommended)
Sample Rate: 22050 Hz (recommended)
📂 Structure
all_data/
│
├── wavs/
│ ├── sample_0.wav
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman2922/arabic-tts-saudi-multi-speaker-xtts.hollow-knight-speakatcosim-speaker-disjoint-splits
ATCOSIM speaker-disjoint splits (metadata only)
This dataset contains no audio and no transcripts. It is a split definition:
one row per ATCOSIM utterance, giving its speaker, its recording session, its
duration, and which half of a speaker-disjoint evaluation it belongs to.
The audio and transcriptions are not here because they cannot be redistributed.
The ATCOSIM corpus
manual §5.2 states that the corpus is "provided free of charge" and "permitted
to use ... for research and… See the full description on the dataset page: https://huggingface.co/datasets/moonshine-ai/atcosim-speaker-disjoint-splits.vat-rates-french-speaking-world-2026
VAT rates in the French-speaking world 2026 (France, Belgium, Switzerland, Quebec)
Official 2026 VAT/GST rates for four French-speaking jurisdictions, with category, legal basis and government source.
Jurisdiction
Standard
Reduced / special
France
20%
10% / 5.5% / 2.1%
Belgium
21%
12% / 6% / 0%
Switzerland
8.1%
3.8% (lodging) / 2.6%
Quebec
14.975% (GST 5% + QST 9.975%)
-
Sources: DGFiP, SPF Finances, Swiss Federal Tax Administration, Revenu Quebec, CRA.… See the full description on the dataset page: https://huggingface.co/datasets/tresor2k/vat-rates-french-speaking-world-2026.coherence-speaking-part2-ieltsru_librispeech_for_speaker_separationDataset for source audio separation task based on Russian LibriSpeech (RuLS) dataset. Dataset contains 50 000 audio mixtures with 2 speakers for train part; 12500 audio mixtures for test part.
Dataset also containts metadata files with audio duration (sec), source 1 and source 2 filepaths for each audio mixture.
source: https://www.openslr.org/96/
SpeakGer_sampleThis data set contains all speeches of all German federal state parliaments as well as the Bundestag from 2022, that are not interjections from the crowd or comments by the chair of the plenary session. Some additional meta data is provided, such as the date of the speech as well as the party/parties of the speaker.
This is a small test sample of the SpeakGer data set, restricted to the year 2022 and with limited meta data. For more information, please visit the official GitHub page. When… See the full description on the dataset page: https://huggingface.co/datasets/K-RLange/SpeakGer_sample.BookTubeSpeech-SpeakersSpeaker_Identificationtoefl_speaking
