datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
musan
MUSAN: A Music, Speech, and Noise Corpus
MUSAN is a corpus of music, speech, and noise recordings designed for training models for voice activity detection and music/speech discrimination. This is a comprehensive collection suitable for various audio processing tasks.
Dataset Structure
The dataset is organized into three main categories:
1. Music (~42 hours)
Subcategories: Classical, Pop/Rock, Jazz, and more
Sources: Free Music Archive, Jamendo, and others… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/musan.fleurs-full
FLEURS Full - Test Set for ASR Benchmarking
Complete test set of Google FLEURS for all 30 languages supported by Qwen3-ASR, prepared for benchmarking with FluidAudio.
Languages (30)
Asian Languages (13)
Code
Language
Samples
cmn_hans_cn
Chinese (Mandarin)
945
yue_hant_hk
Cantonese
819
ja_jp
Japanese
650
ko_kr
Korean
382
vi_vn
Vietnamese
857
th_th
Thai
1,021
id_id
Indonesian
687
ms_my
Malay
749
hi_in
Hindi
418
ar_eg
Arabic (Egyptian)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs-full.ami-corpus-mirror
AMI Corpus Mirror
Mirror of the subset of the AMI Meeting Corpus used by
FluidAudio diarization benchmarks. Hosted here so CI and
local benchmark runs do not depend on the availability of the upstream groups.inf.ed.ac.uk server
(see FluidAudio#752).
Contents
annotations/ami_public_manual_1.6.2.zip — AMI public manual annotations v1.6.2
(repackaged from the official archive; identical content, including segments/, words/,
corpusResources/meetings.xml)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/ami-corpus-mirror.THCHS-30-tests
THCHS-30 Test Set
THCHS-30 test split for Mandarin Chinese speech recognition benchmarking.
Dataset Info
Language: Mandarin Chinese (zh-CN)
Samples: 2,495
Speakers: 10
Sample Rate: 16 kHz
License: Apache 2.0
Usage
from datasets import load_dataset
# After uploading to HuggingFace
dataset = load_dataset("your-username/thchs30-test")
# Example
print(dataset['train'][0])
# {
# 'audio': {'array': [...], 'sampling_rate': 16000, 'path': 'audio/D11_750.wav'},
#… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/THCHS-30-tests.JSUT-basic5000
JSUT (Japanese Speech Corpus) - Test Subset
A test subset of the JSUT corpus containing 500 Japanese utterances from the basic5000 dataset (BASIC5000_4501-5000).
Dataset Structure
jsut_ver1.1/
└── basic5000/
├── wav/ # WAV audio files (500 files, 48kHz)
├── transcript_utf8.txt # Transcriptions
└── recording_info.txt # Recording dates
File Formats
transcript_utf8.txt
BASIC5000_4501:だが、エーアイセンター稼動を快く思わない...… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/JSUT-basic5000.librispeechfleurs
FLEURS Test Dataset
Reorganized FLEURS test dataset with audio and transcripts together.
Structure
fleurs-test/
├── en_us/
│ ├── en_us_0000.wav
│ ├── en_us_0001.wav
│ ├── ...
│ ├── en_us.trans.txt (LibriSpeech format)
│ ├── en_us.csv (detailed metadata)
│ └── en_us.json (JSON metadata)
├── fr_fr/
│ └── ...
└── ...
Languages
bg_bg: 350 test samples
cs_cz: 350 test samples
da_dk: 930 test samples
de_de: 350 test samples
el_gr: 650… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs.tooltalk-samples
ToolTalk Samples - High Quality Duplex Speech and Tool-calling in Customer Service Domain
Two people improvise realistic customer-service calls while one operates a live, stateful tool environment—with synchronized speaker-separated audio, tool calls, and outcomes.
▶ Listen to Clean · ▶ Listen to Noisy · Discuss the full dataset
In this sample: 26 calls · 90.7 minutes · 7 sample domains · 207 tool calls
Technical specs: 48 kHz / 32-bit PCM speaker-separated source… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/tooltalk-samples.cv-corpus-25.0-ja
Mozilla Common Voice 25.0 - Japanese Test Set (Complete)
Dataset Description
Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace.
Key Features
Size: 9,019 validated test utterances
Coverage: 100% of official Common Voice 25.0 Japanese test split
Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.friend-bench
Can a model — or a human — tell how two people are related from a 20-second clip of how they interact?
🌐 Built on Seamless Interaction
FriendBench is a suite of benchmarks for social perception from thin-slice dyadic
interaction — inferring facts about two people's relationship from a brief clip of how they
interact, built on the Seamless Interaction
dataset. Each released set is a config of this repository.
🎧 Multi-modal — text, audio, and video for every clip
🎯 Objective label —… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/friend-bench.multimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.multimodal-expert-instruction-samples
Multimodal Expert Instruction Samples - Musical Instrument Lessons with Channel-separated Audio and Video
A music teacher and a student work through two one-on-one lessons: both voices and both instruments on separate tracks, the student on camera, with the lesson plans, the instructions given to each side and both sides' post-lesson ratings alongside.
▶ Watch the lessons · See Peer Collaboration samples · Discuss the full collection
Sister collection: Peer Collaboration… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-expert-instruction-samples.aec-challenge-synthetic-mini
AEC-Challenge synthetic mini (200 examples)
First 200 examples (shard 0, dataset order) of the Microsoft AEC-Challenge
synthetic set (https://github.com/microsoft/AEC-Challenge/tree/main/datasets/synthetic,
Sridhar et al., ICASSP 2021, arXiv:2009.04972), exported from the
PandaLT/microsoft-AEC-dataset parquet mirror as 16 kHz 16-bit mono WAV:
fileid_<id>_mic.wav near-end microphone signal (near-end speech + echo, optionally noise)
fileid_<id>_lpb.wav far-end / loopback… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/aec-challenge-synthetic-mini.audiofluid-2-sft-asr
Fluid 2 — synthetic dictation cleanup
Fluid 2 is an English supervised-fine-tuning corpus for models that turn noisy automatic-speech-recognition output into the written insertion a user intended. It contains 354,549 rows in official document-grouped 96/2/2 splits, 861.3 hours of processed 16 kHz speech, and 8.48M target-side loss tokens in 355 Parquet shards (49.25 GiB).
This is not an ordinary transcription dataset. The model sees document context plus an ASR hypothesis and… See the full description on the dataset page: https://huggingface.co/datasets/johnbean393/fluid-2-sft-asr.
