datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stt-sampler-v1
stt-sampler-v1
Licensing: clips inherit their source dataset's license — CC-BY-4.0
for MInDS-14 and FLEURS clips, CC-BY-NC-SA-4.0 for Speech-MASSIVE
clips (source_dataset column identifies each clip's origin).
A small, balanced, representative multilingual ASR eval sampler for the
OVOS Plugin Arena:
100 clips per language x 20 locales = 2000 clips, 16 kHz mono float32,
one config per locale (load_dataset("OpenVoiceOS/stt-sampler-v1", "<lang>")).
Designed to seed every STT… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/stt-sampler-v1.muscat-merged-samples
MUSCAT — Merged Long-Form Samples
This dataset is a merged, long-form reformatting of
goodpiku/muscat-eval
(MUSCAT: A Multi-Device Dataset for Code-Switching ASR and Segmentation
Evaluation).
The original MUSCAT release stores each conversation as many short,
single-language segments. Here those segments are concatenated back into one
continuous recording per conversation, so each row is a single long-form
code-switching audio with inline language/timing markers. The layout… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/muscat-merged-samples.JA_audio_JA_text_180k_samples-Noise and silence have been removed from the begining and end of each sample.
-Unnecessary and inaccurate punctuation have been removed.
-Text has been normalized.
Normalization is based on the neologd's rules: https://github.com/neologd/mecab-ipadic-neologd/wiki/Regexp.ja.
reviewed_sample
Silencio Network: Multilingual Accent Speech Dataset (Sample)
Overview
Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/reviewed_sample.english-x-code-switching-samples
Synthetic English Code-Switching Evaluation Set Samples
This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset.
Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row.
Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.indic-audio-natural-conversations-sample
Dataset Card for Indic Audio Natural Conversations Sample Dataset
Dataset Details
Dataset Description
The IndicAudioNaturalConversations Dataset is a multilingual, multichannel, source-separated conversational speech dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi.
Curated by: snorbyte
Funded by: snorbyte
Shared by: snorbyte
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-natural-conversations-sample.english-en-x-code-switching-main-lang-samples-merged
English EN-X Code-Switching Main-Language Merged Samples
This dataset contains contiguous same-language segments from the paired mixed dataset.
Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly.
Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples-merged.kinyarwanda-speech-sample
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains a sample from the 500 hours of Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-sample.indic-audio-dialog-sample
Dataset Card for Indic Dialog Sample Dataset
Dataset Details
Dataset Description
The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.cv17_sw_kenyan_sample
Common Voice 17.0 — Swahili (Kenyan Sample)
This dataset is a filtered sample of the Mozilla Common Voice 17.0 corpus, focusing on Swahili (sw) speech with Kenyan voices.
It has been subsetted for experimentation and prototyping in ASR (Automatic Speech Recognition) models targeting speech-impaired users in Kenya, covering Kenyan English and Kiswahili.
Dataset Summary
Language: Kiswahili (Swahili, sw)
Accent/Region: Kenyan speakers
Domain: Conversational… See the full description on the dataset page: https://huggingface.co/datasets/Veronica1NW/cv17_sw_kenyan_sample.corpus5-t2inserts-sample-200-20260916indic-tts-sample-snac-encoded
Dataset Card for Indic TTS Sample SNAC Encoded Dataset
Dataset Details
Dataset Description
The IndicTTSSampleSNACEncoded Dataset is a multilingual, text-speech pair sample dataset. It features ~135 hours of human-voiced recordings of transcripts in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi, across multiple speakers and other metadata.
Curated by: snorbyte
Funded by: snorbyte
Shared by: snorbyte… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-tts-sample-snac-encoded.corpus5-t2inserts-sample-1000-20260916stillalive-sample-1000-uniform-20260919
stillalive sample 1000 (uniform over 625,217)
1000 chunks drawn uniformly at random, without replacement, from all 625,217 verified
chunks merged so far (batches saq_b0001..saq_b0040, partA final). Same schema as the full
dataset (audio embedded in the audio struct column).
Distinct recordings: 913; audio: 2.01 h
Randomness: secrets.SystemRandom().sample (os.urandom CSPRNG, Fisher-Yates). No seed.
Audit: sample_manifest.json holds sha256 over the sorted (recording_id, chunk_id)… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/stillalive-sample-1000-uniform-20260919.stillalive-sample-1000-20260919
stillalive-sample-1000-20260919
1000 chunks drawn uniformly at random (seed 20260919) from the 129,681 chunks merged so far into the local
stillalive part-A dataset (stillalive-t2inserts-20260919-combined-publication-stillalive-sbpn-diarized-t2inserts-demucs-320k-local-20260919-partA), i.e. from batches saq_b0001..saq_b0010 (9 batches).
The full run is still in progress, so this sample covers only the batches processed so far.
Recordings represented: 706; audio: 2.16 h… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/stillalive-sample-1000-20260919.tibetan-audio-english-6datasets-sample2
Tibetan Audio-English Sentence Dataset (Sample)
This is a sample dataset containing 5 rows from a merged collection of 6 Tibetan audio datasets with English translations.
📊 Dataset Details
Total Samples in Full Dataset: 17,278
Samples in This Preview: 5
Format: Audio + English sentence pairs
Audio Sampling Rate: 16,000 Hz
Languages: Tibetan (audio) → English (text)
🗂️ Source Datasets
This sample is merged from 6 datasets:… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/tibetan-audio-english-6datasets-sample2.english-en-x-code-switching-main-lang-samples
English EN-X Code-Switching Main-Language Samples
This dataset contains the individual full FLEURS utterance chunks used to build the paired mixed dataset.
Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly.
Each selected utterance is RMS-normalized to -20.0… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples.stillalive-sample-1000-uniform-20260920
stillalive sample 1000 (uniform over partA final + partB merged)
1000 chunks drawn uniformly at random, without replacement, from every verified
merged row available at sampling time (partA finalized + partB merged so far).
Same schema as the full dataset (audio embedded in the audio struct column).
Distinct recordings: 970; audio: 1.93 h
Randomness: secrets.SystemRandom().sample (os.urandom CSPRNG, Fisher-Yates). No seed.
Audit: sample_manifest.json holds sha256 over the… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/stillalive-sample-1000-uniform-20260920.world-audio-natural-conversations-sample
Dataset Card for World Audio Natural Conversations Sample Dataset
Dataset Details
Dataset Description
The WorldAudioNaturalConversations Dataset is a multilingual, multichannel, source-separated conversational speech dataset. It features human-voiced recordings of dialogues in various world languages.
Curated by: snorbyte
Funded by: snorbyte
Shared by: snorbyte
Language(s) (NLP): it, es
License: CC BY 4.0
Dataset Sources
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/world-audio-natural-conversations-sample.indic-text-audio-sample
Dataset Card for Indic Text Audio Sample Dataset
Dataset Details
Dataset Description
The IndicTextAudioSample Dataset is a multilingual, text-speech pair sample dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi.
Curated by: snorbyte
Funded by: snorbyte
Shared by: snorbyte
Language(s) (NLP): hi, ta, te, pa, ml, kn, bn, gu, mr
License: CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-text-audio-sample.anv-za-sot-1h-sample-dataset
Sesotho Sample Dataset - Next Voices-ZA (South Africa) - Multilingual Speech Dataset - Sesotho
This dataset includes scripted and unscripted speech across various domains such as agriculture, health, finance, sports, transport, culture, society and general topics. It is primarily designed for automatic speech recognition (ASR).
Use Restriction:
The persons whose voices are included in this dataset, and the creators and owners of this dataset* do not give consent in… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/anv-za-sot-1h-sample-dataset.split_sample
Silencio Network: Multilingual Accent Speech Dataset (Sample)
Overview
Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/split_sample.tibetan-audio-english-6datasets-sample
Tibetan Audio-English Sentence Dataset (Sample)
This is a sample dataset containing 5 rows from a merged collection of 6 Tibetan audio datasets with English translations.
📊 Dataset Details
Total Samples in Full Dataset: 17,278
Samples in This Preview: 5
Format: Audio + English sentence pairs
Audio Sampling Rate: 16,000 Hz
Languages: Tibetan (audio) → English (text)
🗂️ Source Datasets
This sample is merged from 6 datasets:… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/tibetan-audio-english-6datasets-sample.
