datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
safi-diction-sample
Safi Diction Sample
This dataset is a sample speech dataset containing short audio recordings paired with expected transcription text from respondents.
The dataset was created to show a sample of the type of data that can be collected with Safi's collection engine. Responses were collected remotely within the span of 12 hours.
This is just a sample for development or testing - contact hq@safidata.com for the full dataset or custom datasets, or visit https://www.safidata.com/.… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-diction-sample.stt-sampler-v1
stt-sampler-v1
Licensing: clips inherit their source dataset's license — CC-BY-4.0
for MInDS-14 and FLEURS clips, CC-BY-NC-SA-4.0 for Speech-MASSIVE
clips (source_dataset column identifies each clip's origin).
A small, balanced, representative multilingual ASR eval sampler for the
OVOS Plugin Arena:
100 clips per language x 20 locales = 2000 clips, 16 kHz mono float32,
one config per locale (load_dataset("OpenVoiceOS/stt-sampler-v1", "<lang>")).
Designed to seed every STT… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/stt-sampler-v1.JA_audio_JA_text_180k_samples-Noise and silence have been removed from the begining and end of each sample.
-Unnecessary and inaccurate punctuation have been removed.
-Text has been normalized.
Normalization is based on the neologd's rules: https://github.com/neologd/mecab-ipadic-neologd/wiki/Regexp.ja.
muscat-merged-samples
MUSCAT — Merged Long-Form Samples
This dataset is a merged, long-form reformatting of
goodpiku/muscat-eval
(MUSCAT: A Multi-Device Dataset for Code-Switching ASR and Segmentation
Evaluation).
The original MUSCAT release stores each conversation as many short,
single-language segments. Here those segments are concatenated back into one
continuous recording per conversation, so each row is a single long-form
code-switching audio with inline language/timing markers. The layout… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/muscat-merged-samples.nichevault-afrikaans-asr-sample
NicheVault Afrikaans ASR — Free Sample
Overview
NicheVault Afrikaans ASR — Free Sample is a 20-clip preview of a 61.5-hour Whisper-ready Afrikaans ASR training dataset. Every clip is 16kHz mono WAV with a human-verified sentence-level transcript, formatted as JSONL metadata. All clips are pure CC BY — no ShareAlike, no copyleft obligations.
This sample is drawn from 20 clips across the full dataset's train/validation/test splits. It is provided unauthenticated… See the full description on the dataset page: https://huggingface.co/datasets/NicheVault/nichevault-afrikaans-asr-sample.reviewed_sample
Silencio Network: Multilingual Accent Speech Dataset (Sample)
Overview
Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/reviewed_sample.sundanese-spontaneous-conversation-sample
Sundanese Spontaneous Conversation — Free Sample
Unscripted two-speaker Sundanese (su-ID) conversation recorded in
West Java, Indonesia. This is a free 30-minute sample of a larger
commercially licensable corpus.
Why this exists
Sundanese has approximately 40 million native speakers, yet
spontaneous conversational data is almost absent:
Resource
Type
Volume
Common Voice Spontaneous 4.0
Spontaneous, 2 speakers
0.62 h
OpenSLR SLR36 / SLR44
Read speech
—… See the full description on the dataset page: https://huggingface.co/datasets/Indospeech/sundanese-spontaneous-conversation-sample.kare-codeswitch-samples
Kare — Code-Switching Illustrative Samples
Eight short audio clips demonstrating the English/Yoruba/Hausa/Igbo/Pidgin
code-switching that Kare, a voice-first
AI health assistant for Nigeria, is built to understand — submitted as part
of Kare's entry to the Sahara CodeSwitch Africa Challenge.
What this is — and isn't
Is: eight original sentences, written by the Kare team specifically
for this demonstration, mixing everyday Yoruba / Hausa / Igbo / Nigerian
Pidgin… See the full description on the dataset page: https://huggingface.co/datasets/vikkyblacq/kare-codeswitch-samples.english-x-code-switching-samples
Synthetic English Code-Switching Evaluation Set Samples
This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset.
Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row.
Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.medical-audio-sample-brazilian-portuguese
Julia's Data: Brazilian Portuguese Medical Audio Sample
Public sample of a Brazilian Portuguese medical audio dataset built for ASR,
TTS, and conversational AI evaluation. This repository contains deidentified
clinical source material transformed into five spoken content types and
recorded by a human speaker.
This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about
5.26 minutes of audio.
Full dataset and commercial licensing: juliasdata.com
Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.indic-audio-natural-conversations-sample
Dataset Card for Indic Audio Natural Conversations Sample Dataset
Dataset Details
Dataset Description
The IndicAudioNaturalConversations Dataset is a multilingual, multichannel, source-separated conversational speech dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi.
Curated by: snorbyte
Funded by: snorbyte
Shared by: snorbyte
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-natural-conversations-sample.english-en-x-code-switching-main-lang-samples-merged
English EN-X Code-Switching Main-Language Merged Samples
This dataset contains contiguous same-language segments from the paired mixed dataset.
Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly.
Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples-merged.personaplex-finetuning-pharma-data-sample
PersonaPlex Finetuning — Pharma Data Sample
A 10-example slice of the synthetic patient-support / medication
adherence dataset used to train
demegire/personaplex-finetune-pharma.
The on-disk layout below is exactly what the trainer in
emotion-machine-org/personaplex-finetune
consumes — use this as a template when building your own.
Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at
sample scale).
Layout
.
├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.dexterity-pidgin-voice-dataset-sample-v2Dexterity Pidgin Voice Dataset — Sample v2
Creator: Dexterity Learn.
Dataset Description
49 human-created Nigerian Pidgin English (Naija) utterances with aligned home studio audio, covering 8 domains: faith, wisdom, money, business, relationships, community, social commentary, and everyday conversation. Three utterances (U20–U22) address consent and anti-harassment themes, included for social good NLP use.
What sets this dataset apart is its dual-translation structure — every utterance… See the full description on the dataset page: https://huggingface.co/datasets/DexterityLearn/dexterity-pidgin-voice-dataset-sample-v2.ibibio-efik-speech-corpus-sample
Scuba Voice Dataset: Ibibio & Efik Sample (1 hour)
Scuba is a voice data infrastructure company collecting, labeling, and licensing speech corpora for low-resource languages. This repository contains a 1-hour sample of labeled Ibibio and Efik speech data collected from native speakers in Akwa Ibom, Nigeria.
This sample is provided for evaluation purposes only. For access to the full dataset (~200 hours across Ibibio and Efik), contact us at info@scubavoice.com.… See the full description on the dataset page: https://huggingface.co/datasets/scubavoice/ibibio-efik-speech-corpus-sample.kinyarwanda-speech-sample
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains a sample from the 500 hours of Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-sample.bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-sample-v1.inzwi-shona-sample
Inzwi — Shona Speech Corpus (Cleaned Sample)
A small, fully-documented sample of the Inzwi speech corpus: consented, peer-validated
Shona audio paired with ground-truth transcripts, prepared to be AI-ready for
automatic speech recognition (ASR). Built for the POTRAZ AI for Impact (AI4I) Challenge —
Data Track, to demonstrate the Inzwi data pipeline end to end.
This is a representative sample (the validated slice of an early, un-incentivised run),
not the full corpus. It exists… See the full description on the dataset page: https://huggingface.co/datasets/nigeLbasa/inzwi-shona-sample.indic-audio-dialog-sample
Dataset Card for Indic Dialog Sample Dataset
Dataset Details
Dataset Description
The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/1infinity0/bhasaflow-khasi-english-parallel-sample-v1.corpus5-t2inserts-sample-200-20260916saas-english-corporate-voice-dataset-sample
PROFESSIONAL CONVERSATIONAL AI VOICE DATASET - SAAS CORPORATE SERIES
Format: LJ Speech Standard Compliance | 24-bit Signed Linear PCM Mono WAV | 44.1kHz / 48kHz
Thank you for downloading this Developer Compatibility Sample Pack. This repository contains enterprise-grade, high-fidelity, human voice data optimized specifically for low-latency conversational AI UI/UX interfaces, intent-mapped software pipelines, and automated conversational SaaS call bots. This dataset is risk-free… See the full description on the dataset page: https://huggingface.co/datasets/MarieDeVox/saas-english-corporate-voice-dataset-sample.clinical-speech-samples
🏥 Clinical Speech Samples Dataset
High-Quality Clinical Audio Dataset for Medical AI Research
Curated dataset for training and evaluating clinical speech processing models
📖 Dataset Description
This dataset contains de-identified clinical audio samples for research and development of medical speech processing systems. It's designed to support training and evaluation of:
🎤 Speech enhancement in clinical settings
🔊 Audio source separation for medical… See the full description on the dataset page: https://huggingface.co/datasets/bhaskarvilles/clinical-speech-samples.cv17_sw_kenyan_sample
Common Voice 17.0 — Swahili (Kenyan Sample)
This dataset is a filtered sample of the Mozilla Common Voice 17.0 corpus, focusing on Swahili (sw) speech with Kenyan voices.
It has been subsetted for experimentation and prototyping in ASR (Automatic Speech Recognition) models targeting speech-impaired users in Kenya, covering Kenyan English and Kiswahili.
Dataset Summary
Language: Kiswahili (Swahili, sw)
Accent/Region: Kenyan speakers
Domain: Conversational… See the full description on the dataset page: https://huggingface.co/datasets/Veronica1NW/cv17_sw_kenyan_sample.bengali-multi-speaker-speech-samples
Bengali Speech: Multi-Speaker Samples
This sample shows Bengali multi-speaker speech with aligned ground-truth transcripts. It is meant to help buyers review conversational structure, speaker overlap, transcript quality, and audio consistency before scoping a larger delivery.
What This Shows
Multi-speaker Bengali speech with transcript alignment
Conversation-style audio rather than isolated prompt reading
Metadata that distinguishes language, format, and speaker… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-multi-speaker-speech-samples.mandarin-speech-samples
Mandarin Speech Samples
This sample shows Mandarin Chinese speech with clip-level metadata and preview transcripts. It is meant to help buyers review language fit, recording quality, and sample structure before scoping a larger delivery.
What This Shows
Mandarin speech audio with consistent metadata
Clip-level transcript fields for content review
Language and format signals for procurement review
Dataset Specifications
Field
Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/mandarin-speech-samples.korean-speech-samples
Korean Speech Samples
This sample shows Korean contributor speech in a consistent audio format. It is meant to help buyers review recording quality, language coverage, and metadata structure before scoping a larger delivery.
What This Shows
Korean speech recordings from contributor collection workflows
Clip-level metadata for format and review context
Ground-truth transcripts for understanding sample content
Dataset Specifications
Field… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/korean-speech-samples.corpus5-t2inserts-sample-1000-20260916luel-multilingual-tts-samples
Multilingual TTS Samples (Luel)
License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE.
A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.tamil-speech-samples
Tamil Speech Samples
This sample shows Tamil speech with ground-truth transcripts and consistent audio metadata. It is meant to help buyers review language fit, transcript quality, and capture format before scoping a larger delivery.
What This Shows
Tamil speech recordings with paired transcripts
Ground-truth labels at the clip level
Format metadata for review and delivery planning
Dataset Specifications
Field
Value
Modality
Audio… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/tamil-speech-samples.
