CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenVoiceOS /stt-sampler-v1 stt-sampler-v1 Licensing: clips inherit their source dataset's license — CC-BY-4.0 for MInDS-14 and FLEURS clips, CC-BY-NC-SA-4.0 for Speech-MASSIVE clips (source_dataset column identifies each clip's origin). A small, balanced, representative multilingual ASR eval sampler for the OVOS Plugin Arena: 100 clips per language x 20 locales = 2000 clips, 16 kHz mono float32, one config per locale (load_dataset("OpenVoiceOS/stt-sampler-v1", "<lang>")). Designed to seed every STT… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/stt-sampler-v1.audioautomatic-speech-recognition1K<n<10K0 likes209 downloads1mo agoHugging Face02martinturuta /safi-diction-sample Safi Diction Sample This dataset is a sample speech dataset containing short audio recordings paired with expected transcription text from respondents. The dataset was created to show a sample of the type of data that can be collected with Safi's collection engine. Responses were collected remotely within the span of 12 hours. This is just a sample for development or testing - contact hq@safidata.com for the full dataset or custom datasets, or visit https://www.safidata.com/.… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-diction-sample.audioautomatic-speech-recognitionn<1K0 likes207 downloads19d agoHugging Face03Sin2pi /JA_audio_JA_text_180k_samples-Noise and silence have been removed from the begining and end of each sample. -Unnecessary and inaccurate punctuation have been removed. -Text has been normalized. Normalization is based on the neologd's rules: https://github.com/neologd/mecab-ipadic-neologd/wiki/Regexp.ja. audioautomatic-speech-recognition100K<n<1M9 likes173 downloads8mo agoHugging Face04BrunoHays /muscat-merged-samples MUSCAT — Merged Long-Form Samples This dataset is a merged, long-form reformatting of goodpiku/muscat-eval (MUSCAT: A Multi-Device Dataset for Code-Switching ASR and Segmentation Evaluation). The original MUSCAT release stores each conversation as many short, single-language segments. Here those segments are concatenated back into one continuous recording per conversation, so each row is a single long-form code-switching audio with inline language/timing markers. The layout… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/muscat-merged-samples.audioautomatic-speech-recognitionn<1K0 likes127 downloads16d agoHugging Face05demegire /personaplex-finetuning-pharma-data-sample PersonaPlex Finetuning — Pharma Data Sample A 10-example slice of the synthetic patient-support / medication adherence dataset used to train demegire/personaplex-finetune-pharma. The on-disk layout below is exactly what the trainer in emotion-machine-org/personaplex-finetune consumes — use this as a template when building your own. Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at sample scale). Layout . ├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.audiotext-to-speechn<1K0 likes65 downloads4mo agoHugging Face06NicheVault /nichevault-afrikaans-asr-sample NicheVault Afrikaans ASR — Free Sample Overview NicheVault Afrikaans ASR — Free Sample is a 20-clip preview of a 61.5-hour Whisper-ready Afrikaans ASR training dataset. Every clip is 16kHz mono WAV with a human-verified sentence-level transcript, formatted as JSONL metadata. All clips are pure CC BY — no ShareAlike, no copyleft obligations. This sample is drawn from 20 clips across the full dataset's train/validation/test splits. It is provided unauthenticated… See the full description on the dataset page: https://huggingface.co/datasets/NicheVault/nichevault-afrikaans-asr-sample.audioautomatic-speech-recognitionn<1K0 likes55 downloads2mo agoHugging Face07jml2026 /reviewed_sample Silencio Network: Multilingual Accent Speech Dataset (Sample) Overview Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/reviewed_sample.audioautomatic-speech-recognition1K<n<10K0 likes54 downloads8mo agoHugging Face08Indospeech /sundanese-spontaneous-conversation-samplegated Sundanese Spontaneous Conversation — Free Sample Unscripted two-speaker Sundanese (su-ID) conversation recorded in West Java, Indonesia. This is a free 30-minute sample of a larger commercially licensable corpus. Why this exists Sundanese has approximately 40 million native speakers, yet spontaneous conversational data is almost absent: Resource Type Volume Common Voice Spontaneous 4.0 Spontaneous, 2 speakers 0.62 h OpenSLR SLR36 / SLR44 Read speech —… See the full description on the dataset page: https://huggingface.co/datasets/Indospeech/sundanese-spontaneous-conversation-sample.audioautomatic-speech-recognitionn<1K1 likes52 downloads10d agoHugging Face09vikkyblacq /kare-codeswitch-samples Kare — Code-Switching Illustrative Samples Eight short audio clips demonstrating the English/Yoruba/Hausa/Igbo/Pidgin code-switching that Kare, a voice-first AI health assistant for Nigeria, is built to understand — submitted as part of Kare's entry to the Sahara CodeSwitch Africa Challenge. What this is — and isn't Is: eight original sentences, written by the Kare team specifically for this demonstration, mixing everyday Yoruba / Hausa / Igbo / Nigerian Pidgin… See the full description on the dataset page: https://huggingface.co/datasets/vikkyblacq/kare-codeswitch-samples.audioautomatic-speech-recognitionn<1K0 likes46 downloads8d agoHugging Face10BrunoHays /english-x-code-switching-samples Synthetic English Code-Switching Evaluation Set Samples This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.audioautomatic-speech-recognition10K<n<100K0 likes43 downloads5mo agoHugging Face11juliasdata /medical-audio-sample-brazilian-portuguese Julia's Data: Brazilian Portuguese Medical Audio Sample Public sample of a Brazilian Portuguese medical audio dataset built for ASR, TTS, and conversational AI evaluation. This repository contains deidentified clinical source material transformed into five spoken content types and recorded by a human speaker. This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about 5.26 minutes of audio. Full dataset and commercial licensing: juliasdata.com Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.audioautomatic-speech-recognitionn<1K1 likes41 downloads6mo agoHugging Face12snorbyte /indic-audio-natural-conversations-samplegated Dataset Card for Indic Audio Natural Conversations Sample Dataset Dataset Details Dataset Description The IndicAudioNaturalConversations Dataset is a multilingual, multichannel, source-separated conversational speech dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi. Curated by: snorbyte Funded by: snorbyte Shared by: snorbyte Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-natural-conversations-sample.audioaudio-to-audion<1K1 likes38 downloads1y agoHugging Face13BrunoHays /english-en-x-code-switching-main-lang-samples-merged English EN-X Code-Switching Main-Language Merged Samples This dataset contains contiguous same-language segments from the paired mixed dataset. Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly. Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples-merged.audioautomatic-speech-recognitionn<1K0 likes33 downloads5mo agoHugging Face14scubavoice /ibibio-efik-speech-corpus-sample Scuba Voice Dataset: Ibibio & Efik Sample (1 hour) Scuba is a voice data infrastructure company collecting, labeling, and licensing speech corpora for low-resource languages. This repository contains a 1-hour sample of labeled Ibibio and Efik speech data collected from native speakers in Akwa Ibom, Nigeria. This sample is provided for evaluation purposes only. For access to the full dataset (~200 hours across Ibibio and Efik), contact us at info@scubavoice.com.… See the full description on the dataset page: https://huggingface.co/datasets/scubavoice/ibibio-efik-speech-corpus-sample.audioautomatic-speech-recognitionn<1K1 likes27 downloads3mo agoHugging Face15MEDHARVIX-SYSTEMS /bhasaflow-khasi-english-parallel-sample-v1 BhasaFlow Khasi-English Parallel Sample v1 A professionally curated, gold-standard parallel speech and text corpus for the Khasi language. Published by Medharvix Systems Private Limited Part of the BhasaFlow Low-Resource Language Technology Initiative Overview This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-sample-v1.audioautomatic-speech-recognitionn<1K23 likes25 downloads5mo agoHugging Face16badrex /kinyarwanda-speech-sample Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains a sample from the 500 hours of Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~500 hours of transcribed speech Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-sample.audioautomatic-speech-recognition1K<n<10K0 likes24 downloads1y agoHugging Face17nigeLbasa /inzwi-shona-sample Inzwi — Shona Speech Corpus (Cleaned Sample) A small, fully-documented sample of the Inzwi speech corpus: consented, peer-validated Shona audio paired with ground-truth transcripts, prepared to be AI-ready for automatic speech recognition (ASR). Built for the POTRAZ AI for Impact (AI4I) Challenge — Data Track, to demonstrate the Inzwi data pipeline end to end. This is a representative sample (the validated slice of an early, un-incentivised run), not the full corpus. It exists… See the full description on the dataset page: https://huggingface.co/datasets/nigeLbasa/inzwi-shona-sample.audioautomatic-speech-recognitionn<1K0 likes24 downloads2mo agoHugging Face18snorbyte /indic-audio-dialog-samplegated Dataset Card for Indic Dialog Sample Dataset Dataset Details Dataset Description The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.audioaudio-to-audio1K<n<10K1 likes23 downloads1y agoHugging Face19Cybrpgs /corpus5-t2inserts-sample-200-20260916gatedaudioautomatic-speech-recognitionn<1K0 likes22 downloads7d agoHugging Face20Veronica1NW /cv17_sw_kenyan_sample Common Voice 17.0 — Swahili (Kenyan Sample) This dataset is a filtered sample of the Mozilla Common Voice 17.0 corpus, focusing on Swahili (sw) speech with Kenyan voices. It has been subsetted for experimentation and prototyping in ASR (Automatic Speech Recognition) models targeting speech-impaired users in Kenya, covering Kenyan English and Kiswahili. Dataset Summary Language: Kiswahili (Swahili, sw) Accent/Region: Kenyan speakers Domain: Conversational… See the full description on the dataset page: https://huggingface.co/datasets/Veronica1NW/cv17_sw_kenyan_sample.audioautomatic-speech-recognition1K<n<10K0 likes21 downloads1y agoHugging Face211infinity0 /bhasaflow-khasi-english-parallel-sample-v1 BhasaFlow Khasi-English Parallel Sample v1 A professionally curated, gold-standard parallel speech and text corpus for the Khasi language. Published by Medharvix Systems Private Limited Part of the BhasaFlow Low-Resource Language Technology Initiative Overview This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/1infinity0/bhasaflow-khasi-english-parallel-sample-v1.audioautomatic-speech-recognitionn<1K1 likes21 downloads5mo agoHugging Face22MarieDeVox /saas-english-corporate-voice-dataset-sample PROFESSIONAL CONVERSATIONAL AI VOICE DATASET - SAAS CORPORATE SERIES Format: LJ Speech Standard Compliance | 24-bit Signed Linear PCM Mono WAV | 44.1kHz / 48kHz Thank you for downloading this Developer Compatibility Sample Pack. This repository contains enterprise-grade, high-fidelity, human voice data optimized specifically for low-latency conversational AI UI/UX interfaces, intent-mapped software pipelines, and automated conversational SaaS call bots. This dataset is risk-free… See the full description on the dataset page: https://huggingface.co/datasets/MarieDeVox/saas-english-corporate-voice-dataset-sample.audioautomatic-speech-recognitionn<1K1 likes20 downloads4mo agoHugging Face23psdn-ai /bengali-multi-speaker-speech-samplesgated Bengali Speech: Multi-Speaker Samples This sample shows Bengali multi-speaker speech with aligned ground-truth transcripts. It is meant to help buyers review conversational structure, speaker overlap, transcript quality, and audio consistency before scoping a larger delivery. What This Shows Multi-speaker Bengali speech with transcript alignment Conversation-style audio rather than isolated prompt reading Metadata that distinguishes language, format, and speaker… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-multi-speaker-speech-samples.audioautomatic-speech-recognitionn<1K0 likes16 downloads3mo agoHugging Face24psdn-ai /mandarin-speech-samplesgated Mandarin Speech Samples This sample shows Mandarin Chinese speech with clip-level metadata and preview transcripts. It is meant to help buyers review language fit, recording quality, and sample structure before scoping a larger delivery. What This Shows Mandarin speech audio with consistent metadata Clip-level transcript fields for content review Language and format signals for procurement review Dataset Specifications Field Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/mandarin-speech-samples.audioautomatic-speech-recognitionn<1K0 likes16 downloads3mo agoHugging Face25psdn-ai /korean-speech-samplesgated Korean Speech Samples This sample shows Korean contributor speech in a consistent audio format. It is meant to help buyers review recording quality, language coverage, and metadata structure before scoping a larger delivery. What This Shows Korean speech recordings from contributor collection workflows Clip-level metadata for format and review context Ground-truth transcripts for understanding sample content Dataset Specifications Field… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/korean-speech-samples.audioautomatic-speech-recognitionn<1K0 likes16 downloads3mo agoHugging Face26Cybrpgs /corpus5-t2inserts-sample-1000-20260916gatedaudioautomatic-speech-recognition1K<n<10K0 likes16 downloads7d agoHugging Face27lilgoose777 /tibetan-audio-english-6datasets-sample2 Tibetan Audio-English Sentence Dataset (Sample) This is a sample dataset containing 5 rows from a merged collection of 6 Tibetan audio datasets with English translations. 📊 Dataset Details Total Samples in Full Dataset: 17,278 Samples in This Preview: 5 Format: Audio + English sentence pairs Audio Sampling Rate: 16,000 Hz Languages: Tibetan (audio) → English (text) 🗂️ Source Datasets This sample is merged from 6 datasets:… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/tibetan-audio-english-6datasets-sample2.audioautomatic-speech-recognitionn<1K0 likes15 downloads8mo agoHugging Face28Luel-ai /luel-multilingual-tts-samplesgated Multilingual TTS Samples (Luel) License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE. A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.audiotext-to-speechn<1K0 likes15 downloads5mo agoHugging Face29psdn-ai /tamil-speech-samplesgated Tamil Speech Samples This sample shows Tamil speech with ground-truth transcripts and consistent audio metadata. It is meant to help buyers review language fit, transcript quality, and capture format before scoping a larger delivery. What This Shows Tamil speech recordings with paired transcripts Ground-truth labels at the clip level Format metadata for review and delivery planning Dataset Specifications Field Value Modality Audio… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/tamil-speech-samples.audioautomatic-speech-recognitionn<1K0 likes14 downloads3mo agoHugging Face30Cybrpgs /stillalive-sample-1000-20260919gated stillalive-sample-1000-20260919 1000 chunks drawn uniformly at random (seed 20260919) from the 129,681 chunks merged so far into the local stillalive part-A dataset (stillalive-t2inserts-20260919-combined-publication-stillalive-sbpn-diarized-t2inserts-demucs-320k-local-20260919-partA), i.e. from batches saq_b0001..saq_b0010 (9 batches). The full run is still in progress, so this sample covers only the batches processed so far. Recordings represented: 706; audio: 2.16 h… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/stillalive-sample-1000-20260919.audioautomatic-speech-recognition1K<n<10K0 likes14 downloads4d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.