CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01liva-ai /code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here. Code-Switching ASR Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.audioautomatic-speech-recognitionn<1K0 likes181 downloads2mo agoHugging Face02Seif-Eldeen-Sameh /asr_codeswitched_dataset Arabic/English Code-Switched ASR Dataset Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical vocabulary alternate within sentences. Composition Source Description EJUST custom recordings Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs) MohamedRashad/arabic-english-code-switching ~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.audioautomatic-speech-recognition10K<n<100K1 likes164 downloads4mo agoHugging Face03devrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes103 downloads7mo agoHugging Face04Tim2190 /kazakh-codeswitch-asr Kazakh Code-Switching ASR Benchmark A benchmark for evaluating ASR systems on natural Kazakh speech that code-switches with Russian — the everyday Kazakh–Russian mixing found in stand-up, interviews and vlogs, not scripted read speech. This is, to our knowledge, the first speech/ASR resource targeting the Kazakh–Russian code-switching pair (existing Kazakh–Russian NLP resources are text-only). Code, scoring harness and full analysis:… See the full description on the dataset page: https://huggingface.co/datasets/Tim2190/kazakh-codeswitch-asr.audioautomatic-speech-recognitionn<1K0 likes96 downloads2mo agoHugging Face05Panhapich /khmer-english-codeswitch-tts Khmer–English Code-Switch Synthetic Speech 8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech, generated with VoxCPM2 from code-switch text manufactured by confirmed lexical substitution over a Khmer–English parallel corpus. ⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a TTS model, and the code-switch sentences were manufactured by word substitution — they are not transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.audioautomatic-speech-recognition1K<n<10K2 likes79 downloads1mo agoHugging Face06Kimyayd /vocal-money-codeswitch-asr-benchmark Vocal Money — Yoruba–English Code-Switched ASR Benchmark A 210-clip evaluation subset used to benchmark five speech recognition systems on naturally code-switched Yoruba–English speech, together with the reference transcriptions and the output of every system on every clip, so that the published results can be recomputed or contradicted. Produced for the MLC (Africa) × Intron Agentic Voice AI Challenge, Deep Learning Indaba 2026. Team Vocal Money — Hospice Hounfodji, Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/Kimyayd/vocal-money-codeswitch-asr-benchmark.audioautomatic-speech-recognitionn<1K0 likes76 downloads2mo agoHugging Face07shangeth /ljspeech-mimi-codes LJSpeech — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for the LJSpeech corpus — 13,100 English utterances from a single female speaker reading public-domain audiobook passages (~24 hours). This dataset contains codes only, not audio. For waveforms, go to the original LJSpeech release; these codes are designed to be loaded alongside it for training Mimi-based speech models without paying the ~1 hour of GPU extraction cost. Schema One row per utterance:… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/ljspeech-mimi-codes.tabulartext-to-speech10K<n<100K0 likes67 downloads5mo agoHugging Face08Panhapich /khmer-english-codeswitch-tts-llm Khmer–English Code-Switch Synthetic Speech (LLM-authored) 19,825 utterances / 21.7 hours of synthetic Khmer–English code-switched speech at 16 kHz, generated with VoxCPM2 from code-switch sentences written by an LLM and validated programmatically. ⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a TTS model, and every sentence was written by a language model — they are not transcripts of anything a person said. It is intended as an… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts-llm.automatic-speech-recognition10K<n<100K0 likes62 downloads1mo agoHugging Face09Atufa /codeswitch-fr-en-kyutai-stt Code-Switched French–English STT Probe Dataset Dataset Summary This dataset contains 10 audio clips of French–English code-switched speech, each designed as a strict probe targeting a distinct acoustic or linguistic failure axis of Kyutai STT (kyutai/stt-1b-en_fr-trfs) — a streaming bilingual speech-to-text model. Probes cover non-native phonology, fast speech rate, word-level and phrase-level code-switching, intrasentential switching, disfluency with switching, and… See the full description on the dataset page: https://huggingface.co/datasets/Atufa/codeswitch-fr-en-kyutai-stt.audioautomatic-speech-recognitionn<1K0 likes58 downloads7mo agoHugging Face10shangeth /mls-mimi-codes Multilingual LibriSpeech (MLS) — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for Multilingual LibriSpeech — LibriVox audiobooks in 7 non-English languages. English is intentionally excluded. For English Mimi codes, use: shangeth/librispeech-mimi-codes — LibriSpeech (~280k rows, 7 splits) shangeth/libritts-r-mimi-codes — LibriTTS-R (~360k rows, 7 splits, 24 kHz native) shangeth/vctk-mimi-codes — VCTK (~44k rows, 110 speakers w/ accents) shangeth/jenny-mimi-codes — Jenny… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/mls-mimi-codes.tabulartext-to-speech1M<n<10M0 likes58 downloads5mo agoHugging Face11fiewor /gradrai-viva-codeswitch-benchmark GradrAI Viva Code-Switched Oral Benchmark Consented, de-identified classroom-style oral answer clips used to benchmark GradrAI Viva for the Sahara CodeSwitch Africa challenge. Contents metadata.csv / metadata.jsonl: one row per clip. audio/: 16 kHz mono WAV files for Hugging Face dataset preview and ASR reuse. audio_original/: original submitted browser/Opus/WebM audio files. benchmark/: benchmark outputs (results.md, results.json) and manifest used by GradrAI… See the full description on the dataset page: https://huggingface.co/datasets/fiewor/gradrai-viva-codeswitch-benchmark.automatic-speech-recognitionn<1K0 likes57 downloads6d agoHugging Face12shangeth /librispeech-mimi-codes LibriSpeech — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for the LibriSpeech corpus — multi-speaker English audiobook readings from the LibriVox project. This dataset contains codes only, not audio. For waveforms, use any of the LibriSpeech mirrors (e.g. openslr/librispeech_asr); these codes let you skip the ~hours of GPU extraction needed to train Mimi-based speech models. Schema One row per utterance: Column Type Notes id string… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/librispeech-mimi-codes.tabulartext-to-speech100K<n<1M0 likes51 downloads5mo agoHugging Face13vikkyblacq /kare-codeswitch-samples Kare — Code-Switching Illustrative Samples Eight short audio clips demonstrating the English/Yoruba/Hausa/Igbo/Pidgin code-switching that Kare, a voice-first AI health assistant for Nigeria, is built to understand — submitted as part of Kare's entry to the Sahara CodeSwitch Africa Challenge. What this is — and isn't Is: eight original sentences, written by the Kare team specifically for this demonstration, mixing everyday Yoruba / Hausa / Igbo / Nigerian Pidgin… See the full description on the dataset page: https://huggingface.co/datasets/vikkyblacq/kare-codeswitch-samples.audioautomatic-speech-recognitionn<1K0 likes46 downloads7d agoHugging Face141uckyan /code-switch_chunks Dataset Summary This dataset is a curated compilation of SECoMiCSC, DevCECoMiCSC, and BAAI/CS-Dialogue, specifically processed for Code-Switching ASR research. root/ ├── audio/ │ ├── SECoMiCSC/ # Chunked segments from SECoMiCSC │ ├── DevCECoMiCSC/ # Chunked segments from DevCECoMiCSC │ └── CS_Dialogue/ # Extracted <MIX> segments from BAAI/CS-Dialogue ├── metadata.jsonl # Universal index containing paths, transcripts, and metadata └──… See the full description on the dataset page: https://huggingface.co/datasets/1uckyan/code-switch_chunks.audioautomatic-speech-recognition10K<n<100K0 likes36 downloads8mo agoHugging Face15mosesdaudu /switchboard-tierb-codeswitch SwitchBoard Tier B — African code-switched speech 87 consented utterances of intra-sentential code-switching — Nigerian Pidgin, Yorùbá, Hausa and Kiswahili each mixed with English inside a single sentence — recorded from 8 bilingual volunteers at the Deep Learning Indaba 2026, Lagos. Collected for the MLC (Africa) × Intron Agentic Voice AI Challenge as an evaluation set for telco/fintech voice agents. 8.75 minutes total. What this is for Measuring whether a speech… See the full description on the dataset page: https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch.audioautomatic-speech-recognitionn<1K0 likes36 downloads2mo agoHugging Face16Abhisingh-18 /hindi-english-codeswitch-dataset Hindi-English Code-Switch ASR Transcripts Text transcripts and metadata for a large bilingual Hindi-English code-switch ASR training corpus, used to train Abhisingh-18/hindi-english-codeswitch-asr. This release contains transcripts and metadata only — no audio files. Audio was sourced from multiple corpora and institutions and is not redistributed here. Credits Speech data collection and curation credit: SPRING Lab, IIT Madras. Contents File… See the full description on the dataset page: https://huggingface.co/datasets/Abhisingh-18/hindi-english-codeswitch-dataset.textautomatic-speech-recognition1M<n<10M0 likes35 downloads1mo agoHugging Face17shangeth /expresso-mimi-codes Expresso — Mimi Codes (k = 32) Pre-extracted Kyutai Mimi tokens (all 32 codebooks) for both the read and conversational subsets of Expresso. Source audio + transcripts live in shangeth/expresso; this dataset publishes the discrete-token version for training Mimi-based speech models without re-extracting. ⚠️ License: CC-BY-NC-4.0 — non-commercial use only. Why Expresso for Wren? Expresso is the most directly relevant dataset for speech disentanglement research — the… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/expresso-mimi-codes.tabulartext-to-speech10K<n<100K1 likes31 downloads5mo agoHugging Face18shulhaaja /id-en-codeswitch-dataset-alternative Indonesian–English Code-Switching Synthetic Speech Dataset Synthetic speech generated for the undergraduate final project "Handling Code-Switching in Automatic Speech Recognition for Low-Resource Language Pairs: An Indonesian–English Case Study", School of Electrical Engineering and Informatics, Institut Teknologi Bandung. This dataset contains synthetic audio produced from the Indonesian–English code-switching text corpora released in the companion repository below. It was used… See the full description on the dataset page: https://huggingface.co/datasets/shulhaaja/id-en-codeswitch-dataset-alternative.audioautomatic-speech-recognition10K<n<100K0 likes28 downloads1mo agoHugging Face19WTFO /codeswitchinggated WTFO Code-Switching Speech Code-switching speech dataset prepared for automatic speech recognition training. Dataset fields audio: self-contained 16 kHz mono PCM16 FLAC bytes using the Hugging Face Audio feature text: transcript duration: audio duration in seconds (float64) Split summary Split: train Examples: 98,662 Total duration: 567422.698 seconds (157.62 hours) The Parquet shards embed the audio bytes. They do not depend on the original… See the full description on the dataset page: https://huggingface.co/datasets/WTFO/codeswitching.audioautomatic-speech-recognition10K<n<100K0 likes24 downloads1mo agoHugging Face20LauelKills /sugidanon-hil-codeswitch Sugidanon 🎙️ — Code-Switch Hiligaynon Speech Benchmark The first openly-licensed, code-switch-labeled speech benchmark for Hiligaynon (Ilonggo) — a language spoken by 9M+ Filipinos yet nearly invisible to modern speech technology. &nbsp;·&nbsp; Code: https://github.com/Jazztinn/tinig-sa-liwanag &nbsp;·&nbsp; License: CC BY 4.0 Real Hiligaynon-English-Tagalog speech, labeled per word, with a scorer that measures what generic models miss: accuracy at the moment the language… See the full description on the dataset page: https://huggingface.co/datasets/LauelKills/sugidanon-hil-codeswitch.automatic-speech-recognitionn<1K0 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.