datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Codemixed_New
Codemixed ASR Dataset
Unified collection of code-mixed ASR datasets.
ASR_Code_Switch
ASR Code-Switching Benchmark
A curated benchmark of 1,200 code-switching utterances (300 per language pair)
for evaluating commercial ASR systems on multilingual speech with intra-sentential
language switching.
Paper
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
arXiv link
Language pairs
Split
Language pair
Samples
Scripts
egyptian_arabic_english
Egyptian Arabic–English
300
Arabic + Latin… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.Ghana_English-Twi_Code-switching_Speech
Dataset Card for KasaSpeech
Dataset Summary
KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi.
The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios
With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/Kennethdot/Ghana_English-Twi_Code-switching_Speech.arabic-english-code-switching
Thanks to ahmedheakl/arzen-llm-speech-ds as this dataset was built upon it ✨
The dataset was constructed using ahmed's dataset and different videos from the youtube. The scraped data doubled the initial dataset size after deduplication and cleaning.
Citation
If you use this dataset, please cite it as follows:
@misc{rashad2024arabic,
author = {Mohamed Rashad},
title = {arabic-english-code-switching},
year = {2024},
publisher = {Hugging Face},
url =… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-english-code-switching.code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here.
Code-Switching ASR
Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.asr_codeswitched_dataset
Arabic/English Code-Switched ASR Dataset
Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to
fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical
vocabulary alternate within sentences.
Composition
Source
Description
EJUST custom recordings
Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs)
MohamedRashad/arabic-english-code-switching
~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.arabic-english-code-switching-synthetic-asr
Synthetic Arabic-English Code-Switched Speech for ASR
This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository.
Configurations
Configuration
Train
Test
Publication status
synthetic
8,655
962
Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.Ghana_English-Twi_Code-switching_Speech
Dataset Card for KasaSpeech
Dataset Summary
KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi.
The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios
With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/Ghana_English-Twi_Code-switching_Speech.english-x-code-switching
Synthetic English Code-Switching Evaluation Set
This dataset contains synthetic long-form English code-switching audio samples built from ML-SUPERB hybrid data.
Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row.
Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching.english-en-x-code-switching-main-lang
English EN-X Code-Switching Main-Language
This dataset contains synthetic English-plus-one-language code-switching samples built from FLEURS.
Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly.
Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang.Ghana_English-Twi_Code-switching_Speech
Dataset Card for KasaSpeech
Dataset Summary
KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi.
The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios
With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech.ne-en-codeswitching-asr-technical-interview
Dataset Summary
This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar.
It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.kazakh-codeswitch-asr
Kazakh Code-Switching ASR Benchmark
A benchmark for evaluating ASR systems on natural Kazakh speech that
code-switches with Russian — the everyday Kazakh–Russian mixing found in
stand-up, interviews and vlogs, not scripted read speech. This is, to our
knowledge, the first speech/ASR resource targeting the Kazakh–Russian
code-switching pair (existing Kazakh–Russian NLP resources are text-only).
Code, scoring harness and full analysis:… See the full description on the dataset page: https://huggingface.co/datasets/Tim2190/kazakh-codeswitch-asr.Ar-En-Code-Switching-Textual-Dataset
ArE-CSTD: Arabic-English Code-Switching Textual Dataset
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "ArE-CSTD" dataset, which stands for "Arabic-English Code-Switching Textual Dataset”.
This dataset contains 330K dialectical Arabic-English code-swithing sentences generated by the large language model GPT-4.
TXT Files
There are 6 txt files. 2 files for Modern Standard Arabic(MSA) train and… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Ar-En-Code-Switching-Textual-Dataset.khmer-english-codeswitch-tts
Khmer–English Code-Switch Synthetic Speech
8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech,
generated with VoxCPM2 from code-switch text
manufactured by confirmed lexical substitution over a Khmer–English parallel corpus.
⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a
TTS model, and the code-switch sentences were manufactured by word substitution — they are not
transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.chichewa_english_code_switch_dataset
Chichewa-English Code-Switched Speech Dataset
Dataset Description
A speech dataset containing 247 audio recordings of Chichewa-English code-switched phrases. Code-switching — the practice of alternating between two or more languages within a single conversation — is extremely common in Malawi and across multilingual African communities. This dataset captures that natural linguistic behavior in spoken form.
Purpose
This dataset is designed to support research… See the full description on the dataset page: https://huggingface.co/datasets/suru8-ai/chichewa_english_code_switch_dataset.Ghana_English-Twi_Code-switching_Speech-ipa
KasaSpeech English–Twi Code-Switching Speech — IPA
A phonemised version of
ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech
(KasaSpeech) with one added column: ipa.
Every other column — audio included — is carried over byte-for-byte, and row
order is unchanged, so this dataset aligns one-to-one with the original.
The ipa column
Each transcript is converted to a phoneme sequence with
ghanag2p-uni, the Twi-only
grapheme-to-phoneme library built on
ghana-g2p… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech-ipa.vocal-money-codeswitch-asr-benchmark
Vocal Money — Yoruba–English Code-Switched ASR Benchmark
A 210-clip evaluation subset used to benchmark five speech recognition systems on naturally
code-switched Yoruba–English speech, together with the reference transcriptions and the output of
every system on every clip, so that the published results can be recomputed or contradicted.
Produced for the MLC (Africa) × Intron Agentic Voice AI Challenge, Deep Learning Indaba 2026.
Team Vocal Money — Hospice Hounfodji, Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/Kimyayd/vocal-money-codeswitch-asr-benchmark.CoDeTT
CoDeTT: A Context-Aware Decision Benchmark for Turn-Taking Evaluation
🌐 Dataset Summary
CoDeTT is a benchmark dataset for turn-taking decision evaluation in full-duplex spoken dialogue systems.It evaluates not only what action a model should take at the current moment, but also whether the underlying semantic intent is aligned.
Core action space (4 classes):
Maintain
Stop & Listen
Takeover
Dismiss
Fine-grained intent space:
14 scenario labels across two system states:… See the full description on the dataset page: https://huggingface.co/datasets/YingaoWang-casia/CoDeTT.ljspeech-mimi-codes
LJSpeech — Mimi Codes
Pre-extracted Kyutai Mimi neural-codec tokens for the
LJSpeech corpus — 13,100 English utterances
from a single female speaker reading public-domain audiobook passages (~24 hours).
This dataset contains codes only, not audio. For waveforms, go to the original LJSpeech
release; these codes are designed to be loaded alongside it for training Mimi-based speech
models without paying the ~1 hour of GPU extraction cost.
Schema
One row per utterance:… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/ljspeech-mimi-codes.slue-sqa-code-l22-c500
SLUE-SQA-5 HuBERT Layer-22 K=500 Discrete Units
Packed discrete-unit files for SpeechGR experiments on SLUE-SQA-5.
The units were produced with HuBERT layer 22 and a K=500 k-means model, then deduplicated with consecutive counts retained. The packed format avoids one .code and .cnt file per utterance.
Files
documents.npz: packed document/passage units
train.npz: packed train question units
validation.npz: packed validation question units
test.npz: packed test question… See the full description on the dataset page: https://huggingface.co/datasets/dodofk/slue-sqa-code-l22-c500.khmer-english-codeswitch-tts-llm
Khmer–English Code-Switch Synthetic Speech (LLM-authored)
19,825 utterances / 21.7 hours of synthetic Khmer–English code-switched speech at
16 kHz, generated with VoxCPM2 from code-switch
sentences written by an LLM and validated programmatically.
⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a TTS
model, and every sentence was written by a language model — they are not transcripts of anything
a person said. It is intended as an… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts-llm.Saudilang-Code-Switch-Corpus
SCC - Saudilang Code-Switch Corpus
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”.
This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.codeswitch-fr-en-kyutai-stt
Code-Switched French–English STT Probe Dataset
Dataset Summary
This dataset contains 10 audio clips of French–English code-switched speech, each designed as a strict probe targeting a distinct acoustic or linguistic failure axis of Kyutai STT (kyutai/stt-1b-en_fr-trfs) — a streaming bilingual speech-to-text model. Probes cover non-native phonology, fast speech rate, word-level and phrase-level code-switching, intrasentential switching, disfluency with switching, and… See the full description on the dataset page: https://huggingface.co/datasets/Atufa/codeswitch-fr-en-kyutai-stt.mls-mimi-codes
Multilingual LibriSpeech (MLS) — Mimi Codes
Pre-extracted Kyutai Mimi neural-codec tokens
for Multilingual LibriSpeech —
LibriVox audiobooks in 7 non-English languages.
English is intentionally excluded. For English Mimi codes, use:
shangeth/librispeech-mimi-codes — LibriSpeech (~280k rows, 7 splits)
shangeth/libritts-r-mimi-codes — LibriTTS-R (~360k rows, 7 splits, 24 kHz native)
shangeth/vctk-mimi-codes — VCTK (~44k rows, 110 speakers w/ accents)
shangeth/jenny-mimi-codes — Jenny… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/mls-mimi-codes.Korean-Japanese-Code-Switching-Speech
Korean-Japanese-Code-Switching-Speech
This dataset contains Korean-Japanese code-switching speech recordings with sentence-level transcriptions. It was introduced in the paper Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs.
Since there is an extremely small amount of Korean-Japanese code-switching data available, this dataset was designed to be used as a small-scale evaluation dataset.
The dataset consists of code-switching recordings… See the full description on the dataset page: https://huggingface.co/datasets/thetaone-ai/Korean-Japanese-Code-Switching-Speech.BSCs_Code_Switching_CA-ES_ASR_TestThe BSC's Code-Switching Catalan-Spanish ASR Test is a speech dataset of 4 hours and 9 minutes. It consists of carefully selected recordings that feature code-switching between Catalan and Spanish. This dataset is designed to be a test set for Catalan ASR systems that need to handle code-switching to Spanish.gradrai-viva-codeswitch-benchmark
GradrAI Viva Code-Switched Oral Benchmark
Consented, de-identified classroom-style oral answer clips used to benchmark GradrAI Viva for the Sahara CodeSwitch Africa challenge.
Contents
metadata.csv / metadata.jsonl: one row per clip.
audio/: 16 kHz mono WAV files for Hugging Face dataset preview and ASR reuse.
audio_original/: original submitted browser/Opus/WebM audio files.
benchmark/: benchmark outputs (results.md, results.json) and manifest used by GradrAI… See the full description on the dataset page: https://huggingface.co/datasets/fiewor/gradrai-viva-codeswitch-benchmark.
