datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.code_switch_yodas_zh
Dataset Card for code-switching yodas
This dataset is derived from espnet/yodas, more details can be found here: https://huggingface.co/datasets/espnet/yodas
This is a subset of the zh000 subset of espnet/yodas dataset, which selects videos with Mandarin-English code-switching phenomenon.
Note that code-switching is only gauranteed per video rather than per utterance. Therefore, not every utterance in the dataset contains code-switching.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/code_switch_yodas_zh.code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here.
Code-Switching ASR
Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.text-summarizationasr_codeswitchedquestion-answerasr_codeswitched_dataset
Arabic/English Code-Switched ASR Dataset
Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to
fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical
vocabulary alternate within sentences.
Composition
Source
Description
EJUST custom recordings
Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs)
MohamedRashad/arabic-english-code-switching
~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.naturalnesscafe-algerian-codeswitch-speech
CAFE Algerian Codeswitch Speech
This dataset contains Algerian Arabic and French code-switched speech.
Repository Path: FatimahEmadEldin/cafe-algerian-codeswitch-speech
ne-en-codeswitching-asr-technical-interview
Dataset Summary
This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar.
It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.kazakh-codeswitch-asr
Kazakh Code-Switching ASR Benchmark
A benchmark for evaluating ASR systems on natural Kazakh speech that
code-switches with Russian — the everyday Kazakh–Russian mixing found in
stand-up, interviews and vlogs, not scripted read speech. This is, to our
knowledge, the first speech/ASR resource targeting the Kazakh–Russian
code-switching pair (existing Kazakh–Russian NLP resources are text-only).
Code, scoring harness and full analysis:… See the full description on the dataset page: https://huggingface.co/datasets/Tim2190/kazakh-codeswitch-asr.id-en-codeswitch-dataset-alternative
Indonesian–English Code-Switching Synthetic Speech Dataset
Synthetic speech generated for the undergraduate final project "Handling
Code-Switching in Automatic Speech Recognition for Low-Resource Language
Pairs: An Indonesian–English Case Study", School of Electrical Engineering
and Informatics, Institut Teknologi Bandung.
This dataset contains synthetic audio produced from the Indonesian–English
code-switching text corpora released in the companion repository below. It
was used… See the full description on the dataset page: https://huggingface.co/datasets/shulhaaja/id-en-codeswitch-dataset-alternative.neuromoyo-sahara-codeswitch-benchmark
NEUROMOYO — Sahara CodeSwitch Africa Benchmark
🔗 Live Benchmark Results
Interactive benchmark:
https://www.neuromoyo.app/benchmark
This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech.
🚀 Live NEUROMOYO Demo
Live application:
https://www.neuromoyo.app
The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.vocal-money-codeswitch-asr-benchmark
Vocal Money — Yoruba–English Code-Switched ASR Benchmark
A 210-clip evaluation subset used to benchmark five speech recognition systems on naturally
code-switched Yoruba–English speech, together with the reference transcriptions and the output of
every system on every clip, so that the published results can be recomputed or contradicted.
Produced for the MLC (Africa) × Intron Agentic Voice AI Challenge, Deep Learning Indaba 2026.
Team Vocal Money — Hospice Hounfodji, Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/Kimyayd/vocal-money-codeswitch-asr-benchmark.khmer-english-codeswitch-tts
Khmer–English Code-Switch Synthetic Speech
8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech,
generated with VoxCPM2 from code-switch text
manufactured by confirmed lexical substitution over a Khmer–English parallel corpus.
⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a
TTS model, and the code-switch sentences were manufactured by word substitution — they are not
transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.topic-classificationcodeswitch-pairs-lase
Codeswitch Pairs LASE — training corpus
1118 same-voice cross-script utterance pairs (8 ElevenLabs Multilingual voices × en/hi/te/ta) used to train the LASE r1 speaker encoder.
Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair).
Schema (manifest.jsonl)
{
"voice_id": "21m00Tcm4TlvDq8ikWAM",
"lang": "en | hi | te | ta",
"text": "the prompt text"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase.codeswitch-fr-en-kyutai-stt
Code-Switched French–English STT Probe Dataset
Dataset Summary
This dataset contains 10 audio clips of French–English code-switched speech, each designed as a strict probe targeting a distinct acoustic or linguistic failure axis of Kyutai STT (kyutai/stt-1b-en_fr-trfs) — a streaming bilingual speech-to-text model. Probes cover non-native phonology, fast speech rate, word-level and phrase-level code-switching, intrasentential switching, disfluency with switching, and… See the full description on the dataset page: https://huggingface.co/datasets/Atufa/codeswitch-fr-en-kyutai-stt.kare-codeswitch-samples
Kare — Code-Switching Illustrative Samples
Eight short audio clips demonstrating the English/Yoruba/Hausa/Igbo/Pidgin
code-switching that Kare, a voice-first
AI health assistant for Nigeria, is built to understand — submitted as part
of Kare's entry to the Sahara CodeSwitch Africa Challenge.
What this is — and isn't
Is: eight original sentences, written by the Kare team specifically
for this demonstration, mixing everyday Yoruba / Hausa / Igbo / Nigerian
Pidgin… See the full description on the dataset page: https://huggingface.co/datasets/vikkyblacq/kare-codeswitch-samples.code-switch_chunks
Dataset Summary
This dataset is a curated compilation of SECoMiCSC, DevCECoMiCSC, and BAAI/CS-Dialogue, specifically processed for Code-Switching ASR research.
root/
├── audio/
│ ├── SECoMiCSC/ # Chunked segments from SECoMiCSC
│ ├── DevCECoMiCSC/ # Chunked segments from DevCECoMiCSC
│ └── CS_Dialogue/ # Extracted <MIX> segments from BAAI/CS-Dialogue
├── metadata.jsonl # Universal index containing paths, transcripts, and metadata
└──… See the full description on the dataset page: https://huggingface.co/datasets/1uckyan/code-switch_chunks.Moroccan-Codeswitching
Moroccan Darija Code-Switched Corpus (Sentence-level TSV)
Dataset Summary
This dataset contains sentence/post-level code-switched Moroccan Darija text with a single label per text unit. It is intended to support NLP research on Moroccan Darija (Darija), an under-resourced Arabic variety, and on sentence-level code-switching / language identification in Moroccan online text.
Languages
The corpus may contain Moroccan Darija (often ary) and code-switching with:… See the full description on the dataset page: https://huggingface.co/datasets/samihyounes/Moroccan-Codeswitching.jaen-codeswitch-tts
jaen-codeswitch-tts
Japanese/English code-switch synthetic speech: 10k varied-length conversational utterances, single voice (Qwen3-TTS clone of one English-male reference, spk_male). Per-utterance language routed to the dominant script.
Columns: audio, text, speaker_id, language.
nemo-codeswitch-reasoning-debate
Overview
This is a synthetic, multilingual code-switching dataset. Each record contains:
a realistic user query
a long-form reasoning section
a debate / counterargument section
a concise final_answer
It is designed for experiments in multilingual generation, code-switch robustness, and reasoning/debate style responses.
This snapshot contains 574,977 rows and 10 string columns.
Data provenance
Important:
Verify that your intended usage and redistribution complies… See the full description on the dataset page: https://huggingface.co/datasets/lxyuan/nemo-codeswitch-reasoning-debate.adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-moroccan_darija_prompts & trilingual_codeswitch_chat (augmented)
This dataset consists of short conversational prompts written in Moroccan Darija, covering topics like shopping, social interactions, and daily inquiries. Each entry contains a single prompt with a null completion, indicating it is likely intended for instruction tuning or completion generation tasks. The content… See the full description on the dataset page: https://huggingface.co/datasets/oumayma03/adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented.arazn_codeSwitched_mp3_full_notLowercode-switching-codesaviours-si26-hamzacode-switching-codesaviours-si26-SanaCode-switchSpeechRecognition_NTUML2021
Dataset Card for "Code-switchSpeechRecognition_NTUML2021"
More Information needed
ASR_En_Ar_CodeSwitchingNADI2026_subtask1.3_codeswitched_asr
