datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.code_switch_yodas_zh
Dataset Card for code-switching yodas
This dataset is derived from espnet/yodas, more details can be found here: https://huggingface.co/datasets/espnet/yodas
This is a subset of the zh000 subset of espnet/yodas dataset, which selects videos with Mandarin-English code-switching phenomenon.
Note that code-switching is only gauranteed per video rather than per utterance. Therefore, not every utterance in the dataset contains code-switching.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/code_switch_yodas_zh.text-summarizationasr_codeswitchedasr_codeswitched_dataset
Arabic/English Code-Switched ASR Dataset
Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to
fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical
vocabulary alternate within sentences.
Composition
Source
Description
EJUST custom recordings
Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs)
MohamedRashad/arabic-english-code-switching
~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.ne-en-codeswitching-asr-technical-interview
Dataset Summary
This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar.
It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.khmer-english-codeswitch-tts
Khmer–English Code-Switch Synthetic Speech
8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech,
generated with VoxCPM2 from code-switch text
manufactured by confirmed lexical substitution over a Khmer–English parallel corpus.
⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a
TTS model, and the code-switch sentences were manufactured by word substitution — they are not
transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.topic-classificationcodeswitch-fr-en-kyutai-stt
Code-Switched French–English STT Probe Dataset
Dataset Summary
This dataset contains 10 audio clips of French–English code-switched speech, each designed as a strict probe targeting a distinct acoustic or linguistic failure axis of Kyutai STT (kyutai/stt-1b-en_fr-trfs) — a streaming bilingual speech-to-text model. Probes cover non-native phonology, fast speech rate, word-level and phrase-level code-switching, intrasentential switching, disfluency with switching, and… See the full description on the dataset page: https://huggingface.co/datasets/Atufa/codeswitch-fr-en-kyutai-stt.jaen-codeswitch-tts
jaen-codeswitch-tts
Japanese/English code-switch synthetic speech: 10k varied-length conversational utterances, single voice (Qwen3-TTS clone of one English-male reference, spk_male). Per-utterance language routed to the dominant script.
Columns: audio, text, speaker_id, language.
nemo-codeswitch-reasoning-debate
Overview
This is a synthetic, multilingual code-switching dataset. Each record contains:
a realistic user query
a long-form reasoning section
a debate / counterargument section
a concise final_answer
It is designed for experiments in multilingual generation, code-switch robustness, and reasoning/debate style responses.
This snapshot contains 574,977 rows and 10 string columns.
Data provenance
Important:
Verify that your intended usage and redistribution complies… See the full description on the dataset page: https://huggingface.co/datasets/lxyuan/nemo-codeswitch-reasoning-debate.arazn_codeSwitched_mp3_full_notLowerASR_En_Ar_CodeSwitchingNADI2026_subtask1.3_codeswitched_asrCode-Switching-Testafrican-codeswitching
Pan-African Code-Switching Dataset
Built with Adaptive Data by Adaption | Crane AI Labs
Submitted to the Uncharted Data Challenge 2026 by Adaption Labs
Overview
The first open-source, richly-annotated pan-African code-switching dataset covering authentic language mixing patterns across 4 African regions — East Africa, South Africa, and West Africa — in 10 distinct language-pair configurations.
Code-switching (mixing two or more languages in a single utterance) is how… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/african-codeswitching.indic-codeswitched-social-bias
Dataset Description
This dataset contains Indian code-switched sentences — informal text where
speakers fluidly mix English with Indian regional languages — annotated for
social bias across multiple dimensions.
Languages covered include (but may not be limited to):
Hindi-English (Hinglish) — hi-en
Tamil-English (Tanglish) — ta-en
Punjabi-English — pa-en
Bengali-English — bn-en
Gujarati-English — gu-en
Telugu-English — te-en
Kannada-English — kn-en
Marathi-English — mr-en… See the full description on the dataset page: https://huggingface.co/datasets/necrosyth/indic-codeswitched-social-bias.CodeSwitchingSpeechIdentification_ASCENDSinhala-NewsCategory-Codeswitched75tts-jaen-codeswitch-6model
tts-jaen-codeswitch-6model
Japanese-English code-switch text spoken by 6 open-source TTS models, all cloning the SAME single reference voice (en-male). 10 mixed utterances (varied length). Each row: the text + reference audio + one audio column per model.
Model columns: qwen3 · moss · voxcpm · omnivoice · zonos · cosyvoice.
Per-utterance language routed to the dominant script (JA-heavy→Japanese, else English). No quality metrics included.
code-switched-annotatedCodeSwitchingSemanticGrammarAcceptabilityComparison_CSZS-zh-encode-switching-codesaviours-si26-qadeesanoorDataset Summary
This dataset contains Roman Urdu–English code-switched sentences, the way mixed-language text actually gets written in everyday Pakistani texting, tweeting, and casual conversation (e.g. "Aaj ka din bohot busy tha, had 3 meetings back to back"). Each sentence is broken down word-by-word, and every word is tagged with a language label. No existing Roman Urdu NLP resource handles this kind of within-sentence language mixing well — this dataset is a step toward building tools… See the full description on the dataset page: https://huggingface.co/datasets/qadeesanoor/code-switching-codesaviours-si26-qadeesanoor.minerva-ar-en-edu-codeswitch-datasetIMDA_Codeswitch30minerva-ar-en-codeswitch-topic-summarycodeswitching
WTFO Code-Switching Speech
Code-switching speech dataset prepared for automatic speech recognition training.
Dataset fields
audio: self-contained 16 kHz mono PCM16 FLAC bytes using the Hugging Face Audio feature
text: transcript
duration: audio duration in seconds (float64)
Split summary
Split: train
Examples: 98,662
Total duration: 567422.698 seconds (157.62 hours)
The Parquet shards embed the audio bytes. They do not depend on the original… See the full description on the dataset page: https://huggingface.co/datasets/WTFO/codeswitching.punjabi-sentiment-tagger-stratified-codeswitchedCode_switching
🇰🇿 Kazakh-Russian Code-Switching Normalization Dataset
Dataset Summary
Kazakh-Russian Code-Switching Normalization Dataset is a bilingual instruction-following dataset designed for identifying and rewriting Kazakh-Russian mixed-language text into clean Kazakh.
The dataset focuses on informal communication, where Kazakh speakers may naturally mix Russian and Kazakh in one message. Each sample contains a prompt with code-switching, a response that identifies the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Code_switching.arazn_codeSwitched_mp3_full_notLower_notMultiDots
