datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.text-summarizationne-en-codeswitching-asr-technical-interview
Dataset Summary
This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar.
It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.topic-classificationASR_En_Ar_CodeSwitchingCodeSwitchingSpeechIdentification_ASCENDafrican-codeswitching
Pan-African Code-Switching Dataset
Built with Adaptive Data by Adaption | Crane AI Labs
Submitted to the Uncharted Data Challenge 2026 by Adaption Labs
Overview
The first open-source, richly-annotated pan-African code-switching dataset covering authentic language mixing patterns across 4 African regions — East Africa, South Africa, and West Africa — in 10 distinct language-pair configurations.
Code-switching (mixing two or more languages in a single utterance) is how… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/african-codeswitching.Code-Switching-Testcodeswitching
WTFO Code-Switching Speech
Code-switching speech dataset prepared for automatic speech recognition training.
Dataset fields
audio: self-contained 16 kHz mono PCM16 FLAC bytes using the Hugging Face Audio feature
text: transcript
duration: audio duration in seconds (float64)
Split summary
Split: train
Examples: 98,662
Total duration: 567422.698 seconds (157.62 hours)
The Parquet shards embed the audio bytes. They do not depend on the original… See the full description on the dataset page: https://huggingface.co/datasets/WTFO/codeswitching.CodeSwitchingSemanticGrammarAcceptabilityComparison_CSZS-zh-encode-switching-codesaviours-si26-qadeesanoorDataset Summary
This dataset contains Roman Urdu–English code-switched sentences, the way mixed-language text actually gets written in everyday Pakistani texting, tweeting, and casual conversation (e.g. "Aaj ka din bohot busy tha, had 3 meetings back to back"). Each sentence is broken down word-by-word, and every word is tagged with a language label. No existing Roman Urdu NLP resource handles this kind of within-sentence language mixing well — this dataset is a step toward building tools… See the full description on the dataset page: https://huggingface.co/datasets/qadeesanoor/code-switching-codesaviours-si26-qadeesanoor.Code_switching
🇰🇿 Kazakh-Russian Code-Switching Normalization Dataset
Dataset Summary
Kazakh-Russian Code-Switching Normalization Dataset is a bilingual instruction-following dataset designed for identifying and rewriting Kazakh-Russian mixed-language text into clean Kazakh.
The dataset focuses on informal communication, where Kazakh speakers may naturally mix Russian and Kazakh in one message. Each sample contains a prompt with code-switching, a response that identifies the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Code_switching.
