datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.saudi-english-code-switching-datasetko_commongen_v2_code_switching
🇰🇷🇺🇸🇯🇵🇨🇳🇪🇸 KoCommonGEN v2 Code-switching
This KoCommonGEN v2 Code-switching dataset consists of 99 samples for numerical commonsense reasoning, which were created relying on machine translation.
The dataset can be found on Hugging Face at: nlpai-lab/ko_commongen_v2_code_switching
This dataset contains code-switching data for the following languages:
Korean (korean)
English (english)
Japanese (japan)
Chinese (china)
Spanish (espanol)
(The code-switching data relies on… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/ko_commongen_v2_code_switching.arabic-english-code-switching
Thanks to ahmedheakl/arzen-llm-speech-ds as this dataset was built upon it ✨
The dataset was constructed using ahmed's dataset and different videos from the youtube. The scraped data doubled the initial dataset size after deduplication and cleaning.
Citation
If you use this dataset, please cite it as follows:
@misc{rashad2024arabic,
author = {Mohamed Rashad},
title = {arabic-english-code-switching},
year = {2024},
publisher = {Hugging Face},
url =… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-english-code-switching.Ghana_English-Twi_Code-switching_Speech
Dataset Card for KasaSpeech
Dataset Summary
KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi.
The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios
With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/Kennethdot/Ghana_English-Twi_Code-switching_Speech.text-summarizationarabic-english-code-switching-synthetic-asr
Synthetic Arabic-English Code-Switched Speech for ASR
This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository.
Configurations
Configuration
Train
Test
Publication status
synthetic
8,655
962
Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.Ghana_English-Twi_Code-switching_Speech
Dataset Card for KasaSpeech
Dataset Summary
KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi.
The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios
With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/Ghana_English-Twi_Code-switching_Speech.english-x-code-switching
Synthetic English Code-Switching Evaluation Set
This dataset contains synthetic long-form English code-switching audio samples built from ML-SUPERB hybrid data.
Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row.
Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching.english-en-x-code-switching-main-lang
English EN-X Code-Switching Main-Language
This dataset contains synthetic English-plus-one-language code-switching samples built from FLEURS.
Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly.
Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang.Ghana_English-Twi_Code-switching_Speech
Dataset Card for KasaSpeech
Dataset Summary
KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi.
The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios
With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech.ne-en-codeswitching-asr-technical-interview
Dataset Summary
This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar.
It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.fleurs_code_switching_test
FLEURS Code-Switching Evaluation Set
Dataset Summary
This dataset is a synthetic code-switching evaluation set built from the google/fleurs corpus.Each sample is a single long-form audio sequence (minimum 5 minutes by default) composed by concatenating short utterances from multiple languages.
The goal is to provide a controlled benchmark for testing ASR robustness when language switches happen frequently inside one recording.
How The Dataset Was Curated… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/fleurs_code_switching_test.Ghana_English-Twi_Code-switching_Speech-ipa
KasaSpeech English–Twi Code-Switching Speech — IPA
A phonemised version of
ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech
(KasaSpeech) with one added column: ipa.
Every other column — audio included — is carried over byte-for-byte, and row
order is unchanged, so this dataset aligns one-to-one with the original.
The ipa column
Each transcript is converted to a phoneme sequence with
ghanag2p-uni, the Twi-only
grapheme-to-phoneme library built on
ghana-g2p… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech-ipa.Korean-Japanese-Code-Switching-Speech
Korean-Japanese-Code-Switching-Speech
This dataset contains Korean-Japanese code-switching speech recordings with sentence-level transcriptions. It was introduced in the paper Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs.
Since there is an extremely small amount of Korean-Japanese code-switching data available, this dataset was designed to be used as a small-scale evaluation dataset.
The dataset consists of code-switching recordings… See the full description on the dataset page: https://huggingface.co/datasets/thetaone-ai/Korean-Japanese-Code-Switching-Speech.topic-classificationASR_En_Ar_CodeSwitchingenglish-x-code-switching-samples
Synthetic English Code-Switching Evaluation Set Samples
This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset.
Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row.
Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.CodeSwitchingSpeechIdentification_ASCENDenglish-en-x-code-switching-main-lang-samples-merged
English EN-X Code-Switching Main-Language Merged Samples
This dataset contains contiguous same-language segments from the paired mixed dataset.
Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly.
Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples-merged.arabic-english-code-switching-text
Arabic-English Code-Switching Dataset (Text Only)
This dataset is a text-only version of Arabic-English Code-Switching Dataset dataset,
created by this notebook.
Changes Made
Extracted only the text column from the original dataset.
Usage
from datasets import load_dataset
dataset = load_dataset("MagedSaeed/arabic-english-code-switching-text")
Citation
Please reference/cite the original dataset when using this data.
african-codeswitching
Pan-African Code-Switching Dataset
Built with Adaptive Data by Adaption | Crane AI Labs
Submitted to the Uncharted Data Challenge 2026 by Adaption Labs
Overview
The first open-source, richly-annotated pan-African code-switching dataset covering authentic language mixing patterns across 4 African regions — East Africa, South Africa, and West Africa — in 10 distinct language-pair configurations.
Code-switching (mixing two or more languages in a single utterance) is how… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/african-codeswitching.Code-Switching-Testcode-switching-codesaviours-si26-qadeesanoorDataset Summary
This dataset contains Roman Urdu–English code-switched sentences, the way mixed-language text actually gets written in everyday Pakistani texting, tweeting, and casual conversation (e.g. "Aaj ka din bohot busy tha, had 3 meetings back to back"). Each sentence is broken down word-by-word, and every word is tagged with a language label. No existing Roman Urdu NLP resource handles this kind of within-sentence language mixing well — this dataset is a step toward building tools… See the full description on the dataset page: https://huggingface.co/datasets/qadeesanoor/code-switching-codesaviours-si26-qadeesanoor.english-en-x-code-switching-main-lang-samples
English EN-X Code-Switching Main-Language Samples
This dataset contains the individual full FLEURS utterance chunks used to build the paired mixed dataset.
Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly.
Each selected utterance is RMS-normalized to -20.0… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples.training_dataset_8K_code-switchingcodeswitching
WTFO Code-Switching Speech
Code-switching speech dataset prepared for automatic speech recognition training.
Dataset fields
audio: self-contained 16 kHz mono PCM16 FLAC bytes using the Hugging Face Audio feature
text: transcript
duration: audio duration in seconds (float64)
Split summary
Split: train
Examples: 98,662
Total duration: 567422.698 seconds (157.62 hours)
The Parquet shards embed the audio bytes. They do not depend on the original… See the full description on the dataset page: https://huggingface.co/datasets/WTFO/codeswitching.Dialectal_Speech_Code_SwitchingCodeSwitchingSemanticGrammarAcceptabilityComparison_CSZS-zh-enCode_switching
🇰🇿 Kazakh-Russian Code-Switching Normalization Dataset
Dataset Summary
Kazakh-Russian Code-Switching Normalization Dataset is a bilingual instruction-following dataset designed for identifying and rewriting Kazakh-Russian mixed-language text into clean Kazakh.
The dataset focuses on informal communication, where Kazakh speakers may naturally mix Russian and Kazakh in one message. Each sample contains a prompt with code-switching, a response that identifies the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Code_switching.
