datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here.
Code-Switching ASR
Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.text-summarizationquestion-answernaturalnessne-en-codeswitching-asr-technical-interview
Dataset Summary
This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar.
It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.CodeSwitching-TE-ENAnCora Catalan NER.
This is a dataset for Named Eentity Reacognition (NER) from Ancora corpus adapted for
Machine Learning and Language Model evaluation purposes.
Since multiwords (including Named Entites) in the original Ancora corpus are aggregated as
a single lexical item using underscores (e.g. "Ajuntament_de_Barcelona")
we splitted them to align with word-per-line format, and added conventional Begin-Inside-Outside (IOB)
tags to mark and classify Named Entites.
We did not filter out the different categories of NEs from Ancora (weak and strong).
We did 6 minor edits by hand.
AnCora corpus is used under [CC-by] (https://creativecommons.org/licenses/by/4.0/) licence.
This dataset was developed by BSC TeMU as part of the AINA project, and to enrich the Catalan Language Understanding Benchmark (CLUB).topic-classificationcode-switching-codesaviours-si26-hamzaMoroccan-Codeswitching
Moroccan Darija Code-Switched Corpus (Sentence-level TSV)
Dataset Summary
This dataset contains sentence/post-level code-switched Moroccan Darija text with a single label per text unit. It is intended to support NLP research on Moroccan Darija (Darija), an under-resourced Arabic variety, and on sentence-level code-switching / language identification in Moroccan online text.
Languages
The corpus may contain Moroccan Darija (often ary) and code-switching with:… See the full description on the dataset page: https://huggingface.co/datasets/samihyounes/Moroccan-Codeswitching.code-switching-codesaviours-si26-muhammadahmad
Code-Switching Codesaviours SI26 — Muhammad Ahmad
Dataset Description
This dataset contains 155 naturally occurring Roman Urdu–English
code-switched sentences (1,400+ word-level entries), reflecting how
Roman Urdu and English are mixed in everyday informal communication by
Pakistani speakers online (Twitter/X, WhatsApp, YouTube comments, Reddit).
Code-switching — alternating between two or more languages within a single
sentence or conversation — is extremely… See the full description on the dataset page: https://huggingface.co/datasets/Muhammad-Ahmad-1263/code-switching-codesaviours-si26-muhammadahmad.ASR_En_Ar_CodeSwitchingcode-switching-codesaviours-si26-Sanacode-switching-codesaviours-si26-warood
Code-Switching Codesaviours SI-26 Dataset
Dataset Description
This dataset contains 150 naturally occurring Roman Urdu–English code-switched sentences, commonly spoken by Pakistani speakers in casual, everyday communication. Each sentence has been broken down into individual words, and every word is labeled by language.
Roman Urdu–English code-switching is extremely common in Pakistan (spoken/written by an estimated 230+ million people) but is poorly handled by… See the full description on the dataset page: https://huggingface.co/datasets/waroodzkhan/code-switching-codesaviours-si26-warood.CodeSwitchingSpeechIdentification_ASCENDcode-switching-codesaviours-si26-bilal
Roman Urdu-English Code Switching Dataset
Dataset Description
A manually labeled dataset of 160+ code-switching sentences where Roman Urdu and English are naturally mixed — reflecting how 230 million Pakistanis actually communicate online.
Each word in every sentence is tagged with a language label, making this dataset suitable for token-level language identification and code-switching NLP research.
Label Meanings
Label
Description
Examples… See the full description on the dataset page: https://huggingface.co/datasets/Noisy77/code-switching-codesaviours-si26-bilal.code-switching-codesaviours-si26-tehreemcode-switching-codesaviours-si26-amnaCode Switching Dataset — Code Saviours SI-26 (Amna)
Dataset Description
Roman Urdu mixed with English — e.g. "Aaj mera mood nahi hai for anything" — is how a huge share of Pakistanis actually write online, but almost no existing NLP resource labels this kind of code-switching at the word level. This dataset provides 160 real, naturally occurring mixed-language sentences, tokenised and labelled word-by-word as Roman Urdu, English, or a genuine hybrid blend.
How It Was Collected
Sentences were… See the full description on the dataset page: https://huggingface.co/datasets/AmnaNoor123/code-switching-codesaviours-si26-amna.Code-Switching_English-Hindi_Syntheticcodeswitching
WTFO Code-Switching Speech
Code-switching speech dataset prepared for automatic speech recognition training.
Dataset fields
audio: self-contained 16 kHz mono PCM16 FLAC bytes using the Hugging Face Audio feature
text: transcript
duration: audio duration in seconds (float64)
Split summary
Split: train
Examples: 98,662
Total duration: 567422.698 seconds (157.62 hours)
The Parquet shards embed the audio bytes. They do not depend on the original… See the full description on the dataset page: https://huggingface.co/datasets/WTFO/codeswitching.code-switching-codesaviours-si26-zainab
Roman Urdu–English Code-Switching Dataset
Dataset Description
This dataset contains 1,901 sentences and 21,370 word-level language labels, built to capture how Roman Urdu and English are naturally mixed together in everyday Pakistani online communication.
Code-switching — blending two languages within a single sentence — is how the vast majority of Pakistanis actually write and speak online, on platforms like Twitter/X, Facebook, YouTube, Reddit, and WhatsApp. A… See the full description on the dataset page: https://huggingface.co/datasets/Zainab-Binte-Khalid/code-switching-codesaviours-si26-zainab.code-switching-codesaviours-si26-MuhammadHassaan
Roman Urdu-English Code-Switching Dataset
Description
This dataset contains 150 sentences that mix Roman Urdu and English.
The purpose of this dataset is to study code-switching between Roman Urdu and English.
Labels
URD: Roman Urdu word
ENG: English word
MIX: Mixed or unclear word
Dataset Format
Each row contains:
sentence
word
label
Collection
The sentences were prepared as natural Roman Urdu-English… See the full description on the dataset page: https://huggingface.co/datasets/Hassaanatif992/code-switching-codesaviours-si26-MuhammadHassaan.african-codeswitching
Pan-African Code-Switching Dataset
Built with Adaptive Data by Adaption | Crane AI Labs
Submitted to the Uncharted Data Challenge 2026 by Adaption Labs
Overview
The first open-source, richly-annotated pan-African code-switching dataset covering authentic language mixing patterns across 4 African regions — East Africa, South Africa, and West Africa — in 10 distinct language-pair configurations.
Code-switching (mixing two or more languages in a single utterance) is how… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/african-codeswitching.code-switching-codesaviours-si26-Hania-Emaancode-switching-codesaviours-si26-Moazam
Roman Urdu-English Code-Switching Dataset
Description
This dataset contains naturally occurring Roman Urdu / English code-switched sentences,
collected to reflect how Pakistani speakers actually communicate online — mixing
Roman Urdu and English within the same sentence (e.g. "Aaj mera mood nahi hai for anything").
Each sentence is broken down word-by-word, with every word labeled by language.
Collection Method
Sentences were collected from a mix of… See the full description on the dataset page: https://huggingface.co/datasets/Moazamzf/code-switching-codesaviours-si26-Moazam.code-switching-codesaviours-si26-sheeza
Code Switching NLP Dataset
Dataset Description
This dataset contains 161 code-switching sentences that combine Roman Urdu and English. The dataset represents informal Pakistani online communication, including social media posts, comments, messages, and everyday digital conversations.
Each sentence is labelled at the word level.
Labels
URD: Roman Urdu words
ENG: English words
MIX: Tokens that combine Roman Urdu and English within the same word… See the full description on the dataset page: https://huggingface.co/datasets/sheezariaz2315/code-switching-codesaviours-si26-sheeza.code-switching-codesaviours-si26-qadeesanoorDataset Summary
This dataset contains Roman Urdu–English code-switched sentences, the way mixed-language text actually gets written in everyday Pakistani texting, tweeting, and casual conversation (e.g. "Aaj ka din bohot busy tha, had 3 meetings back to back"). Each sentence is broken down word-by-word, and every word is tagged with a language label. No existing Roman Urdu NLP resource handles this kind of within-sentence language mixing well — this dataset is a step toward building tools… See the full description on the dataset page: https://huggingface.co/datasets/qadeesanoor/code-switching-codesaviours-si26-qadeesanoor.Code-Switching-Testcode-switching-codesaviours-si26-Amnacode-switching-codesaviours-si26-samaika
Code Switching NLP | Code Saviours SI-26 | Samaika
About the Dataset
This dataset contains Roman Urdu and English code-switching sentences collected from social media comments and online content.
The purpose of this dataset is to represent the way Pakistani users naturally mix Roman Urdu and English while communicating online.
Data Collection
The sentences were collected from:
Instagram comments
YouTube comments
Twitter/X
The collected sentences… See the full description on the dataset page: https://huggingface.co/datasets/samaikaimran/code-switching-codesaviours-si26-samaika.
