datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
question-answernaturalnessMoroccan-Codeswitching
Moroccan Darija Code-Switched Corpus (Sentence-level TSV)
Dataset Summary
This dataset contains sentence/post-level code-switched Moroccan Darija text with a single label per text unit. It is intended to support NLP research on Moroccan Darija (Darija), an under-resourced Arabic variety, and on sentence-level code-switching / language identification in Moroccan online text.
Languages
The corpus may contain Moroccan Darija (often ary) and code-switching with:… See the full description on the dataset page: https://huggingface.co/datasets/samihyounes/Moroccan-Codeswitching.code-switching-codesaviours-si26-hamzacode-switching-codesaviours-si26-muhammadahmad
Code-Switching Codesaviours SI26 — Muhammad Ahmad
Dataset Description
This dataset contains 155 naturally occurring Roman Urdu–English
code-switched sentences (1,400+ word-level entries), reflecting how
Roman Urdu and English are mixed in everyday informal communication by
Pakistani speakers online (Twitter/X, WhatsApp, YouTube comments, Reddit).
Code-switching — alternating between two or more languages within a single
sentence or conversation — is extremely… See the full description on the dataset page: https://huggingface.co/datasets/Muhammad-Ahmad-1263/code-switching-codesaviours-si26-muhammadahmad.code-switching-codesaviours-si26-Sanacode-switching-codesaviours-si26-warood
Code-Switching Codesaviours SI-26 Dataset
Dataset Description
This dataset contains 150 naturally occurring Roman Urdu–English code-switched sentences, commonly spoken by Pakistani speakers in casual, everyday communication. Each sentence has been broken down into individual words, and every word is labeled by language.
Roman Urdu–English code-switching is extremely common in Pakistan (spoken/written by an estimated 230+ million people) but is poorly handled by… See the full description on the dataset page: https://huggingface.co/datasets/waroodzkhan/code-switching-codesaviours-si26-warood.code-switching-codesaviours-si26-zainab
Roman Urdu–English Code-Switching Dataset
Dataset Description
This dataset contains 1,901 sentences and 21,370 word-level language labels, built to capture how Roman Urdu and English are naturally mixed together in everyday Pakistani online communication.
Code-switching — blending two languages within a single sentence — is how the vast majority of Pakistanis actually write and speak online, on platforms like Twitter/X, Facebook, YouTube, Reddit, and WhatsApp. A… See the full description on the dataset page: https://huggingface.co/datasets/Zainab-Binte-Khalid/code-switching-codesaviours-si26-zainab.code-switching-codesaviours-si26-Moazam
Roman Urdu-English Code-Switching Dataset
Description
This dataset contains naturally occurring Roman Urdu / English code-switched sentences,
collected to reflect how Pakistani speakers actually communicate online — mixing
Roman Urdu and English within the same sentence (e.g. "Aaj mera mood nahi hai for anything").
Each sentence is broken down word-by-word, with every word labeled by language.
Collection Method
Sentences were collected from a mix of… See the full description on the dataset page: https://huggingface.co/datasets/Moazamzf/code-switching-codesaviours-si26-Moazam.code-switching-codesaviours-si26-MuhammadHassaan
Roman Urdu-English Code-Switching Dataset
Description
This dataset contains 150 sentences that mix Roman Urdu and English.
The purpose of this dataset is to study code-switching between Roman Urdu and English.
Labels
URD: Roman Urdu word
ENG: English word
MIX: Mixed or unclear word
Dataset Format
Each row contains:
sentence
word
label
Collection
The sentences were prepared as natural Roman Urdu-English… See the full description on the dataset page: https://huggingface.co/datasets/Hassaanatif992/code-switching-codesaviours-si26-MuhammadHassaan.code-switching-codesaviours-si26-bilal
Roman Urdu-English Code Switching Dataset
Dataset Description
A manually labeled dataset of 160+ code-switching sentences where Roman Urdu and English are naturally mixed — reflecting how 230 million Pakistanis actually communicate online.
Each word in every sentence is tagged with a language label, making this dataset suitable for token-level language identification and code-switching NLP research.
Label Meanings
Label
Description
Examples… See the full description on the dataset page: https://huggingface.co/datasets/Noisy77/code-switching-codesaviours-si26-bilal.code-switching-codesaviours-si26-Hania-Emaancode-switching-codesaviours-si26-Amnacode-switching-codesaviours-si26-amnaCode Switching Dataset — Code Saviours SI-26 (Amna)
Dataset Description
Roman Urdu mixed with English — e.g. "Aaj mera mood nahi hai for anything" — is how a huge share of Pakistanis actually write online, but almost no existing NLP resource labels this kind of code-switching at the word level. This dataset provides 160 real, naturally occurring mixed-language sentences, tokenised and labelled word-by-word as Roman Urdu, English, or a genuine hybrid blend.
How It Was Collected
Sentences were… See the full description on the dataset page: https://huggingface.co/datasets/AmnaNoor123/code-switching-codesaviours-si26-amna.code-switching-codesaviours-si26-Uswa
Code-Switching Dataset: Roman Urdu ↔ English (Pakistan)
Dataset Description
This dataset contains 191 naturally code-switched Roman Urdu–English sentences** (well above
the 150 minimum) blending Roman Urdu and English, the way Pakistani
speakers actually write online. Every word in every sentence is labelled at
the token level, making this a word-level sequence labelling / language
identification dataset for code-switched text.
Roman Urdu–English mixing is… See the full description on the dataset page: https://huggingface.co/datasets/122Uswa/code-switching-codesaviours-si26-Uswa.code-switching-codesaviours-si26-sheeza
Code Switching NLP Dataset
Dataset Description
This dataset contains 161 code-switching sentences that combine Roman Urdu and English. The dataset represents informal Pakistani online communication, including social media posts, comments, messages, and everyday digital conversations.
Each sentence is labelled at the word level.
Labels
URD: Roman Urdu words
ENG: English words
MIX: Tokens that combine Roman Urdu and English within the same word… See the full description on the dataset page: https://huggingface.co/datasets/sheezariaz2315/code-switching-codesaviours-si26-sheeza.Code-switching-codesaviours-si26-Sumair
Code Switching (Roman Urdu - English) Dataset
1. What it is
This dataset is a specialized NLP corpus created for Token-Classification tasks on bilingual Code-Switching text. It features natural, everyday sentences combining Roman Urdu and English (how Pakistani speakers communicate online on social platforms). Every individual word/token in the dataset is annotated with word-level language identification tags.
2. Collection Method
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sumair-Parveiz/Code-switching-codesaviours-si26-Sumair.code-switching-codesaviours-si26-saima
Code-Switching Urdu-English Dataset
Dataset Description
This dataset is created for Urdu-English code-switching language identification. It contains sentences that include Urdu words, English words, and a small number of mixed-language entries.
The dataset was prepared by collecting and organizing code-switched Urdu-English sentences. Each sentence was divided into individual words, and every word was assigned a language label. The data was then converted into a… See the full description on the dataset page: https://huggingface.co/datasets/Saima-Manzoor/code-switching-codesaviours-si26-saima.code-switching-codesaviours-si26-ushna
Code Switching NLP Dataset (Roman Urdu + English)
Project Overview
This dataset contains natural code-switched sentences combining Roman Urdu and English, collected for Code Saviours SI-26 internship (Week 6).
Dataset Summary
Total Sentences: 175
Total Word Tokens: 1333
Format: CSV (sentence, word, label)
Label Meanings
URD: Roman Urdu words (e.g., aaj, mera, hai)
ENG: English words (e.g., busy, presentation, test)
MIX:… See the full description on the dataset page: https://huggingface.co/datasets/Ushna-Alam219/code-switching-codesaviours-si26-ushna.code-switching-codesaviours-si26-samaika
Code Switching NLP | Code Saviours SI-26 | Samaika
About the Dataset
This dataset contains Roman Urdu and English code-switching sentences collected from social media comments and online content.
The purpose of this dataset is to represent the way Pakistani users naturally mix Roman Urdu and English while communicating online.
Data Collection
The sentences were collected from:
Instagram comments
YouTube comments
Twitter/X
The collected sentences… See the full description on the dataset page: https://huggingface.co/datasets/samaikaimran/code-switching-codesaviours-si26-samaika.code-switching-codesaviours-si26-izza
Code Switching NLP Dataset
Overview
This dataset was created for the Code Saviours SI-26 Week 6 internship project.
It contains Roman Urdu-English code-switching sentences collected from public online reviews and converted into a word-level labeled dataset.
Dataset Statistics
Total Sentences: 150
Total Word Entries: 3911
Labels
URD = Roman Urdu
ENG = English
MIX = Mixed / Other
Files
dataset.csv… See the full description on the dataset page: https://huggingface.co/datasets/izzazahid/code-switching-codesaviours-si26-izza.code-switching-codesaviours-si26-maryam
Roman Urdu & English Code-Switching Dataset
Description
This dataset contains 152 real-world sentences where users naturally mix Roman Urdu and English. It was built to train models to handle how 230 million Pakistanis actually communicate online.
Data Collection
The raw text was scraped directly from natural conversations on Pakistani Twitter/X and anonymized WhatsApp messages.
Label Meanings
Every single word in this dataset has… See the full description on the dataset page: https://huggingface.co/datasets/Maryam657775/code-switching-codesaviours-si26-maryam.code-switching-codesaviours-si26-hadia
Code Switching Dataset
Description
This dataset was created as part of the Code Saviours Summer Internship 2026 (Project 2).
It contains Roman Urdu and English code-switched sentences. Each word is annotated with one of the following labels:
URD – Roman Urdu
ENG – English
MIX – Mixed-language token (if applicable)
Dataset Structure
The dataset contains the following columns:
sentence – Complete code-switched sentence
word – Individual token from… See the full description on the dataset page: https://huggingface.co/datasets/hadia-tech/code-switching-codesaviours-si26-hadia.code-switching-codesaviours-si26-saliha
Code Switching Dataset (SI-26)
This dataset contains code-switching sentences (Roman Urdu and English) annotated at the word level.
Dataset Structure
sentence: Full code-switched sentence
word: Individual tokenized word
label: Token classification tag
URD: Roman Urdu word
ENG: English word
code-switching-codesaviours-si26-eman
Code Switching NLP Dataset
Code Saviours SI-26 | Week 6
Description
This dataset contains Roman Urdu-English code-switching sentences collected for Code Saviours SI-26 Week 6.
The purpose of this dataset is to provide examples of how Pakistani users naturally mix Roman Urdu and English in informal digital communication.
Dataset Contents
The dataset contains:
157 unique sentences
1,593 word-level entries
1,035 URD labels
558 ENG… See the full description on the dataset page: https://huggingface.co/datasets/emanfatimaa05/code-switching-codesaviours-si26-eman.code-switching-codesaviours-si26-humna
Code Switching NLP Dataset — Code Saviours SI-26, Humna Imran
Word-level labeled dataset of naturally code-switched Roman Urdu + English
sentences, as commonly written by Pakistani social media users.
What it is
220 sentences (2,490 word-level rows). Each word is labeled URD (Roman Urdu),
ENG (English), or MIX (hybrid token, e.g. hyphenated compounds).
How it was built
Source sentences come from Smat26/Roman-Urdu-Dataset (GitHub), a public… See the full description on the dataset page: https://huggingface.co/datasets/hamnaheh/code-switching-codesaviours-si26-humna.code-switching-codesaviours-si26-qandeel
Code Switching NLP Dataset | Code Saviours SI-26
Dataset Description
A word-level labelled Roman Urdu–English code-switching dataset (200 sentences,
3182 word entries). Labels: URD (Roman Urdu), ENG (English), MIX (nativized loanword).
Source
Sentences filtered from the Roman Urdu Data Set (Sharf, 2017, UCI ML Repository,
CC BY 4.0), originally collected from e-commerce reviews, Facebook comments, and
Twitter posts. Filtered for genuine… See the full description on the dataset page: https://huggingface.co/datasets/qandeelasim13/code-switching-codesaviours-si26-qandeel.code-switching-codesaviours-si26-usama
Code Switching Dataset
Description
This dataset contains Roman Urdu and English mixed sentences collected during the Code Saviours ML/AI Internship (SI-26).
The purpose of this dataset is to help train NLP models that can understand code-switched language used by Pakistani people in daily conversations.
Dataset Details
Total sentences: 150+
Language: Roman Urdu + English
Format: CSV
Columns:
sentence
word
label
Label Meanings… See the full description on the dataset page: https://huggingface.co/datasets/Usamasarfraz/code-switching-codesaviours-si26-usama.code-switching-codesaviours-si26-halimaCode Switching Dataset
Description
This dataset was created as part of the Code Saviours SI-26 Project 2. It contains 150 Roman Urdu–English code-switching sentences. Each sentence is split into individual words, and every word is labeled according to the language it belongs to. The dataset is intended for educational and research purposes in Natural Language Processing (NLP), especially for language identification and code-switching detection.
Data Collection
The sentences were collected from… See the full description on the dataset page: https://huggingface.co/datasets/halimajaved592/code-switching-codesaviours-si26-halima.code-switching-codesaviours-si26-fatima
Code-Switching Dataset — Roman Urdu / English
What is this?
A dataset of 200 sentences that naturally mix Roman Urdu and English,
as commonly spoken/written by Pakistanis on social media, chat apps,
and daily conversation.
How it was collected
Sentences were collected from Pakistani Twitter/X, Reddit (r/pakistan),
YouTube comments, WhatsApp messages, and Facebook public pages, focusing
on naturally occurring Roman Urdu-English code-switching.… See the full description on the dataset page: https://huggingface.co/datasets/Fatimasajid/code-switching-codesaviours-si26-fatima.
