Muhammad-Ahmad-1263/code-switching-codesaviours-si26-muhammadahmad
Code-Switching Codesaviours SI26 — Muhammad Ahmad Dataset Description This dataset contains 155 naturally occurring Roman Urdu–English code-switched sentences (1,400+ word-level entries), reflecting how Roman Urdu and English are mixed in everyday informal communication by Pakistani speakers online (Twitter/X, WhatsApp, YouTube comments, Reddit). Code-switching — alternating between two or more languages within a single sentence or conversation — is extremely… See the full description on the dataset page: https://huggingface.co/datasets/Muhammad-Ahmad-1263/code-switching-codesaviours-si26-muhammadahmad.
Code-Switching Codesaviours SI26 — Muhammad Ahmad
Dataset Description
This dataset contains 155 naturally occurring Roman Urdu–English code-switched sentences (1,400+ word-level entries), reflecting how Roman Urdu and English are mixed in everyday informal communication by Pakistani speakers online (Twitter/X, WhatsApp, YouTube comments, Reddit).
Code-switching — alternating between two or more languages within a single sentence or conversation — is extremely common in Pakistani digital communication, but is poorly handled by most existing NLP models, which are trained overwhelmingly on monolingual English or formal Urdu text.
How it was collected
Sentences were sourced by observing genuine patterns of Roman Urdu–English mixing common on Pakistani social media (Twitter/X, r/pakistan, YouTube comments) and everyday informal chat, then written out and hand-labelled word by word. This is a first version aimed at bootstrapping a resource that does not currently exist in a clean, labelled form; a natural next step is expanding it with directly scraped, timestamped social-media posts.
Fields
Label meanings
- URD — Roman Urdu word (Urdu written in Latin script), e.g.
bohot,kya,hai - ENG — English word, e.g.
busy,meeting,honestly - MIX — an assimilated loanword that Pakistani speakers use interchangeably in both languages without treating it as a "foreign" word, e.g.
mobile,internet,data(as in phone data)
Intended Use
This dataset is intended for:
- Training/evaluating word-level language identification (LID) models for code-switched text
- Building or fine-tuning NLU models better suited to Roman Urdu–English text
- Linguistic research on code-switching patterns in South Asian digital communication
Limitations
- Sample size (155 sentences) is a starting point, not a comprehensive corpus
- Sentences were authored/transcribed to reflect real usage patterns rather than scraped in bulk with timestamps/usernames; no personally identifiable information is included
- Label boundaries between MIX and ENG/URD can be subjective for heavily assimilated loanwords
License
MIT
