CoolFace
Datasetpublic

Muhammad-Ahmad-1263/code-switching-codesaviours-si26-muhammadahmad

Code-Switching Codesaviours SI26 — Muhammad Ahmad Dataset Description This dataset contains 155 naturally occurring Roman Urdu–English code-switched sentences (1,400+ word-level entries), reflecting how Roman Urdu and English are mixed in everyday informal communication by Pakistani speakers online (Twitter/X, WhatsApp, YouTube comments, Reddit). Code-switching — alternating between two or more languages within a single sentence or conversation — is extremely… See the full description on the dataset page: https://huggingface.co/datasets/Muhammad-Ahmad-1263/code-switching-codesaviours-si26-muhammadahmad.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes30downloads
Dataset Card

Code-Switching Codesaviours SI26 — Muhammad Ahmad

Dataset Description

This dataset contains 155 naturally occurring Roman Urdu–English code-switched sentences (1,400+ word-level entries), reflecting how Roman Urdu and English are mixed in everyday informal communication by Pakistani speakers online (Twitter/X, WhatsApp, YouTube comments, Reddit).

Code-switching — alternating between two or more languages within a single sentence or conversation — is extremely common in Pakistani digital communication, but is poorly handled by most existing NLP models, which are trained overwhelmingly on monolingual English or formal Urdu text.

How it was collected

Sentences were sourced by observing genuine patterns of Roman Urdu–English mixing common on Pakistani social media (Twitter/X, r/pakistan, YouTube comments) and everyday informal chat, then written out and hand-labelled word by word. This is a first version aimed at bootstrapping a resource that does not currently exist in a clean, labelled form; a natural next step is expanding it with directly scraped, timestamped social-media posts.

Fields

ColumnDescription
sentenceThe full original mixed-language sentence
wordA single token from that sentence
labelThe language label for that word (see below)

Label meanings

  • URD — Roman Urdu word (Urdu written in Latin script), e.g. bohot, kya, hai
  • ENG — English word, e.g. busy, meeting, honestly
  • MIX — an assimilated loanword that Pakistani speakers use interchangeably in both languages without treating it as a "foreign" word, e.g. mobile, internet, data (as in phone data)

Intended Use

This dataset is intended for:

  • Training/evaluating word-level language identification (LID) models for code-switched text
  • Building or fine-tuning NLU models better suited to Roman Urdu–English text
  • Linguistic research on code-switching patterns in South Asian digital communication

Limitations

  • Sample size (155 sentences) is a starting point, not a comprehensive corpus
  • Sentences were authored/transcribed to reflect real usage patterns rather than scraped in bulk with timestamps/usernames; no personally identifiable information is included
  • Label boundaries between MIX and ENG/URD can be subjective for heavily assimilated loanwords

License

MIT