CoolFace
Datasetpublic

samaikaimran/code-switching-codesaviours-si26-samaika

Code Switching NLP | Code Saviours SI-26 | Samaika About the Dataset This dataset contains Roman Urdu and English code-switching sentences collected from social media comments and online content. The purpose of this dataset is to represent the way Pakistani users naturally mix Roman Urdu and English while communicating online. Data Collection The sentences were collected from: Instagram comments YouTube comments Twitter/X The collected sentences… See the full description on the dataset page: https://huggingface.co/datasets/samaikaimran/code-switching-codesaviours-si26-samaika.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes16downloads
Dataset Card

Code Switching NLP | Code Saviours SI-26 | Samaika

About the Dataset

This dataset contains Roman Urdu and English code-switching sentences collected from social media comments and online content.

The purpose of this dataset is to represent the way Pakistani users naturally mix Roman Urdu and English while communicating online.

Data Collection

The sentences were collected from:

  • Instagram comments
  • YouTube comments
  • Twitter/X

The collected sentences were selected to represent natural Roman Urdu-English code-switching patterns used in online communication.

Labels

Each word in the dataset is labelled as one of the following:

  • URD — Roman Urdu word
  • ENG — English word
  • MIX — A word containing a mixture of Roman Urdu and English

Dataset Format

The dataset contains three columns:

  • sentence — Complete sentence
  • word — Individual word from the sentence
  • label — Language label for that word

Dataset Statistics

  • Total sentences: 150
  • Total word entries: 1,435
  • URD words: 842
  • ENG words: 593
  • MIX words: 0

Project

This dataset was created as part of Code Saviours SI-26 — Project 2: Code Switching Dataset Collection.