samaikaimran/code-switching-codesaviours-si26-samaika
Code Switching NLP | Code Saviours SI-26 | Samaika About the Dataset This dataset contains Roman Urdu and English code-switching sentences collected from social media comments and online content. The purpose of this dataset is to represent the way Pakistani users naturally mix Roman Urdu and English while communicating online. Data Collection The sentences were collected from: Instagram comments YouTube comments Twitter/X The collected sentences… See the full description on the dataset page: https://huggingface.co/datasets/samaikaimran/code-switching-codesaviours-si26-samaika.
Code Switching NLP | Code Saviours SI-26 | Samaika
About the Dataset
This dataset contains Roman Urdu and English code-switching sentences collected from social media comments and online content.
The purpose of this dataset is to represent the way Pakistani users naturally mix Roman Urdu and English while communicating online.
Data Collection
The sentences were collected from:
- Instagram comments
- YouTube comments
- Twitter/X
The collected sentences were selected to represent natural Roman Urdu-English code-switching patterns used in online communication.
Labels
Each word in the dataset is labelled as one of the following:
- URD — Roman Urdu word
- ENG — English word
- MIX — A word containing a mixture of Roman Urdu and English
Dataset Format
The dataset contains three columns:
sentence— Complete sentenceword— Individual word from the sentencelabel— Language label for that word
Dataset Statistics
- Total sentences: 150
- Total word entries: 1,435
- URD words: 842
- ENG words: 593
- MIX words: 0
Project
This dataset was created as part of Code Saviours SI-26 — Project 2: Code Switching Dataset Collection.
