halimajaved592/code-switching-codesaviours-si26-halima
Code Switching Dataset Description This dataset was created as part of the Code Saviours SI-26 Project 2. It contains 150 Roman Urdu–English code-switching sentences. Each sentence is split into individual words, and every word is labeled according to the language it belongs to. The dataset is intended for educational and research purposes in Natural Language Processing (NLP), especially for language identification and code-switching detection. Data Collection The sentences were collected from… See the full description on the dataset page: https://huggingface.co/datasets/halimajaved592/code-switching-codesaviours-si26-halima.
Code Switching Dataset Description
This dataset was created as part of the Code Saviours SI-26 Project 2. It contains 150 Roman Urdu–English code-switching sentences. Each sentence is split into individual words, and every word is labeled according to the language it belongs to. The dataset is intended for educational and research purposes in Natural Language Processing (NLP), especially for language identification and code-switching detection.
Data Collection
The sentences were collected from publicly available online sources where Roman Urdu and English are commonly mixed in everyday communication. The dataset was then cleaned, organized, tokenized into words, and labeled manually.
Labels URD – Roman Urdu word ENG – English word MIX – Mixed-language token (used when a single token contains both Roman Urdu and English) Dataset Information Total Sentences: 150 Language: Roman Urdu + English Format: CSV Columns: sentence – Original mixed-language sentence word – Individual word/token label – Language label (URD, ENG, or MIX) Purpose
This dataset can be used for code-switching detection, token-level language identification, and other Natural Language Processing (NLP) tasks involving Roman Urdu and English text.
