qandeelasim13/code-switching-codesaviours-si26-qandeel
Code Switching NLP Dataset | Code Saviours SI-26 Dataset Description A word-level labelled Roman Urdu–English code-switching dataset (200 sentences, 3182 word entries). Labels: URD (Roman Urdu), ENG (English), MIX (nativized loanword). Source Sentences filtered from the Roman Urdu Data Set (Sharf, 2017, UCI ML Repository, CC BY 4.0), originally collected from e-commerce reviews, Facebook comments, and Twitter posts. Filtered for genuine… See the full description on the dataset page: https://huggingface.co/datasets/qandeelasim13/code-switching-codesaviours-si26-qandeel.
license: cc-by-4.0 language:
- ur
- en task_categories:
- token-classification ---
Code Switching NLP Dataset | Code Saviours SI-26
Dataset Description
A word-level labelled Roman Urdu–English code-switching dataset (200 sentences, 3182 word entries). Labels: URD (Roman Urdu), ENG (English), MIX (nativized loanword).
Source
Sentences filtered from the Roman Urdu Data Set (Sharf, 2017, UCI ML Repository, CC BY 4.0), originally collected from e-commerce reviews, Facebook comments, and Twitter posts. Filtered for genuine code-switching and word-labelled.
Attribution
Sharf, Z. (2017). Roman Urdu Data Set [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C58325
