122Uswa/code-switching-codesaviours-si26-Uswa
Code-Switching Dataset: Roman Urdu ↔ English (Pakistan) Dataset Description This dataset contains 191 naturally code-switched Roman Urdu–English sentences** (well above the 150 minimum) blending Roman Urdu and English, the way Pakistani speakers actually write online. Every word in every sentence is labelled at the token level, making this a word-level sequence labelling / language identification dataset for code-switched text. Roman Urdu–English mixing is… See the full description on the dataset page: https://huggingface.co/datasets/122Uswa/code-switching-codesaviours-si26-Uswa.
This repository belongs to 122Uswa on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
