CoolFace
Datasetpublic

122Uswa/code-switching-codesaviours-si26-Uswa

Code-Switching Dataset: Roman Urdu ↔ English (Pakistan) Dataset Description This dataset contains 191 naturally code-switched Roman Urdu–English sentences** (well above the 150 minimum) blending Roman Urdu and English, the way Pakistani speakers actually write online. Every word in every sentence is labelled at the token level, making this a word-level sequence labelling / language identification dataset for code-switched text. Roman Urdu–English mixing is… See the full description on the dataset page: https://huggingface.co/datasets/122Uswa/code-switching-codesaviours-si26-Uswa.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes14downloads

122Uswa/code-switching-codesaviours-si26-Uswa · main · files are served by the source, never re-hosted here