CoolFace
Datasetpublic

hamnaheh/code-switching-codesaviours-si26-humna

Code Switching NLP Dataset — Code Saviours SI-26, Humna Imran Word-level labeled dataset of naturally code-switched Roman Urdu + English sentences, as commonly written by Pakistani social media users. What it is 220 sentences (2,490 word-level rows). Each word is labeled URD (Roman Urdu), ENG (English), or MIX (hybrid token, e.g. hyphenated compounds). How it was built Source sentences come from Smat26/Roman-Urdu-Dataset (GitHub), a public… See the full description on the dataset page: https://huggingface.co/datasets/hamnaheh/code-switching-codesaviours-si26-humna.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes6downloads

hamnaheh/code-switching-codesaviours-si26-humna · main · files are served by the source, never re-hosted here