hamnaheh/code-switching-codesaviours-si26-humna
Code Switching NLP Dataset — Code Saviours SI-26, Humna Imran Word-level labeled dataset of naturally code-switched Roman Urdu + English sentences, as commonly written by Pakistani social media users. What it is 220 sentences (2,490 word-level rows). Each word is labeled URD (Roman Urdu), ENG (English), or MIX (hybrid token, e.g. hyphenated compounds). How it was built Source sentences come from Smat26/Roman-Urdu-Dataset (GitHub), a public… See the full description on the dataset page: https://huggingface.co/datasets/hamnaheh/code-switching-codesaviours-si26-humna.
06
