CoolFace
Datasetpublic

Tanushreeeeee/COMI-LINGUA

Dataset Details COMI-LINGUA (COde-MIxing and LINGuistic Insights on Natural Hinglish Usage and Annotation) is a high-quality Hindi-English code-mixed dataset, manually annotated by three annotators. It serves as a benchmark for multilingual NLP models by covering multiple foundational tasks. COMI-LINGUA provides annotations for several key NLP tasks: Language Identification (LID): Token-wise classification of Hindi, English, and other linguistic units. Initial predictions were… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/COMI-LINGUA.

sourceHugging Facecc-by-4.0updated 10mo agoView on Hugging Face
0likes77downloads
settings

This repository belongs to Tanushreeeeee on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameCOMI-LINGUA
visibilitypublic
licencecc-by-4.0
gatedno
ownerTanushreeeeee
Account settings
Tanushreeeeee/COMI-LINGUA · CoolFace