CoolFace
Datasetpublic

Tanushreeeeee/COMI-LINGUA

Dataset Details COMI-LINGUA (COde-MIxing and LINGuistic Insights on Natural Hinglish Usage and Annotation) is a high-quality Hindi-English code-mixed dataset, manually annotated by three annotators. It serves as a benchmark for multilingual NLP models by covering multiple foundational tasks. COMI-LINGUA provides annotations for several key NLP tasks: Language Identification (LID): Token-wise classification of Hindi, English, and other linguistic units. Initial predictions were… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/COMI-LINGUA.

sourceHugging Facecc-by-4.0updated 10mo agoView on Hugging Face
0likes78downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Tanushreeeeee/COMI-LINGUA · CoolFace