CoolFace
Datasetpublic

Tanushreeeeee/COMI-LINGUA

Dataset Details COMI-LINGUA (COde-MIxing and LINGuistic Insights on Natural Hinglish Usage and Annotation) is a high-quality Hindi-English code-mixed dataset, manually annotated by three annotators. It serves as a benchmark for multilingual NLP models by covering multiple foundational tasks. COMI-LINGUA provides annotations for several key NLP tasks: Language Identification (LID): Token-wise classification of Hindi, English, and other linguistic units. Initial predictions were… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/COMI-LINGUA.

sourceHugging Facecc-by-4.0updated 10mo agoView on Hugging Face
0likes77downloads

Tanushreeeeee/COMI-LINGUA · main · files are served by the source, never re-hosted here