datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tanglishmedbench
TanglishMedBench
Dataset Summary
TanglishMedBench is a gold-standard evaluation benchmark of 700 medical question-answer pairs in Tanglish — romanized, WhatsApp-style Tamil-English code-mixed text, as commonly typed by Tamil speakers in everyday messaging. The dataset is designed to evaluate how well language models handle medical queries posed in this informal, code-mixed register, which differs substantially from the formal Tamil or English text most models are… See the full description on the dataset page: https://huggingface.co/datasets/anthea1407/tanglishmedbench.TanglishSTS
TanglishSTS
The first human-annotated Semantic Textual Similarity benchmark for Romanised Tamil-English code-mixed text.
Overview
TanglishSTS is a 325-pair evaluation benchmark for measuring how well sentence embedding models understand Tanglish — Tamil and English mixed in Roman script, the way 80+ million Tamil speakers actually communicate on WhatsApp, YouTube, Instagram, and Reddit.
No such benchmark existed before this work. TanglishSTS fills the gap.… See the full description on the dataset page: https://huggingface.co/datasets/vishnu-n/TanglishSTS.
