tanglish
Datasets
All datasets matching “tanglish”tanglish-tamil
Translation of Tanglish to tamil
Source: karky.in
To use
python import datasets s = datasets.load_dataset('Deepakvictor/tanglish-tamil') print(s) """DatasetDict({ train: Dataset({ features: ['Movie', 'FileName', 'Song', 'Tamillyrics', 'Tanglishlyrics', 'Mood', 'Genre'], num_rows: 597 }) })"""
Credits and Source: https://karky.in/
For simpler version
Visit this dataset --> "Deepakvictor/tan-tam"
tanglishmedbench
TanglishMedBench
Dataset Summary
TanglishMedBench is a gold-standard evaluation benchmark of 700 medical question-answer pairs in Tanglish — romanized, WhatsApp-style Tamil-English code-mixed text, as commonly typed by Tamil speakers in everyday messaging. The dataset is designed to evaluate how well language models handle medical queries posed in this informal, code-mixed register, which differs substantially from the formal Tamil or English text most models are… See the full description on the dataset page: https://huggingface.co/datasets/anthea1407/tanglishmedbench.Tanglish-Corpus-185k
Tanglish-Corpus-185k
The largest Romanised Tamil-English code-mixed sentence corpus ever assembled — 185,973 clean sentences.
Overview
This corpus contains 185,973 clean Tanglish sentences — Tamil and English mixed in Roman script, scraped from real Indian social media. It is 11.8x larger than DravidianCodeMix (15,744 sentences), the previous largest publicly available Tanglish corpus.
Built from scratch using a custom morphological filtering pipeline — no… See the full description on the dataset page: https://huggingface.co/datasets/vishnu-n/Tanglish-Corpus-185k.tanglish_conversation_dataset_listenglish-to-tanglish-translation-datasetTanglishSTS
TanglishSTS
The first human-annotated Semantic Textual Similarity benchmark for Romanised Tamil-English code-mixed text.
Overview
TanglishSTS is a 325-pair evaluation benchmark for measuring how well sentence embedding models understand Tanglish — Tamil and English mixed in Roman script, the way 80+ million Tamil speakers actually communicate on WhatsApp, YouTube, Instagram, and Reddit.
No such benchmark existed before this work. TanglishSTS fills the gap.… See the full description on the dataset page: https://huggingface.co/datasets/vishnu-n/TanglishSTS.
