CoolFace
Datasetpublic

sarvamai/tatoeba-indic

Tatoeba Benchmark (Indian languages only) This benchmark is prepared from the 2023 Tatoeba Challenge, by extracting the dev and test sets for languages spoken in the Indian Republic. The code to download and process the data can be found here in the repo: data_prep/original_v1/extract.py Note: This is not the official version of Tatoeba benchmark. Just a processed mirror for Indian languages, made available in HuggingFace for ease of use. Languages Language… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/tatoeba-indic.

sourceHugging Faceupdated 1y agoView on Hugging Face
2likes250downloads
Dataset Card

Tatoeba Benchmark (Indian languages only)

This benchmark is prepared from the 2023 Tatoeba Challenge, by extracting the dev and test sets for languages spoken in the Indian Republic.

The code to download and process the data can be found here in the repo: data_prep/original_v1/extract.py

Note: This is not the official version of Tatoeba benchmark. Just a processed mirror for Indian languages, made available in HuggingFace for ease of use.

Languages

Language codeLanguage name
asmAssamese
awaAwadhi
benBengali
bhoBhojpuri
brxBodo
gujGujarati
hinHindi
kanKannada
khaKhasi
kokKonkani
lahLahnda
maiMaithili
malMalayalam
marMarathi
mniManipuri
nepNepali
oriOdia
panPanjabi
pliPali
sanSanskrit
satSantali
sndSindhi
tamTamil
telTelugu
urdUrdu