En-Hi
Datasets
All datasets matching “En-Hi”quickmt-train.hi-en
quickmt hi-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
IITB-hien_dev-1.5-hin-eng
Neulab-tedtalks_test-1-eng-hin
Google-wmt24pp-1-eng-hin_IN
IITB-hien_test-1.5-hin-eng
Statmt-news_commentary-14-eng-hin
Statmt-news_commentary-15-eng-hin
Statmt-news_commentary-16-eng-hin
Statmt-news_commentary-17-eng-hin
Statmt-news_commentary-18-eng-hin
Statmt-news_commentary-18.1-eng-hin
Statmt-pmindia-1-eng-hin… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.hi-en.HINMIX_hi-en
Dataset Card for Hindi English Codemix Dataset - HINMIX
HINMIX is a massive parallel codemixed dataset for Hindi-English code switching.
See the 📚 paper on arxiv to dive deep into this synthetic codemix data generation pipeline.
Dataset contains 4.2M fully parallel sentences in 6 Hindi-English forms.
Further, we release gold standard codemix dev and test set manually translated by proficient bilingual annotators.
Dev Set consists of 280 examples
Test set consists of 2507 examples… See the full description on the dataset page: https://huggingface.co/datasets/kartikagg98/HINMIX_hi-en.magpie-en-eu-reasoning-instructions-qwen3
Dataset Card for magpie-en-eu-reasoning-instructions-qwen3
Dataset Summary
The magpie-en-eu-reasoning-instructions-qwen3 dataset is a large-scale, high-quality, bilingual instruction and preference dataset developed by the HiTZ Center. It is specifically tailored for training, aligning, and evaluating reasoning-focused Large Language Models (LLMs) in both English and Basque (Euskera).
Built using the self-synthesizing Magpie methodology, the dataset contains a… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/magpie-en-eu-reasoning-instructions-qwen3.wiki_en-docwiki_en_rawindic_hi_en_tts
