kartikagg98/HINMIX_hi-en
Dataset Card for Hindi English Codemix Dataset - HINMIX HINMIX is a massive parallel codemixed dataset for Hindi-English code switching. See the ๐ paper on arxiv to dive deep into this synthetic codemix data generation pipeline. Dataset contains 4.2M fully parallel sentences in 6 Hindi-English forms. Further, we release gold standard codemix dev and test set manually translated by proficient bilingual annotators. Dev Set consists of 280 examples Test set consists of 2507โฆ See the full description on the dataset page: https://huggingface.co/datasets/kartikagg98/HINMIX_hi-en.
This repository belongs to kartikagg98 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
