datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nagamese-english-mt
Nagamese-English Machine Translation Corpus
Dataset Summary
This dataset contains 3,340 parallel sentence pairs in English and
Nagamese (Naga Pidgin), intended for training and evaluating machine
translation systems between the two languages. It has already been used to
fine-tune at least one NLLB-200-based translation model
(agnivamaiti/nllb-200-en-nagamese).
This is the first dataset card written for this dataset — no card
existed on the repository prior to this… See the full description on the dataset page: https://huggingface.co/datasets/agnivamaiti/nagamese-english-mt.Vaani-nagamese-majority-lg-English-no-transcript0Vaani-nagamese-majority-lg-English-with-transcriptEnglish-Nagamesenagamese-synthetic-ttsnagamese2english
