datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
5_uit_paragraphiam_paragraphparagraph_translation_eng_kh
Machine Translation Dataset: Khmer to English
A machine translation dataset project that pairs Khmer document images with their corresponding English translations. This dataset is designed for training machine translation models to convert Khmer text (via OCR from images) to English text.
Project Overview
This project processes bilingual documents from BOP (Bank of PNG) Bulletins, creating a Khmer-English machine translation dataset. Each sample pairs a Khmer image… See the full description on the dataset page: https://huggingface.co/datasets/vichetkao/paragraph_translation_eng_kh.funsd-bank-paragraph-test32_vnondb_paragraph5_uit_paragraph_transform5_uit_paragraph_overlaycc3m-caption-or-paragraph
