liboaccn/OPUS-MIT-5M
Multilingual Image Translation Dataset: OPUS-MIT-5M The OPUS-MIT-5M image translation dataset is constructed by randomly sampling 5M sentence pairs from the OPUS corpus. Figure illustrates the distribution of image-text pairs across 20 language pairs within the OPUS-MIT-5M dataset. A key goal in creating the OPUS-MIT-5M dataset is to ensure a balanced representation across languages to enable robust multilingual image translation. We endeavor to synthesize an equal number… See the full description on the dataset page: https://huggingface.co/datasets/liboaccn/OPUS-MIT-5M.
17
No card is published for this repository, or it could not be fetched from Hugging Face right now.
