CoolFace
Datasetpublic

flax-community/conceptual-12m-multilingual-marian

This dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models. Data distribution is following: train_file_marian_final.tsv: 10010625 captions (2502656 captions of English, German, Spanish, French each) val_file_marian_final.tsv:… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian.

sourceHugging Faceupdated 3y agoView on Hugging Face
1likes212downloads
Dataset Card

This dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models. Data distribution is following:

train_file_marian_final.tsv: 10010625 captions (2502656 captions of English, German, Spanish, French each) <br /> val_file_marian_final.tsv: 110592 captions (27648 captions of English, German, Spanish, French each)