CoolFace
Datasetpublic

flax-community/conceptual-12m-multilingual-marian-128

This dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models (with sequence length 128). Data distribution is following: train_file_marian_final.tsv: 10002432 captions (2500608 captions of English, German, Spanish, French each)… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian-128.

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes184downloads

flax-community/conceptual-12m-multilingual-marian-128 · main · files are served by the source, never re-hosted here