CoolFace
Datasetpublic

flax-community/conceptual-12m-multilingual-marian

This dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models. Data distribution is following: train_file_marian_final.tsv: 10010625 captions (2502656 captions of English, German, Spanish, French each) val_file_marian_final.tsv:… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian.

sourceHugging Faceupdated 3y agoView on Hugging Face
1likes208downloads
5 commits on main
e2ddc4e3y ago

Update README.md (#1)

pcuenq, lbourdois
699a1165y ago

update README

bhavitvyamalik
b3cd5c75y ago

update README

bhavitvyamalik
77b6b915y ago

add data for training MIC

bhavitvyamalik
746505c5y ago

initial commit

system