datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datacomp-medium-pool-translatedli-vdr-translatedlaion2B-multi-joined-translated-to-en-smolcolpali-train-set-splitted-translatedurdu-translated-coco-captions-subset
Research Paper: https://www.arxiv.org/abs/2509.09014
Github: https://github.com/umair-hassan2/COCO-Urdu
Overview
Urdu, spoken by over 250 million people, remains critically under-served in multimodal and vision-language research. COCO-Urdu addresses this gap by providing 59K images and 319K high-quality Urdu captions. Captions were generated via zero-shot translation using SeamlessM4T v2, validated with a hybrid QE pipeline combining COMET-Kiwi, CLIP-based visual grounding, and BERTScore… See the full description on the dataset page: https://huggingface.co/datasets/umairhassan02/urdu-translated-coco-captions-subset.Vietnamese-ShareGPT4Video-ShareGPT4Video-gg-translatedtranslated_visual_puzzles_with_questionVietnamese-OpenGVLab-ShareGPT-4o-gg-translatedtranslated_visual_puzzlesthe_cauldron-vqav2-translated_datasettranslated-th-coco2017
Dataset Card for "translated-th-coco2017"
More Information needed
flickr30k-pt-br-human-translatedtranslated_mmiq_datasetthe_cauldron-vqav2-translated_dataset-sm
Dataset Card for "the_cauldron-vqav2-translated_dataset-sm"
More Information needed
PuzzleVQA-MultipleChoice-TranslatedArabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationcoco-captions-subset-59k-translated-final
License Information
The annotations in this dataset along with this website belong to the COCO Consortium and are licensed under a Creative Commons Attribution 4.0 License.
Images
The COCO Consortium does not own the copyright of the images. Use of the images must abide by the Flickr Terms of Use. The users of the images accept full responsibility for the use of the dataset, including but not limited to the use of any copies of copyrighted images that they may create from the… See the full description on the dataset page: https://huggingface.co/datasets/umairhassan02/coco-captions-subset-59k-translated-final.vdsid-translatedtranslated_mmiq_dataset_with_questiontranslate-DenseFusion-1M
Translated https://huggingface.co/datasets/BAAI/DenseFusion-1M
Translate to Malay using https://mesolitica.com/translation Base model, a nice dataset for OCR with description. We make sure translated text also maintain the same OCR
flickr30k-pt-br-5k-human-translatedArabic_dataset_1M_translated_jsonl_format_ViT-B-16-plus-240This translation done using https://huggingface.co/Helsinki-NLP/opus-mt-en-ar
ccs_synthetic_translated_arabicThe columns inside the dataset as follows:
index
url
caption_en
caption_ar
The dataset size is 12556500 rows × 4 columns
ccs_synthetic_translated_arabic_processedArabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240This dataset repo contains the dataset (CC3M+CC12M+SBU) translated using opus-mt-en-ar and cleaned. Its size about 13M
Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240-fulldata-v2DatasetDict({
train: Dataset({
features: ['index', 'embeddings', 'en_caption', 'ar_caption', 'nr_words', 'url'],
num_rows: 12166802
})
})
clevr1000_rephrased_langsplit_v1_translatedendangered-recipes-translated-500
Endangered Recipes Translated 500
Endangered Recipes Translated 500 is the next phase of the
ELR-1000 research effort. It
brings together 500 community-contributed recipes from endangered and
under-represented Indic languages, with 50 recipes per language. Alongside the
original recipe text, this release includes English translations for recipe
names, ingredients, tools, cultural notes, and recipe steps.
The source collection and research context are described in the ELR-1000 paper… See the full description on the dataset page: https://huggingface.co/datasets/karya/endangered-recipes-translated-500.viz-wiz-train-translatedccs_synthetic_ar_1M-Arabic_dataset_1M_translated_jsonl_format
