datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datacomp-medium-pool-translatedli-vdr-translatedlaion2B-multi-joined-translated-to-en-smolcolpali-train-set-splitted-translatedurdu-translated-coco-captions-subset
Research Paper: https://www.arxiv.org/abs/2509.09014
Github: https://github.com/umair-hassan2/COCO-Urdu
Overview
Urdu, spoken by over 250 million people, remains critically under-served in multimodal and vision-language research. COCO-Urdu addresses this gap by providing 59K images and 319K high-quality Urdu captions. Captions were generated via zero-shot translation using SeamlessM4T v2, validated with a hybrid QE pipeline combining COMET-Kiwi, CLIP-based visual grounding, and BERTScore… See the full description on the dataset page: https://huggingface.co/datasets/umairhassan02/urdu-translated-coco-captions-subset.translated_visual_puzzles_with_questionVietnamese-ShareGPT4Video-ShareGPT4Video-gg-translatedVietnamese-OpenGVLab-ShareGPT-4o-gg-translatedthe_cauldron-vqav2-translated_datasettranslated_visual_puzzlestranslated-th-coco2017
Dataset Card for "translated-th-coco2017"
More Information needed
Flux_translateflickr30k-pt-br-human-translatedtranslated_mmiq_datasetPuzzleVQA-MultipleChoice-Translatedthe_cauldron-vqav2-translated_dataset-sm
Dataset Card for "the_cauldron-vqav2-translated_dataset-sm"
More Information needed
translate-Multi-modal-Self-instruct
Translated https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct
Translate to Malay using https://mesolitica.com/translation Base model, a nice dataset for visual QA charts, tables, simulated maps, dashboards, flowcharts, relation graphs, floor plans, and visual puzzles.
Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationcoco-captions-subset-59k-translated-final
License Information
The annotations in this dataset along with this website belong to the COCO Consortium and are licensed under a Creative Commons Attribution 4.0 License.
Images
The COCO Consortium does not own the copyright of the images. Use of the images must abide by the Flickr Terms of Use. The users of the images accept full responsibility for the use of the dataset, including but not limited to the use of any copies of copyrighted images that they may create from the… See the full description on the dataset page: https://huggingface.co/datasets/umairhassan02/coco-captions-subset-59k-translated-final.translated_mmiq_dataset_with_questiontranslate-DenseFusion-1M
Translated https://huggingface.co/datasets/BAAI/DenseFusion-1M
Translate to Malay using https://mesolitica.com/translation Base model, a nice dataset for OCR with description. We make sure translated text also maintain the same OCR
vdsid-translatedLAION-art-EN-improved-captions-translate
Development Process
source dataset from recastai/LAION-art-EN-improved-captions
We used Qwen/Qwen2-72B-Instruct model to translate.
License
Qwen/Qwen2.5-72B-Instruct : https://huggingface.co/Qwen/Qwen2-72B-Instruct/blob/main/LICENSE
recastai/LAION-art-EN-improved-captions : https://huggingface.co/datasets/choosealicense/licenses/blob/main/markdown/cc-by-4.0.md
Acknowledgement
This research is supported by TPU Research Cloud program.
sensor_translate_s1s2flickr30k-pt-br-5k-human-translatedtranslate_laendangered-recipes-translated-500
Endangered Recipes Translated 500
Endangered Recipes Translated 500 is the next phase of the
ELR-1000 research effort. It
brings together 500 community-contributed recipes from endangered and
under-represented Indic languages, with 50 recipes per language. Alongside the
original recipe text, this release includes English translations for recipe
names, ingredients, tools, cultural notes, and recipe steps.
The source collection and research context are described in the ELR-1000 paper… See the full description on the dataset page: https://huggingface.co/datasets/karya/endangered-recipes-translated-500.ccs_synthetic_translated_arabic_processedArabic_dataset_1M_translated_jsonl_format_ViT-B-16-plus-240This translation done using https://huggingface.co/Helsinki-NLP/opus-mt-en-ar
Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240This dataset repo contains the dataset (CC3M+CC12M+SBU) translated using opus-mt-en-ar and cleaned. Its size about 13M
