CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thaottn /datacomp-medium-pool-translatedimage100M<n<1B0 likes1k downloads1y agoHugging Face02vidore /li-vdr-translatedimage100K<n<1M0 likes514 downloads1y agoHugging Face03lewington /laion2B-multi-joined-translated-to-en-smolimage10M<n<100M2 likes390 downloads2y agoHugging Face04vidore /colpali-train-set-splitted-translatedimage100K<n<1M0 likes202 downloads1y agoHugging Face05umairhassan02 /urdu-translated-coco-captions-subset Research Paper: https://www.arxiv.org/abs/2509.09014 Github: https://github.com/umair-hassan2/COCO-Urdu Overview Urdu, spoken by over 250 million people, remains critically under-served in multimodal and vision-language research. COCO-Urdu addresses this gap by providing 59K images and 319K high-quality Urdu captions. Captions were generated via zero-shot translation using SeamlessM4T v2, validated with a hybrid QE pipeline combining COMET-Kiwi, CLIP-based visual grounding, and BERTScore… See the full description on the dataset page: https://huggingface.co/datasets/umairhassan02/urdu-translated-coco-captions-subset.image10K<n<100K0 likes153 downloads1y agoHugging Face06Berkesule /translated_visual_puzzles_with_questionimagen<1K0 likes81 downloads9mo agoHugging Face075CD-AI /Vietnamese-ShareGPT4Video-ShareGPT4Video-gg-translatedtextvisual-question-answering10K<n<100K0 likes77 downloads2y agoHugging Face085CD-AI /Vietnamese-OpenGVLab-ShareGPT-4o-gg-translatedtextvisual-question-answering10K<n<100K0 likes49 downloads2y agoHugging Face09oddadmix /the_cauldron-vqav2-translated_datasetimage10K<n<100K0 likes41 downloads5mo agoHugging Face10ayganyavuz /translated_visual_puzzlesimage1K<n<10K0 likes40 downloads11mo agoHugging Face11kimmchii /translated-th-coco2017 Dataset Card for "translated-th-coco2017" More Information needed image100K<n<1M0 likes34 downloads2y agoHugging Face12Danielmse /Flux_translateimagen<1K1 likes34 downloads1y agoHugging Face13laicsiifes /flickr30k-pt-br-human-translatedimage10K<n<100K1 likes32 downloads2y agoHugging Face14ayganyavuz /translated_mmiq_datasetimage1K<n<10K1 likes32 downloads11mo agoHugging Face15Berkesule /PuzzleVQA-MultipleChoice-Translatedimage1K<n<10K0 likes30 downloads10mo agoHugging Face16oddadmix /the_cauldron-vqav2-translated_dataset-sm Dataset Card for "the_cauldron-vqav2-translated_dataset-sm" More Information needed image10K<n<100K0 likes27 downloads5mo agoHugging Face17mesolitica /translate-Multi-modal-Self-instruct Translated https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct Translate to Malay using https://mesolitica.com/translation Base model, a nice dataset for visual QA charts, tables, simulated maps, dashboards, flowcharts, relation graphs, floor plans, and visual puzzles. image10K<n<100K2 likes26 downloads2y agoHugging Face18Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationimage1K<n<10K0 likes23 downloads3y agoHugging Face19umairhassan02 /coco-captions-subset-59k-translated-final License Information The annotations in this dataset along with this website belong to the COCO Consortium and are licensed under a Creative Commons Attribution 4.0 License. Images The COCO Consortium does not own the copyright of the images. Use of the images must abide by the Flickr Terms of Use. The users of the images accept full responsibility for the use of the dataset, including but not limited to the use of any copies of copyrighted images that they may create from the… See the full description on the dataset page: https://huggingface.co/datasets/umairhassan02/coco-captions-subset-59k-translated-final.image100K<n<1M0 likes21 downloads1y agoHugging Face20Berkesule /translated_mmiq_dataset_with_questionimage1K<n<10K0 likes20 downloads10mo agoHugging Face21mesolitica /translate-DenseFusion-1M Translated https://huggingface.co/datasets/BAAI/DenseFusion-1M Translate to Malay using https://mesolitica.com/translation Base model, a nice dataset for OCR with description. We make sure translated text also maintain the same OCR image0 likes18 downloads2y agoHugging Face22vidore /vdsid-translatedimage10K<n<100K0 likes18 downloads1y agoHugging Face23jaeyong2 /LAION-art-EN-improved-captions-translate Development Process source dataset from recastai/LAION-art-EN-improved-captions We used Qwen/Qwen2-72B-Instruct model to translate. License Qwen/Qwen2.5-72B-Instruct : https://huggingface.co/Qwen/Qwen2-72B-Instruct/blob/main/LICENSE recastai/LAION-art-EN-improved-captions : https://huggingface.co/datasets/choosealicense/licenses/blob/main/markdown/cc-by-4.0.md Acknowledgement This research is supported by TPU Research Cloud program. image100K<n<1M0 likes17 downloads2y agoHugging Face24gdurkin /sensor_translate_s1s2image10K<n<100K0 likes14 downloads2y agoHugging Face25laicsiifes /flickr30k-pt-br-5k-human-translatedimage1K<n<10K0 likes14 downloads2y agoHugging Face26gdurkin /translate_laimagen<1K0 likes14 downloads1y agoHugging Face27karya /endangered-recipes-translated-500gated Endangered Recipes Translated 500 Endangered Recipes Translated 500 is the next phase of the ELR-1000 research effort. It brings together 500 community-contributed recipes from endangered and under-represented Indic languages, with 50 recipes per language. Alongside the original recipe text, this release includes English translations for recipe names, ingredients, tools, cultural notes, and recipe steps. The source collection and research context are described in the ELR-1000 paper… See the full description on the dataset page: https://huggingface.co/datasets/karya/endangered-recipes-translated-500.imagetranslation0 likes14 downloads5mo agoHugging Face28Arabic-Clip /ccs_synthetic_translated_arabic_processedimage10M<n<100M1 likes13 downloads2y agoHugging Face29Arabic-Clip-Archive /Arabic_dataset_1M_translated_jsonl_format_ViT-B-16-plus-240This translation done using https://huggingface.co/Helsinki-NLP/opus-mt-en-ar image100K<n<1M0 likes12 downloads3y agoHugging Face30Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240This dataset repo contains the dataset (CC3M+CC12M+SBU) translated using opus-mt-en-ar and cleaned. Its size about 13M image1M<n<10M0 likes12 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.