CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes15k downloads2y agoHugging Face02youssef101 /artelingo-dummyArtELingo is a benchmark and dataset introduced in a research paper aimed at promoting work on diversity across languages and cultures. It is an extension of ArtEmis, which is a collection of 80,000 artworks from WikiArt with 450,000 emotion labels and English-only captions. ArtELingo expands this dataset by adding 790,000 annotations in Arabic and Chinese. The purpose of these additional annotations is to evaluate the performance of "cultural-transfer" in AI systems. The dataset in ArtELingo… See the full description on the dataset page: https://huggingface.co/datasets/youssef101/artelingo-dummy.imageimage-to-text10K<n<100K2 likes1k downloads3y agoHugging Face03ppak10 /AdditiveLLM2-OA AdditiveLLM2-OA Dataset Open Access journal articles (up to February 2026) used in domain adapting pretraining and instruction tuning for AdditiveLLM2. Dataset Split by Journal text images vit Vocabulary Overlap Pairwise Jaccard similarity of word-level vocabularies (lowercase, 3+ letter tokens) across the four source journals. Run info/vocabulary/vocabulary_overlap.py to reproduce. Top Phrases by Journal Most frequent bigrams and… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/AdditiveLLM2-OA.imagetext-generation10K<n<100K2 likes561 downloads6mo agoHugging Face04artefactory /Argimi-Ardian-Finance-10k-text-image The ArGiMI Ardian datasets : text and images The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.imagetext-retrieval1M<n<10M14 likes530 downloads7mo agoHugging Face05AlroWilde /olmOCR-mix-1025-Photoreal ⚠️ Important Feedback Invitation If this dataset has brought stronger positive or negative improvements to your model training, I warmly welcome you to email me your feedback at any time. This is very important to me — thank you! 📧Contact: hi@support.alrowilde.com Photorealistic enhancements of document pages from allenai/olmOCR-mix-1025.Clean PDF renders are transformed into realistic scanned/photographed images with natural lighting, paper texture, shadows and capture… See the full description on the dataset page: https://huggingface.co/datasets/AlroWilde/olmOCR-mix-1025-Photoreal.imagetext-generationn<1K0 likes164 downloads2mo agoHugging Face06mishl /StreetVision-10K StreetVision-10K Each sample contains: A system prompt instructing the model to act as an OSINT/geospatial expert A user message with a street-level photo and the instruction to determine coordinates An assistant response with ground-truth coordinates in <direct_lon_lat_output>longitude,latitude</direct_lon_lat_output> format Format Each line is a JSON array of ChatML messages: [ {"role": "system", "content": "..."}, {"role": "user", "content": [… See the full description on the dataset page: https://huggingface.co/datasets/mishl/StreetVision-10K.imagevisual-question-answering10K<n<100K1 likes161 downloads6mo agoHugging Face07Alfaxad /vector-100k VectorOS Vector 100k SimSat VLM Dataset VectorOS Vector 100k is a high-fidelity multimodal instruction dataset for fine-tuning vision-language models on geospatial epidemiology tasks. It was built for the VectorOS hackathon project and targets LiquidAI/LFM2.5-VL-450M. The dataset contains 100,000 chat-style examples derived from 10,000 geospatial chips across 30 AOIs. Every accepted chip has a real SimSat Sentinel-2 true-color view, a real SimSat Sentinel-2 NIR-red-green false-color… See the full description on the dataset page: https://huggingface.co/datasets/Alfaxad/vector-100k.imagevisual-question-answering100K<n<1M1 likes101 downloads5mo agoHugging Face08NewstaR /bleedingheart-pretrain-10MBleedingheart Pretrain Dataset A collaboration between Kaleido and Newstar We collected all the datasets we could find that are in Tagalog or any other Philippine dialect and put them in this repository. This data will be used to train the Bleedingheart model. Bleeding Heart is a stunning bird native to the island of Luzon in the Philippines. It is a medium-sized ground dove with a distinctive red patch of feathers on its chest, which gives it its name. The male's red patch is… See the full description on the dataset page: https://huggingface.co/datasets/NewstaR/bleedingheart-pretrain-10M.imagetext-generation1M<n<10M1 likes94 downloads3y agoHugging Face09Mekasiu1018 /3documenttoken-classificationn<1K0 likes85 downloads4mo agoHugging Face10PJMixers-Images /bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219. The format should be similar to that of liuhaotian/LLaVA-Instruct-150K. Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks. Generation Details If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.imagetext-generation1K<n<10K1 likes66 downloads2y agoHugging Face11Daniel10004 /Nemotron-Personas-Korea Nemotron-Personas-Korea 우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템 A compound AI approach to personas grounded in real-world distributions 데이터셋 개요 (Overview) Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다. Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/Daniel10004/Nemotron-Personas-Korea.imagetext-generation1M<n<10M0 likes61 downloads5mo agoHugging Face12ArkaMukherjee /reasoning-10k-v2 🧠 Dataset for Vision-Language Reasoning Reasoning-10K-v2 is a 10,000-sample multimodal reasoning dataset. Our work unifies diverse sources of vision–language reasoning data spanning code, mathematics, geography, tables, and scientific figures. Using a hybrid of synthesis and filtration strategies inspired by prior baselines (LIMO VL, MM MathInstruct, and Multimodal Open R1), we curated high-quality reasoning examples from open and synthetic data. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/ArkaMukherjee/reasoning-10k-v2.documenttext-generation10K<n<100K2 likes55 downloads11mo agoHugging Face13TurkishCodeMan /recipe-synthetic-images-10k Recipe PDF Dataset A multimodal dataset of 10K+ recipes rendered as PDF images with full metadata. Dataset Description Each sample contains: image: Recipe rendered as a styled PDF page (PNG, ~1654x2339px) name: Recipe title description: Recipe description ingredients: List of ingredients steps: Cooking instructions nutrition: Nutritional values (calories, fat%, sugar%, sodium%, protein%, sat.fat%, carbs%) random_reviews: User reviews minutes: Cooking time tags: Recipe… See the full description on the dataset page: https://huggingface.co/datasets/TurkishCodeMan/recipe-synthetic-images-10k.imageimage-to-text10K<n<100K0 likes54 downloads8mo agoHugging Face14alst10 /gemma4-multimodal-recipe-dataset 🍳 Gemma 4 Multimodal Recipe & Food Dataset A balanced, high-density multimodal dataset curated specifically for fine-tuning compact vision-language models (such as gemma-4-e2b-it) for visual food recognition, recipe generation, and dietary recommendation. 🔗 Upstream & Source Datasets This dataset was created by cleaning, reformatting, and synthesizing samples across the following 5 Hugging Face sources: Dataset Modality Role in Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/alst10/gemma4-multimodal-recipe-dataset.imagevisual-question-answering10K<n<100K0 likes52 downloads1mo agoHugging Face15Coder109 /C3BEnglish | 简体中文 C³B: Comics Cross-Cultural Benchmark Culture In a Frame: C³B as a Comic-Based Benchmark for Multimodal Cultural Awareness ICLR 2026 About C³B C³B (Comics Cross-Cultural Benchmark) is a multicultural, multitask, and multilingual benchmark for evaluating cultural awareness capabilities of Multimodal Large Language Models (MLLMs). Progressive task difficulty: From basic visual recognition, to higher-level cultural conflict understanding, to cultural content… See the full description on the dataset page: https://huggingface.co/datasets/Coder109/C3B.imagevisual-question-answering1K<n<10K1 likes39 downloads7mo agoHugging Face16DemoTest0122 /Demo101 Overview 🚀📚 The first comprehensive dataset for training AI models to write complete novels with sophisticated reasoning. 🧠 Hierarchical Reasoning Architecture — Multi-layered planning traces including character archetypes, story arcs, world rules, and scene breakdowns. A complete cognitive roadmap for long-form narrative construction. 📖 Complete Novel Coverage — From 40,000 to 600,000+ tokens per book, spanning novellas to epic series with consistent quality throughout. ⚡… See the full description on the dataset page: https://huggingface.co/datasets/DemoTest0122/Demo101.imagetext-generationn<1K0 likes34 downloads8mo agoHugging Face17DemoTest0122 /Demo103 Overview 🚀📚 The first comprehensive dataset for training AI models to write complete novels with sophisticated reasoning. 🧠 Hierarchical Reasoning Architecture — Multi-layered planning traces including character archetypes, story arcs, world rules, and scene breakdowns. A complete cognitive roadmap for long-form narrative construction. 📖 Complete Novel Coverage — From 40,000 to 600,000+ tokens per book, spanning novellas to epic series with consistent quality throughout. ⚡… See the full description on the dataset page: https://huggingface.co/datasets/DemoTest0122/Demo103.imagetext-generationn<1K0 likes30 downloads8mo agoHugging Face18DemoTest0122 /Demo102 Overview 🚀📚 The first comprehensive dataset for training AI models to write complete novels with sophisticated reasoning. 🧠 Hierarchical Reasoning Architecture — Multi-layered planning traces including character archetypes, story arcs, world rules, and scene breakdowns. A complete cognitive roadmap for long-form narrative construction. 📖 Complete Novel Coverage — From 40,000 to 600,000+ tokens per book, spanning novellas to epic series with consistent quality throughout. ⚡… See the full description on the dataset page: https://huggingface.co/datasets/DemoTest0122/Demo102.imagetext-generationn<1K0 likes29 downloads8mo agoHugging Face19eosync /synthetic-users-1000 Synthetic Users 1000 🧪 Synthetic Users Dataset (1,000 Records) This dataset contains 1,000 high-quality synthetic user profiles, including realistic bios, usernames, metadata, and profile image filenames — ideal for: UI/UX testing Prototyping dashboards Mock data in SaaS apps User flow demos AI evaluation without PII 🧠 Features ✅ 100% privacy-safe 🧍 Full names, emails, bios, countries 📸 Profile image filenames (S3-ready) 🔍 Gender, age, emotion tags via… See the full description on the dataset page: https://huggingface.co/datasets/eosync/synthetic-users-1000.imagetext-generationn<1K1 likes13 downloads1y agoHugging Face20DBbun /100K_Cerebral_Palsy_v1.0gated Synthetic ROCS Dataset (100K Patients) Overview This dataset is a fully synthetic reproduction of the patient cohort described in A Big Data Approach to Evaluate Receipt of Optimal Care in Childhood Cerebral Palsy. It simulates the Receipt of Optimal Care Score (ROCS) framework across 100K synthetic patients, modeling event sequences, adherence probabilities, and component-level quality-of-care scores. All data were algorithmically generated — no real patient data are… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/100K_Cerebral_Palsy_v1.0.imagefeature-extractionn<1K0 likes2 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.