datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MINT-1T-PDF-CC-2024-10
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.artelingo-dummyArtELingo is a benchmark and dataset introduced in a research paper aimed at promoting work on diversity across languages and cultures. It is an extension of ArtEmis, which is a collection of 80,000 artworks from WikiArt with 450,000 emotion labels and English-only captions. ArtELingo expands this dataset by adding 790,000 annotations in Arabic and Chinese. The purpose of these additional annotations is to evaluate the performance of "cultural-transfer" in AI systems.
The dataset in ArtELingo… See the full description on the dataset page: https://huggingface.co/datasets/youssef101/artelingo-dummy.AdditiveLLM2-OA
AdditiveLLM2-OA Dataset
Open Access journal articles (up to February 2026) used in domain adapting
pretraining and instruction tuning for AdditiveLLM2.
Dataset Split by Journal
text
images
vit
Vocabulary Overlap
Pairwise Jaccard similarity of word-level vocabularies (lowercase, 3+ letter tokens) across the four source journals. Run info/vocabulary/vocabulary_overlap.py to reproduce.
Top Phrases by Journal
Most frequent bigrams and… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/AdditiveLLM2-OA.Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.olmOCR-mix-1025-Photoreal
⚠️ Important Feedback Invitation
If this dataset has brought stronger positive or negative improvements to your model training,
I warmly welcome you to email me your feedback at any time.
This is very important to me — thank you!
📧Contact: hi@support.alrowilde.com
Photorealistic enhancements of document pages from allenai/olmOCR-mix-1025.Clean PDF renders are transformed into realistic scanned/photographed images with natural lighting, paper texture, shadows and capture… See the full description on the dataset page: https://huggingface.co/datasets/AlroWilde/olmOCR-mix-1025-Photoreal.StreetVision-10K
StreetVision-10K
Each sample contains:
A system prompt instructing the model to act as an OSINT/geospatial expert
A user message with a street-level photo and the instruction to determine coordinates
An assistant response with ground-truth coordinates in <direct_lon_lat_output>longitude,latitude</direct_lon_lat_output> format
Format
Each line is a JSON array of ChatML messages:
[
{"role": "system", "content": "..."},
{"role": "user", "content": [… See the full description on the dataset page: https://huggingface.co/datasets/mishl/StreetVision-10K.vector-100k
VectorOS Vector 100k SimSat VLM Dataset
VectorOS Vector 100k is a high-fidelity multimodal instruction dataset for fine-tuning vision-language models on geospatial epidemiology tasks. It was built for the VectorOS hackathon project and targets LiquidAI/LFM2.5-VL-450M.
The dataset contains 100,000 chat-style examples derived from 10,000 geospatial chips across 30 AOIs. Every accepted chip has a real SimSat Sentinel-2 true-color view, a real SimSat Sentinel-2 NIR-red-green false-color… See the full description on the dataset page: https://huggingface.co/datasets/Alfaxad/vector-100k.bleedingheart-pretrain-10MBleedingheart Pretrain Dataset
A collaboration between Kaleido and Newstar
We collected all the datasets we could find that are in Tagalog or any other Philippine dialect and put them in this repository.
This data will be used to train the Bleedingheart model.
Bleeding Heart is a stunning bird native to the island of Luzon in the Philippines. It is a medium-sized ground dove with a distinctive red patch of feathers on its chest, which gives it its name. The male's red patch is… See the full description on the dataset page: https://huggingface.co/datasets/NewstaR/bleedingheart-pretrain-10M.3bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.
Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/Daniel10004/Nemotron-Personas-Korea.reasoning-10k-v2
🧠 Dataset for Vision-Language Reasoning
Reasoning-10K-v2 is a 10,000-sample multimodal reasoning dataset. Our work unifies diverse sources of vision–language reasoning data spanning code, mathematics, geography, tables, and scientific figures. Using a hybrid of synthesis and filtration strategies inspired by prior baselines (LIMO VL, MM MathInstruct, and Multimodal Open R1), we curated high-quality reasoning examples from open and synthetic data. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/ArkaMukherjee/reasoning-10k-v2.recipe-synthetic-images-10k
Recipe PDF Dataset
A multimodal dataset of 10K+ recipes rendered as PDF images with full metadata.
Dataset Description
Each sample contains:
image: Recipe rendered as a styled PDF page (PNG, ~1654x2339px)
name: Recipe title
description: Recipe description
ingredients: List of ingredients
steps: Cooking instructions
nutrition: Nutritional values (calories, fat%, sugar%, sodium%, protein%, sat.fat%, carbs%)
random_reviews: User reviews
minutes: Cooking time
tags: Recipe… See the full description on the dataset page: https://huggingface.co/datasets/TurkishCodeMan/recipe-synthetic-images-10k.gemma4-multimodal-recipe-dataset
🍳 Gemma 4 Multimodal Recipe & Food Dataset
A balanced, high-density multimodal dataset curated specifically for fine-tuning compact vision-language models (such as gemma-4-e2b-it) for visual food recognition, recipe generation, and dietary recommendation.
🔗 Upstream & Source Datasets
This dataset was created by cleaning, reformatting, and synthesizing samples across the following 5 Hugging Face sources:
Dataset
Modality
Role in Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/alst10/gemma4-multimodal-recipe-dataset.C3BEnglish | 简体中文
C³B: Comics Cross-Cultural Benchmark
Culture In a Frame: C³B as a Comic-Based Benchmark for Multimodal Cultural Awareness
ICLR 2026
About C³B
C³B (Comics Cross-Cultural Benchmark) is a multicultural, multitask, and multilingual benchmark for evaluating cultural awareness capabilities of Multimodal Large Language Models (MLLMs).
Progressive task difficulty: From basic visual recognition, to higher-level cultural conflict understanding, to cultural content… See the full description on the dataset page: https://huggingface.co/datasets/Coder109/C3B.Demo101
Overview 🚀📚
The first comprehensive dataset for training AI models to write complete novels with sophisticated reasoning.
🧠 Hierarchical Reasoning Architecture — Multi-layered planning traces including character archetypes, story arcs, world rules, and scene breakdowns. A complete cognitive roadmap for long-form narrative construction.
📖 Complete Novel Coverage — From 40,000 to 600,000+ tokens per book, spanning novellas to epic series with consistent quality throughout.
⚡… See the full description on the dataset page: https://huggingface.co/datasets/DemoTest0122/Demo101.Demo103
Overview 🚀📚
The first comprehensive dataset for training AI models to write complete novels with sophisticated reasoning.
🧠 Hierarchical Reasoning Architecture — Multi-layered planning traces including character archetypes, story arcs, world rules, and scene breakdowns. A complete cognitive roadmap for long-form narrative construction.
📖 Complete Novel Coverage — From 40,000 to 600,000+ tokens per book, spanning novellas to epic series with consistent quality throughout.
⚡… See the full description on the dataset page: https://huggingface.co/datasets/DemoTest0122/Demo103.Demo102
Overview 🚀📚
The first comprehensive dataset for training AI models to write complete novels with sophisticated reasoning.
🧠 Hierarchical Reasoning Architecture — Multi-layered planning traces including character archetypes, story arcs, world rules, and scene breakdowns. A complete cognitive roadmap for long-form narrative construction.
📖 Complete Novel Coverage — From 40,000 to 600,000+ tokens per book, spanning novellas to epic series with consistent quality throughout.
⚡… See the full description on the dataset page: https://huggingface.co/datasets/DemoTest0122/Demo102.synthetic-users-1000
Synthetic Users 1000
🧪 Synthetic Users Dataset (1,000 Records)
This dataset contains 1,000 high-quality synthetic user profiles, including realistic bios, usernames, metadata, and profile image filenames — ideal for:
UI/UX testing
Prototyping dashboards
Mock data in SaaS apps
User flow demos
AI evaluation without PII
🧠 Features
✅ 100% privacy-safe
🧍 Full names, emails, bios, countries
📸 Profile image filenames (S3-ready)
🔍 Gender, age, emotion tags via… See the full description on the dataset page: https://huggingface.co/datasets/eosync/synthetic-users-1000.100K_Cerebral_Palsy_v1.0
Synthetic ROCS Dataset (100K Patients)
Overview
This dataset is a fully synthetic reproduction of the patient cohort described in A Big Data Approach to Evaluate Receipt of Optimal Care in Childhood Cerebral Palsy.
It simulates the Receipt of Optimal Care Score (ROCS) framework across 100K synthetic patients, modeling event sequences, adherence probabilities, and component-level quality-of-care scores.
All data were algorithmically generated — no real patient data are… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/100K_Cerebral_Palsy_v1.0.
