CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face02mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes17k downloads2y agoHugging Face03mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes15k downloads2y agoHugging Face04mlfoundations /MINT-1T-ArXiv 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.imageimage-to-text1M<n<10M61 likes8.3k downloads2y agoHugging Face05mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes6k downloads2y agoHugging Face06BUAADreamer /MMPR-v1.1This dataset is borrowed from OpenGVLab/MMPR-v1.1 imagetext-generation1M<n<10M1 likes5.2k downloads2y agoHugging Face07OpenDataArena /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M126 likes1.8k downloads7mo agoHugging Face08NarsAI /FineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1.4k downloads8mo agoHugging Face09NuTonic /sat-vl-sft-training-ready-v1 Dataset Summary NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills). The goal is to create high-signal, production-shaped supervision for multimodal chat models: Captioning for satellite chips Grounding (bounding boxes in normalized coordinates) for land-cover regions Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.imagetext-generation100K<n<1M2 likes1.3k downloads5mo agoHugging Face10yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.2k downloads7mo agoHugging Face11pin-team /oercommons-v1-optimized OERCommons v1 Optimized Authors: Junjie Wang and Yuhan SunHosted by: PIN TeamDataset: pin-team/oercommons-v1-optimized OERCommons v1 Optimized is a provenance-preserving multimodal pretraining corpus built on the OERCommons subset of The Common Pile v0.1, which serves as its upstream data and licensing baseline. We extend it with full-page recovery, canonical Markdown, ordered image/PDF/link metadata, conservative corrections, and integrity evidence. At a glance… See the full description on the dataset page: https://huggingface.co/datasets/pin-team/oercommons-v1-optimized.documenttext-generation1K<n<10K2 likes1.1k downloads2mo agoHugging Face12youssef101 /artelingo-dummyArtELingo is a benchmark and dataset introduced in a research paper aimed at promoting work on diversity across languages and cultures. It is an extension of ArtEmis, which is a collection of 80,000 artworks from WikiArt with 450,000 emotion labels and English-only captions. ArtELingo expands this dataset by adding 790,000 annotations in Arabic and Chinese. The purpose of these additional annotations is to evaluate the performance of "cultural-transfer" in AI systems. The dataset in ArtELingo… See the full description on the dataset page: https://huggingface.co/datasets/youssef101/artelingo-dummy.imageimage-to-text10K<n<100K2 likes1k downloads3y agoHugging Face13Sandeepthakur /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Sandeepthakur/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1k downloads8mo agoHugging Face14schneewolflabs /ArtemisMix-v1 ArtemisMix-v1 Stage-2 (multimodal instruction fine-tuning) corpus for Artemis, the Schneewolf Labs vision-language flagship that grafts a Qwen3-VL ViT + MLP projector onto the A2 decoder (the A3 lineage). This is the Lite slice (v1): layers L1 + L4 only, 350,000 rows. The planned L2 (multimodal tool/agent) and L3 (custom distill) layers are produced separately and concatenated later. Composition Layer Rows Share Purpose L1 general multimodal instruction… See the full description on the dataset page: https://huggingface.co/datasets/schneewolflabs/ArtemisMix-v1.imagevisual-question-answering100K<n<1M0 likes987 downloads4mo agoHugging Face15Cameronk199 /donald-trump-truth-social-posts Donald Trump Truth Social Posts Archive Archive overview 36,170 public Truth Social posts associated with Donald J. Trump's @realDonaldTrump account. The release preserves source URLs, timestamps, post types, original HTML, extracted plain text, attachment provenance, and analysis-ready tables. It also includes streamable image media plus video metadata and transcripts where the source provides them. The package is source-linked and reconciled by archive ID.… See the full description on the dataset page: https://huggingface.co/datasets/Cameronk199/donald-trump-truth-social-posts.imagetext-generation100K<n<1M3 likes858 downloads10d agoHugging Face16ChrisDing1105 /unified-agent-trajectories Unified Benchmark Agent Trajectories Dataset release: v2.1.1 (2026-09-18)Record format: unified-agent-sft-v1 A growing collection of benchmark agent execution trajectories converted into one transparent, multimodal, tool-aware representation. These are complete recorded benchmark runs—not ordinary chat transcripts—including benchmark tasks, model reasoning and answers, tool calls, tool observations, runtime status, and benchmark scores when available. The directory layout is… See the full description on the dataset page: https://huggingface.co/datasets/ChrisDing1105/unified-agent-trajectories.imagetext-generation1K<n<10K3 likes857 downloads6d agoHugging Face17zabir1996 /mimic-medical-imaging-qa MIMIC Medical Imaging QA Dataset 5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction. License The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.imagequestion-answering1K<n<10K3 likes796 downloads5mo agoHugging Face18Mike1997126 /All-Prompt-Jailbreakimagetext-generationn<1K1 likes621 downloads8mo agoHugging Face19kernel-14 /SemanticAlign-Bench SemanticAlign-Bench A benchmark for evaluating AI agents on structured claim extraction from top-tier ML conference papers. Each paper is decomposed into Semantic Alignment Units (SAU) — atomic, self-contained implementation propositions — across four diagnostic dimensions spanning numerical precision to pipeline-level workflow. Agents are evaluated on whether they can reproduce these claims without hallucination, omission, or misordering. The Four SAU Dimensions… See the full description on the dataset page: https://huggingface.co/datasets/kernel-14/SemanticAlign-Bench.imagequestion-answering1K<n<10K1 likes573 downloads4mo agoHugging Face20ppak10 /AdditiveLLM2-OA AdditiveLLM2-OA Dataset Open Access journal articles (up to February 2026) used in domain adapting pretraining and instruction tuning for AdditiveLLM2. Dataset Split by Journal text images vit Vocabulary Overlap Pairwise Jaccard similarity of word-level vocabularies (lowercase, 3+ letter tokens) across the four source journals. Run info/vocabulary/vocabulary_overlap.py to reproduce. Top Phrases by Journal Most frequent bigrams and… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/AdditiveLLM2-OA.imagetext-generation10K<n<100K2 likes561 downloads6mo agoHugging Face21artefactory /Argimi-Ardian-Finance-10k-text-image The ArGiMI Ardian datasets : text and images The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.imagetext-retrieval1M<n<10M14 likes530 downloads7mo agoHugging Face22OpenDataArena /MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking MMFineReason-SFT-123K The Hardest 7% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0). 🎯 Key Highlights 123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.imagevisual-question-answering100K<n<1M86 likes525 downloads8mo agoHugging Face23rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes511 downloads6mo agoHugging Face24NuTonic /sat-vl-sft-postprocessed-merged-v1 Dataset Summary NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills). The goal is to create high-signal, production-shaped supervision for multimodal chat models: Captioning for satellite chips Grounding (bounding boxes in normalized coordinates) for land-cover regions Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-postprocessed-merged-v1.imagetext-generation100K<n<1M0 likes505 downloads5mo agoHugging Face25daruokta /t5gemma2-indonesia-instruct-v1 T5Gemma-2 Indonesian Instruct — Mono-Repo Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia. Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder, setiap config = folder dan berisi split train + validation (80:20) di level percakapan. Struktur (by fungsi) t5gemma2-indonesia-instruct-v1/ ├── README.md ├── manifest.json ├── chat_idx_map.json ├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.imagetext-generation100K<n<1M0 likes503 downloads1d agoHugging Face26yanivohayon1 /transfertalk-players TransferTalk Players Dataset A fully synthetic, fictional multimodal dataset of 1,000 football (soccer) player profiles, created for a university Data Science capstone project. Each profile pairs a rich text description with a matching illustrated player card image. No real players, teams, or leagues are represented. All names, teams, leagues, and nationalities belong to a fictional universe ("The Meridian League"), generated to avoid any real-world IP, privacy, or… See the full description on the dataset page: https://huggingface.co/datasets/yanivohayon1/transfertalk-players.imagetext-generation1K<n<10K0 likes440 downloads1mo agoHugging Face27ed001 /ds-coder-instruct-v1 Dataset Card for DS Coder Instruct Dataset DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python. The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.imagetext-generation10K<n<100K5 likes392 downloads3y agoHugging Face28MilitaryHospital175 /VNMedical_bv175 ViX-Ray: A Vietnamese Chest X-Ray Dataset for Vision-Language Models ViX-Ray is the first publicly available Vietnamese chest X-ray dataset designed for vision-language model (VLM) research in medical AI. It pairs radiographic images from Vietnamese patients with expert-written clinical annotations in Vietnamese, directly addressing the lack of Vietnamese medical data in existing VLMs. License: CC BY-NC-SA 4.0 Paper: arXiv:2603.15513 Dataset Overview Each sample… See the full description on the dataset page: https://huggingface.co/datasets/MilitaryHospital175/VNMedical_bv175.imagetext-generation1K<n<10K1 likes388 downloads5mo agoHugging Face29abhinav00anand /behavioral-fine-tuning-v1 Why This Dataset Exists "A model that refuses everything is useless. A model that refuses nothing is dangerous. The goal is a model that thinks." The Problem Our Solution Uncensored data → helpful but uncontrolled Surgical 85% helpfulness + 13% safety + 2% eval mix Safety-only data → lobotomized, over-refusing models Calibrated ratio preserves full helpfulness Raw data → PII, leaked secrets, duplicates 7-stage pipeline validates every… See the full description on the dataset page: https://huggingface.co/datasets/abhinav00anand/behavioral-fine-tuning-v1.imagetext-generation100K<n<1M1 likes381 downloads28d agoHugging Face30latam-gpt /LatamGPT-Corpus-1.0gated LatamGPT-Corpus-1.0 🌐 Language versions: English | Español | Português 🔗 Project links: Official LatamGPT website | Corpus dashboard 🤖 Associated model: The complete LatamGPT corpus—of which this repository contains the openly released portion—was used in the training process of Llama-3.1-70B-LatamGPT-SFT-1.0. Dataset description Summary LatamGPT-Corpus-1.0 is the open release of the data corpus assembled for the continued pretraining of… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/LatamGPT-Corpus-1.0.imagetext-generation100M<n<1B8 likes378 downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.