CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dalle-mini /witimage1M<n<10M7 likes137k downloads5y agoHugging Face02mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face03mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes16k downloads2y agoHugging Face04mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes14k downloads2y agoHugging Face05osunlp /Multimodal-Mind2Web Dataset Summary Multimodal-Mind2Web is the multimodal version of Mind2Web, a dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website. In this dataset, we align each HTML document in the dataset with its corresponding webpage screenshot image from the Mind2Web raw dump. This multimodal version addresses the inconvenience of loading images from the ~300GB Mind2Web Raw Dump. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/Multimodal-Mind2Web.image10K<n<100K99 likes9.7k downloads2y agoHugging Face06minhdinh2 /Unsafe2Safeimage10K<n<100K0 likes8.7k downloads6mo agoHugging Face07ckadirt /imagery_mindbridgeimage1K<n<10K0 likes7.6k downloads1y agoHugging Face08mlfoundations /MINT-1T-ArXiv 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.imageimage-to-text1M<n<10M61 likes7.5k downloads2y agoHugging Face09Ming0614 /2026-01-29T15-21-58plus00-00_gdpvaldocumentn<1K0 likes6.8k downloads8mo agoHugging Face10timm /mini-imagenet Dataset Description A mini version of ImageNet-1k with 100 of 1000 classes present. Unlike some 'mini' variants this one includes the original images at their original sizes. Many such subsets downsample to 84x84 or other smaller resolutions. Data Splits Train 50000 samples from ImageNet-1k train split Validation 10000 samples from ImageNet-1k train split Test 5000 samples from ImageNet-1k validation split (all 50 samples per class)… See the full description on the dataset page: https://huggingface.co/datasets/timm/mini-imagenet.imageimage-classification10K<n<100K28 likes6.8k downloads2y agoHugging Face11RL-MIND /XHRBench XHRBench Ultra-High-Resolution Remote Sensing Understanding and Reasoning 🤗 Hugging Face · 🤖 ModelScope · 📄 Paper · 💻 Code English | 中文 📚 Introduction XHRBench evaluates fine-grained perception and complex reasoning in multimodal large language models using ultra-high-resolution remote-sensing imagery. This repository retains the name XHRBench and belongs to the same RSHR benchmark project as RSHR-Bench, with a… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/XHRBench.imageimage-text-to-text1K<n<10K7 likes6.6k downloads8d agoHugging Face12mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes6k downloads2y agoHugging Face13GaussianWorld /scannet_mini_val_set_suiteimage1 likes5.8k downloads1y agoHugging Face14gagandeepreehal /minuszero-indian-autonomous-driving-dataset-v2gated INDUS-AD: Indian Dataset of Unstructured Urban Scenes for Autonomous Driving Overview INDUS-AD is the largest publicly released Indian autonomous-driving dataset for end-to-end autonomous-driving research. Its name expands to Indian Dataset of Unstructured Urban Scenes for Autonomous Driving. This gated dataset is the decoded companion to the Minus Zero Indian Urban Autonomous Driving Dataset. It provides directly usable camera MP4s, normalized sensor tables… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset-v2.imagerobotics10M<n<100M1 likes5.2k downloads5d agoHugging Face15Smith42 /minty-astro-ph MINT-1T ArXiv Astro-ph An astronomy-focused subset of mlfoundations/MINT-1T-ArXiv, filtered to include only papers from the astro-ph arXiv category (including cross-listed papers). Overview Papers ~845k Total size ~804 GB Format WebDataset tar shards Shards 287 (astro-ph-00000.tar to astro-ph-00286.tar) Shard size ~3 GB each Source MINT-1T (Awadalla et al., 2024) Data Format Each tar shard contains paired files per paper:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/minty-astro-ph.imagetext-generation100K<n<1M1 likes4.6k downloads5mo agoHugging Face16pollen-robotics /reachy-mini-wall-data Reachy Mini — wall data (public) posts.json for the Reachy Mini community wall: the AI-filtered posts shown publicly, aggregated from Bluesky, YouTube, LinkedIn, TikTok, X and Reddit by the social-wall pipeline. Fetch it directly (CORS-enabled) from any static site: const url = "https://huggingface.co/datasets/pollen-robotics/reachy-mini-wall-data/resolve/main/posts.json"; const posts = await (await fetch(url)).json(); Each item: id, platform, author, handle, avatar, text… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-wall-data.imagen<1K0 likes3.4k downloads1d agoHugging Face17sled-umich /MindCraftimagen<1K0 likes3.1k downloads3y agoHugging Face18sled-umich /MindCraft2image0 likes2.5k downloads3y agoHugging Face19antofuller /mini-VTAB Mini-VTAB A collection of VTAB (Visual Task Adaptation Benchmark) datasets. We sampled 1K training samples and 1K testing samples for each task. Tasks datasets = [ "caltech101", "cifar10", "cifar100", "dtd", "flowers", "pets", "sun397", "svhn", "pcam", "eurosat", "resisc45", "diabetic_retinopathy", "clevr_count_all", "clevr_closest_object_distance", "dmlab", "dsprites_label_x_position"… See the full description on the dataset page: https://huggingface.co/datasets/antofuller/mini-VTAB.image10K<n<100K1 likes2.3k downloads8mo agoHugging Face20Minoday /Robo4D-200k Paper Project Page Model Repo imagen<1K6 likes2.3k downloads6mo agoHugging Face21minhnguyent546 /TranNhiem-Vietnamese-ImageText-Reasoning TranNhiem Vietnamese Image-Text Reasoning (V-LAION) Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set. Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.imagevisual-question-answering100K<n<1M0 likes2.3k downloads2mo agoHugging Face22badincite /minimax-h3-soup MiniMax H3 Soup Reproducibility archive for a local ComfyUI MiniMax H3 Ref2V benchmark on an RTX 3090. What is included Original benchmark workflow graph (source_prompt.json), manifest, and result table. Every one-second MP4 from the original C1-C11 benchmark grid and its Euler repeat sweep. The separate duration experiments are intentionally not included. Labeled C1-C11 visual contact sheets, Ref2VA stock-control sheets, and a static render-time summary chart.… See the full description on the dataset page: https://huggingface.co/datasets/badincite/minimax-h3-soup.imagen<1K5 likes2.2k downloads13d agoHugging Face23Mingde /PolaRGB PolarFree: Polarization-based Reflection-Free Imaging Dataset Overview PolarFree is a high-quality dataset designed for polarization-based reflection removal tasks, as introduced in the CVPR 2025 paper "PolarFree: Polarization-based Reflection-Free Imaging". The dataset aims to support tasks such as image reflection removal and image enhancement, particularly suitable for training and evaluating polarization-based image reflection removal models. Download Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mingde/PolaRGB.imageimage-to-image1K<n<10K2 likes2.2k downloads10mo agoHugging Face24dalle-mini /open-imagesimage1M<n<10M27 likes2k downloads5y agoHugging Face25Mindykkyan /PittadsDB-AdsPics Pitt Ads Dataset (PittAdsDB) This folder holds the University of Pittsburgh Ads dataset artifacts we can download publicly, and a helper script to fetch images once access is granted. What is included (downloaded now) docs/ readme_images.txt (dataset readme for images) readme_videos.txt (dataset readme for videos) image_annotations/ image/QA_Action.json image/QA_Combined_Action_Reason.json image/QA_Reason.json image/Sentiments.json image/Sentiments_List.txt… See the full description on the dataset page: https://huggingface.co/datasets/Mindykkyan/PittadsDB-AdsPics.image100K<n<1M1 likes1.6k downloads3mo agoHugging Face26Mini-o3 /VisualProbe_trainimage1K<n<10K3 likes1.6k downloads1y agoHugging Face27Ming0614 /2026-01-29T07-51-20plus00-00_gdpvaldocumentn<1K0 likes1.6k downloads8mo agoHugging Face28oscarqjh /MindCube_lmmseval MindCube LMMs Eval Dataset This dataset is formatted for use with lmms-eval framework. Dataset Schema Column Type Description id string Unique identifier for each sample (format: {split}_{scene_id}_{question_id}) category list[string] Category labels (e.g., ["perpendicular", "P-O", "meanwhile", "self"]) type string Question type (e.g., "1_frame", "2_frame", "3_frame", "general") meta_info list[list[string]] Metadata about scene objects and their spatial… See the full description on the dataset page: https://huggingface.co/datasets/oscarqjh/MindCube_lmmseval.image10K<n<100K1 likes1.6k downloads8mo agoHugging Face29minhletran /BDD100K-laneimage10K<n<100K0 likes1.5k downloads1y agoHugging Face30vanyacohen /MET-Bench-Minecraft MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models Vanya Cohen and Raymond Mooney · ICML 2026 Paper · Publication page · Load the dataset · Citation Domains: Chess · Shell Game · Minecraft MET-Bench evaluates entity state tracking across text and image modalities. This repository contains the Minecraft domain. Minecraft Minecraft is a state prediction task involving partial observations, dynamic… See the full description on the dataset page: https://huggingface.co/datasets/vanyacohen/MET-Bench-Minecraft.image1K<n<10K0 likes1.5k downloads5h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.