CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /dcvlm-baseline-200b DCVLM-Baseline (200B tokens) DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token dataset release consisting of 103,985,276 samples, curated from our DCVLM-large data pool. A smaller 6.25B-token version is also available. ⚠️ NOTE: The training data is the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-200b.imageimage-text-to-text10K<n<100K8 likes149k downloads2mo agoHugging Face02mlfoundations /datacomp_pools DataComp Pools This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.image19 likes88k downloads3y agoHugging Face03mlfoundations /datacomp_xlarge DataComp XLarge Pool This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.image10B<n<100B21 likes39k downloads3y agoHugging Face04mlfoundations /dcvlm-balanced-200b DCVLM-Balanced (200B tokens) DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool. The instruction-heavy counterpart (DCVLM-baseline) is available as dcvlm-baseline-200b, along with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.imageimage-text-to-text10K<n<100K1 likes30k downloads2mo agoHugging Face05MLCommons /speech-wikimedia Dataset Card for Speech Wikimedia Dataset Summary The Speech Wikimedia Dataset is a compilation of audiofiles with transcriptions extracted from wikimedia commons that is licensed for academic and commercial usage under CC and Public domain. It includes 2,000+ hours of transcribed speech in different languages with a diverse set of speakers. Each audiofile should have one or more transcriptions in different languages. Transcription languages English German… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/speech-wikimedia.audion<1K14 likes23k downloads3y agoHugging Face06jablonkagroup /chempile-mlift ChemPile-MLIFT A comprehensive multimodal dataset for chemistry property prediction using vision large language models 📋 Dataset Summary ChemPile-MLIFT is a dataset designed for multimodal chemistry property prediction tasks, specifically focusing on the prediction of chemical properties using vision large language models (VLLMs). It is part of the ChemPile project, which aims to create a comprehensive collection of chemistry-related data for training LLMs. The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/chempile-mlift.imagetext-generation10M<n<100M14 likes22k downloads1y agoHugging Face07mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face08mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes16k downloads2y agoHugging Face09mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes10k downloads2y agoHugging Face10mlfoundations /MINT-1T-ArXiv 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.imageimage-to-text1M<n<10M61 likes7.6k downloads2y agoHugging Face11mlfoundations /dcvlm-baseline-6_25b DCVLM-Baseline (6.25B tokens) DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This dataset version is a small 6.25B-token (small-pool) release consisting of 3,253,356 samples. ⚠️ NOTE: The training data is the WebDataset shards under shards/. The preview config shown in the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-6_25b.imageimage-text-to-textn<1K1 likes6.9k downloads2mo agoHugging Face12mlfoundations /datacomp_large DataComp Large Pool This repository contains metadata files for the large pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_large.image1B<n<10B6 likes6.7k downloads3y agoHugging Face13mlfoundations /datacomp_1b DataComp-1B This repository contains metadata files for DataComp-1B. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_1b.image1B<n<10B53 likes6.5k downloads3y agoHugging Face14mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes6k downloads2y agoHugging Face15mlfoundations /DataComp-12M Dataset Card for DataComp-12M This dataset contains a 12M subset of DataComp-1B-BestPool. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Image-text models trained on DataComp-12M are significantly better than on CC-12M/YFCC-15M as well as DataComp-Small/Medium. DataComp-12M was introduced in MobileCLIP paper and along with the reinforced dataset DataCompDR-12M. The UIDs… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/DataComp-12M.imagetext-to-image14 likes5.9k downloads2y agoHugging Face16GD-ML /MAPBench-V2For more details, please check our project page. Paper: https://arxiv.org/abs/2601.05432 Repository: https://github.com/AMAP-ML/Thinking-with-Map image1K<n<10K4 likes5.9k downloads8mo agoHugging Face17ifx-pse-sys-ml /FineVisionConcatShuffleIFXimage10M<n<100M0 likes4.8k downloads8mo agoHugging Face18zr-zhang /MLLM-Generated-Image-Detection-Dataset MLLM-Generated Image Dataset This dataset contains real and AI-generated image samples organized for binary MLLM-generated image detection. Paper | Code Dataset Summary We construct an MLLM-generated image detection benchmark from GPT Image2 and Nano Banana2. This benchmark covers texture-dominated, structure-dominated, and hybrid-dominated. It is designed to evaluate detector performance under the new challenges introduced by large-scale image generation models.… See the full description on the dataset page: https://huggingface.co/datasets/zr-zhang/MLLM-Generated-Image-Detection-Dataset.imageimage-classification1K<n<10K1 likes2.2k downloads2mo agoHugging Face19mlfoundations /datacomp_small DataComp Small Pool This repository contains metadata files for the small pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_small.image10M<n<100M6 likes1.9k downloads3y agoHugging Face20ml-infra-toloka /cord-receipt-imagesimage1K<n<10K0 likes1.5k downloads1mo agoHugging Face21mlfoundations /VisIT-Bench Dataset Card for VisIT-Bench Dataset Description Links Dataset Structure Data Fields Data Splits Data Loading Licensing Information Annotations Considerations for Using the Data Citation Information Dataset Description VisIT-Bench is a dataset and benchmark for vision-and-language instruction following. The dataset is comprised of image-instruction pairs and corresponding example outputs, spanning a wide range of tasks, from simple object recognition to complex… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/VisIT-Bench.imagen<1K16 likes1.4k downloads3y agoHugging Face22mlfoundations-cua-dev /easyr1-grounding-dataset-30k-not_grounded-SE-GUI-3B-2MPimage10K<n<100K1 likes1.3k downloads1y agoHugging Face23mlfoundations /datacomp_medium DataComp Medium Pool This repository contains metadata files for the medium pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_medium.image100M<n<1B3 likes1.3k downloads3y agoHugging Face24TimS-ml /trailogy-na-plantae-sft NA-Plantae SFT data tree Mixed supervised-finetuning corpus for the on-device Gemma 4 E2B "hike companion" VLM. Combines a North-American Plantae image-ID slice (sourced from iNaturalist, label text enriched via GBIF) with general anti-forgetting buckets (LLaVA-style image QA + refusal/negative). Layout inaturalist_na_plantae/ # web-crawled label sources (irreproducible) observations.jsonl # iNaturalist observation metadata… See the full description on the dataset page: https://huggingface.co/datasets/TimS-ml/trailogy-na-plantae-sft.imageimage-text-to-text10K<n<100K0 likes1.3k downloads4mo agoHugging Face25MLLM-CL /UCITUnofficial training-ready fork of HaiyangGuo/UCIT image100K<n<1M1 likes1.2k downloads6mo agoHugging Face26aharoon /fpp-ml-bench FPP-ML-Bench: Fringe Projection Profilometry Benchmarking Dataset The first open-source, photorealistic synthetic dataset for single-shot fringe projection profilometry (FPP), generated using VIRTUS-FPP in NVIDIA Isaac Sim. This dataset enables standardized benchmarking and systematic comparison of deep learning approaches for single-shot 3D depth reconstruction from fringe patterns. Dataset Summary Property Value Total fringe images 15,600 (52 per viewpoint… See the full description on the dataset page: https://huggingface.co/datasets/aharoon/fpp-ml-bench.imagedepth-estimation1K<n<10K3 likes1.1k downloads8mo agoHugging Face27brandonyang /metaworld_ml45-v2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "metaworld", "total_episodes": 4391, "total_frames": 358997, "total_tasks": 44, "total_videos": 0, "total_chunks": 5, "chunks_size": 1000, "fps": 80, "splits": { "train": "0:4391" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/metaworld_ml45-v2.imagerobotics100K<n<1M0 likes1k downloads7mo agoHugging Face28mlech26l /liquidrandom-data liquidrandom-data Diverse seed data for ML/LLM training data generation pipelines. Used by the liquidrandom Python package. Dataset Summary This dataset contains 520,080 seed data samples across 24 categories, generated using a hierarchical taxonomy tree approach with LLM-based quality validation and fuzzy deduplication. Data is stored as Parquet with zstd compression. Categories Category Samples File Coding Tasks 30,069… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/liquidrandom-data.tabulartext-generation100K<n<1M0 likes1k downloads2mo agoHugging Face29ONE-Lab /MLLM-as-a-Judgeimagequestion-answering1K<n<10K4 likes986 downloads2y agoHugging Face30MLLMMU /MLLMU-Bench Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench Abstract Generative models such as Large Language Models (LLM) and Multimodal Large Language models (MLLMs) trained on massive web corpora can memorize and disclose individuals' confidential and private data, raising legal and ethical concerns. While many previous works have addressed this issue in LLM via machine unlearning, it remains largely unexplored for MLLMs. To tackle this challenge, we… See the full description on the dataset page: https://huggingface.co/datasets/MLLMMU/MLLMU-Bench.image1K<n<10K6 likes954 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.