CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mvp-lab /Sekaitext1M<n<10M0 likes24k downloads10mo agoHugging Face02mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face03ma-xu /fine-t2i Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning [arxiv] by Xu Ma, Yitian Zhang, Qihua Dong, Yun Fu Northeastern Univeristy Please see our [Dataset Explore] to view detailed samples (loading is slow, be patient). 🆕 What's New [2026.02.20]: Fine-T2I reaches the #1 spot among Hugging Face Datasets Trending list ⭐️⭐️⭐️ [2026.02.16]: Fine-T2I tops the Hugging Face Datasets Trending list, reaching the #2 spot and #1… See the full description on the dataset page: https://huggingface.co/datasets/ma-xu/fine-t2i.imageimage-to-text100K<n<1M120 likes20k downloads7mo agoHugging Face04mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes16k downloads2y agoHugging Face05mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes14k downloads2y agoHugging Face06sarulab-speech /mls_sidon MLS-Sidon Overview This dataset is a cleansed version of Multilingual LibriSpeech (MLS) with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. The dataset is provided in WebDataset format for efficient large-scale training. Source: Multilingual LibriSpeech Languages: English, German, French, Spanish, Italian, Polish, Dutch, Portuguese Format: WebDataset (.tar shards) License: CC-BY-4.0 Dataset Structure Each sample in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/mls_sidon.audiotext-to-speech10M<n<100M11 likes9.8k downloads1y agoHugging Face07mlfoundations /MINT-1T-ArXiv 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.imageimage-to-text1M<n<10M61 likes7.5k downloads2y agoHugging Face08Amshaker /Mobile-O-Post-Train Mobile-O Post-Training Data Unified Multimodal Post-Training · ~105K Quadruplet Samples 📌 Overview This dataset is used for Stage 3: Unified Multimodal Post-Training of Mobile-O, a unified multimodal model for on-device understanding and generation. The goal of this stage is to jointly improve both image generation and visual understanding through a multi-task objective using quadruplet samples. 📊 Dataset Format Each sample is a quadruplet consisting of:… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Post-Train.imagetext-to-image1K<n<10K13 likes6.7k downloads7mo agoHugging Face09mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes6k downloads2y agoHugging Face10MJJJJ1064 /FTP-1-Dataset FTP-1-Dataset FTP-1-Dataset contains heterogeneous tactile manipulation data for FTP-1 pretraining. This release currently includes 18 dataset archives: FreeTacMan MotionTrans RDP RDP_Bimanual RH20TCfg5Franka RH20TCfg6ATIAxia RH20TCfg7Tactile Unit Unit_Bimanual VLA_touch ViTaMIn VisuoTactile_D-WHEEL VisuoTactile_QINGLOONG exUMI sharpa Each dataset directory contains either a single <dataset>.tar file or split parts named <dataset>.tar.part-*. For split archives, concatenate… See the full description on the dataset page: https://huggingface.co/datasets/MJJJJ1064/FTP-1-Dataset.text1M<n<10M3 likes6k downloads3mo agoHugging Face11Smith42 /minty-astro-ph MINT-1T ArXiv Astro-ph An astronomy-focused subset of mlfoundations/MINT-1T-ArXiv, filtered to include only papers from the astro-ph arXiv category (including cross-listed papers). Overview Papers ~845k Total size ~804 GB Format WebDataset tar shards Shards 287 (astro-ph-00000.tar to astro-ph-00286.tar) Shard size ~3 GB each Source MINT-1T (Awadalla et al., 2024) Data Format Each tar shard contains paired files per paper:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/minty-astro-ph.imagetext-generation100K<n<1M1 likes4.6k downloads5mo agoHugging Face12mitermix /audiosnippetsaudio1M<n<10M6 likes4.4k downloads2y agoHugging Face13Yale-BIDS-Chen /medpmc-11m-dataset_jun24_baseline MedPMC WebDataset MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources. This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.imagezero-shot-image-classification1M<n<10M3 likes3.9k downloads2mo agoHugging Face14MAmmoTH-VL /MAmmoTH-VL-Instruct-12M MAmmoTH-VL-Instruct-12M 🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo Introduction Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses. The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.imagevisual-question-answering10M<n<100M67 likes3.7k downloads2y agoHugging Face15lighthouse-emnlp2024 /Clotho-Moment Clotho-Moment This repository provides wav files used in Language-based Audio Moment Retrieval. Each sample includes long audio containing some audio events with the temporal and textual annotation. Project page: https://h-munakata.github.io/Language-based-Audio-Moment-Retrieval/ Code: https://github.com/line/lighthouse Split Train train/train-{000..715}.tar 37930 audio samples Valid valid/valid-{000..108}.tar 5741 audio samples Test test/test-{000..142}.tar 7569… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/Clotho-Moment.audioaudio-text-to-text10K<n<100K2 likes3.6k downloads8mo agoHugging Face16LLMDH /marianne_pdf_7text10K<n<100K0 likes3.1k downloads2y agoHugging Face17erickfm /melee-ranked-replays Melee Ranked Replays Anonymized Slippi ranked replays (platinum+) from Super Smash Bros. Melee, sharded by character and rank pair. Built for behavior-cloning and other replay-driven ML work on Melee — notably MIMIC. Contents Raw .slp files grouped into tarballs by (character, rank_pair, source_archive), organized into per-character folders: {CHAR}/ {CHAR}_{rank_pair}_a{N}.tar.gz metadata/ metadata_a{N}.json Characters (25): BOWSER, CPTFALCON, DK, DOC, FALCO… See the full description on the dataset page: https://huggingface.co/datasets/erickfm/melee-ranked-replays.textreinforcement-learning1M<n<10M1 likes2.9k downloads3mo agoHugging Face18mitermix /audiosnippets_small_with_detailed_annotationaudio100K<n<1M1 likes2.7k downloads2y agoHugging Face19mitermix /audiosnippets_small_with_detailed_annotation2audio1M<n<10M1 likes2.7k downloads2y agoHugging Face20Amshaker /Mobile-O-Pre-Train Mobile-O Pre-Training Data Cross-Modal Alignment · 9M Text-Image Pairs 📌 Overview This dataset is used for Stage 1: Cross-Modal Alignment pre-training of Mobile-O, a unified multimodal model for on-device understanding and generation. The goal of this stage is to align the DiT diffusion decoder and Mobile Conditioning Projector (MCP) with the frozen VLM backbone using large-scale text-image pairs. 📊 Dataset Composition Source Samples Description… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Pre-Train.imagetext-to-image10M<n<100M12 likes2.6k downloads7mo agoHugging Face21LLMDH /marianne_pdf_9text100K<n<1M0 likes2.6k downloads2y agoHugging Face22ScienceOne-AI /S1-MMAlignS1-MMAlign A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers. Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1-MMAlign aims to… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.imageimage-to-text10M<n<100M106 likes2.5k downloads6mo agoHugging Face23TTS-AGI /majestrino-unified-detailed-captions Majestrino Unified Detailed Captions Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption. Stats 4,658,407 samples 932 tar files (~1.1 GB each) ~1,017 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption caption_type — always unified_detailed_caption transcription — speech transcription (when available, normalized from multiple source keys) duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.audioaudio-classification1M<n<10M3 likes2.4k downloads6mo agoHugging Face24LLMDH /marianne_pdf_5text100K<n<1M0 likes2.2k downloads2y agoHugging Face25LLMDH /marianne_pdf_3text100K<n<1M0 likes2.2k downloads2y agoHugging Face26laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M15 likes2.1k downloads11mo agoHugging Face27gaunernst /ms1mv3-wds MS-Celeb-1M (v3) This dataset is introduced in the Lightweight Face Recognition Challenge at ICCV 2019. Paper. There are 5,179,510 images and 93,431 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112. This was downloaded from https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_ (MS1M-RetinaFace). The original dataset format is MXNet RecordIO. It was converted to WebDataset in this copy here. There are 100… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/ms1mv3-wds.imageimage-classification100K<n<1M0 likes1.9k downloads2y agoHugging Face28allenai /Molmo2-ER-RoboPoint Molmo2-ER · wentao-yuan/robopoint-data 1.43M robotics affordance instruction-tuning examples (pointing + detection + VQA). This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified. Upstream source Original dataset: wentao-yuan/robopoint-data Paper: RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics (arXiv:2406.10721) License:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-RoboPoint.image1M<n<10M1 likes1.9k downloads5mo agoHugging Face29kohei0209 /mls_hq_urgent_track1audio100K<n<1M0 likes1.9k downloads2y agoHugging Face30laion /majestrino-dataaudio1M<n<10M1 likes1.7k downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.