CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes32k downloads10mo agoHugging Face02mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face03mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes16k downloads2y agoHugging Face04adams-story /imagenet1k-256-wdsThis is imagenet1k in webdataset format. Images are stored as jpg files. Every image has been resized to a maximum side length of 256. That means that if an image in the original dataset was 1000 by 500, the new size will be 256 by 128. Images with a maximum side length of under 256 were not resized. The total size of all dataset files is 57.8 GB, there are 1,281,167 rows in the training split and 50,000 rows in the validation split. imageimage-classification100K<n<1M2 likes16k downloads1y agoHugging Face05vaishaal /ImageNetV2image10K<n<100K9 likes14k downloads4y agoHugging Face06mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes14k downloads2y agoHugging Face07gasstation /gs-videos-v2text10K<n<100K0 likes9.7k downloads8mo agoHugging Face08AILab-CVC /obelics_seed2_tokensPart of the OBELISC data set, including 32 Million samples, please refer to dataset.py to use this data text10M<n<100M1 likes9.4k downloads3y agoHugging Face09gasstation /gs-images-v2image100K<n<1M1 likes7.7k downloads8mo agoHugging Face10clip-benchmark /wds_imagenetv2image10K<n<100K0 likes6.8k downloads4y agoHugging Face11mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes6k downloads2y agoHugging Face12timm /imagenet-22k-wdsgated Dataset Summary This is a copy of the full ImageNet dataset consisting of all of the original 21841 clases. It also contains labels in a separate field for the '12k' subset described at at (https://github.com/rwightman/imagenet-12k, https://huggingface.co/datasets/timm/imagenet-12k-wds) This dataset is from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000 synsets containing people, a number of these are of an offensive… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-22k-wds.imageimage-classification100K<n<1M14 likes5.9k downloads3y agoHugging Face13yuanty /robotwin2.0-fastwam RobotWin 2.0 (Preprocessed LeRobot v2.1 Release) This repository releases our preprocessed RoboTwin / RobotWin 2.0 dataset in LeRobot v2.1 format for the open-source release of Fast-WAM: Do World Action Models Need Test-time Future Imagination? This is not the official upstream RoboTwin release. It is our paper-specific processed version prepared to support training, evaluation, and reproducibility for our project. This Hugging Face repository distributes the dataset as split… See the full description on the dataset page: https://huggingface.co/datasets/yuanty/robotwin2.0-fastwam.text10K<n<100K2 likes5.9k downloads6mo agoHugging Face14speechcolab /gigaspeech2gated Dataset Card for GigaSpeech 2 Dataset Description GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. Repository: https://github.com/SpeechColab/GigaSpeech2 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.audioautomatic-speech-recognition10M<n<100M71 likes5.8k downloads6mo agoHugging Face15cschell /boxrr-23 BOXRR-23: Berkeley Open Extended Reality Recording Dataset 2023 This is a copy of the official Berkeley Open Extended Reality Recording Dataset 2023 (BOXRR-23). Please visit the project website for more information. In users/ you find one tarball for each user (which you can untar with tar xvf <path/to/user.tar>), which includes all replays of that user. Each replay is stored in a dedicated file in the XROR format. Metadata The entire dataset is around 5 TB large… See the full description on the dataset page: https://huggingface.co/datasets/cschell/boxrr-23.text1K<n<10K3 likes4.2k downloads2y agoHugging Face16Yale-BIDS-Chen /medpmc-11m-dataset_jun24_baseline MedPMC WebDataset MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources. This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.imagezero-shot-image-classification1M<n<10M3 likes3.9k downloads2mo agoHugging Face17lighthouse-emnlp2024 /Clotho-Moment Clotho-Moment This repository provides wav files used in Language-based Audio Moment Retrieval. Each sample includes long audio containing some audio events with the temporal and textual annotation. Project page: https://h-munakata.github.io/Language-based-Audio-Moment-Retrieval/ Code: https://github.com/line/lighthouse Split Train train/train-{000..715}.tar 37930 audio samples Valid valid/valid-{000..108}.tar 5741 audio samples Test test/test-{000..142}.tar 7569… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/Clotho-Moment.audioaudio-text-to-text10K<n<100K2 likes3.6k downloads8mo agoHugging Face18sayakpaul /pickapic_v2_webdatasetwebdataset archive of yuvalkirstain/pickapic_v2. Dataloading code can be found here. image1K<n<10K2 likes3.5k downloads2y agoHugging Face19Intelligent-Systems /BEDLAM2-depthgated Dataset Mirror of BEDLAM2.0 Dataset (Depth Data Subset) Project site: https://bedlam2.is.tuebingen.mpg.de/ Please register at project site for additional information and data in its Download section. Related Hugging Face dataset mirror: BEDLAM2 Dataset Information Depth maps (Multilayer EXR, 16-bit, available for 44% of images, 15TB) Multilayer EXR details 16-bit float depth in red channel (FinalImageMovieRenderQueue_WorldDepth.R) Color image without motion blur Body… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM2-depth.text1M<n<10M0 likes3.2k downloads7mo agoHugging Face20clip-benchmark /wds_fer2013image10K<n<100K0 likes3.1k downloads4y agoHugging Face21deepghs /danbooru2024-webp-4Mpixelgated 🎨 Danbooru2024 Webp 4MPixel Dataset 📊 Dataset Overview The Danbooru2024-Webp dataset is a comprehensive collection focused on animation and illustration artwork, derived from the official Danbooru platform. It contains approximately 8.05 million high-quality, user-annotated images with corresponding tags and textual descriptions. This dataset is 4MP-focused webp resized-dataset of Danbooru2024. ✨ Features 📋 Metadata Support Includes a Parquet… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/danbooru2024-webp-4Mpixel.textimage-classification100M<n<1B27 likes2.9k downloads2y agoHugging Face22mitermix /audiosnippets_small_with_detailed_annotation2audio1M<n<10M1 likes2.7k downloads2y agoHugging Face23gaunernst /voxceleb2-dev-wds VoxCeleb2 - dev set This is a copy of VoxCeleb2 dev set in WebDataset format. The audio data is the original AAC-encoded files without any transcoding. Refer to https://arxiv.org/abs/1806.05622 for more details about the dataset. There are 1,092,009 samples covering 5,994 unique speakers. The dataset is split into 779 shards of ~100MB. Usage import torchaudio import webdataset as wds from datasets import load_dataset def decode_audio(sample): audio, fs =… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/voxceleb2-dev-wds.textaudio-classification1M<n<10M2 likes2k downloads2y agoHugging Face24allenai /Molmo2-ER-RoboPoint Molmo2-ER · wentao-yuan/robopoint-data 1.43M robotics affordance instruction-tuning examples (pointing + detection + VQA). This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified. Upstream source Original dataset: wentao-yuan/robopoint-data Paper: RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics (arXiv:2406.10721) License:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-RoboPoint.image1M<n<10M1 likes1.9k downloads5mo agoHugging Face25sarulab-speech /commonvoice22_sidongated CV22-Sidon Overview This dataset hosts a release of Mozilla Common Voice 22 restored with the Sidon speech restoration model. Source: Mozilla Common Voice 22.0 Processing: Sidon denoising (sarulab-speech/sidon-v0.1) with 21 s chunks and 48 kHz reconstruction Format: WebDataset shards (.tar.gz) Manifest: paths.yaml enumerates every shard path for Hugging Face–style loading License: Original Common Voice license (CC0 1.0) Languages 137 language folders are… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon.audiotext-to-speech10M<n<100M30 likes1.8k downloads1y agoHugging Face26blowing-up-groundhogs /font-square-pretrain-20M 📚 Citation If you use this dataset in your research, please cite these papers: @article{pippi2023evaluating, title={Evaluating Synthetic Pre-Training for Handwriting Processing Tasks}, author={Pippi, Vittorio and Cascianelli, Silvia and Baraldi, Lorenzo and Cucchiara, Rita}, journal={Pattern Recognition Letters}, year={2023}, publisher={Elsevier} } @InProceedings{pippi2025zeroshot, author = {Pippi, Vittorio and Quattrini, Fabio and Cascianelli, Silvia and Tonioni… See the full description on the dataset page: https://huggingface.co/datasets/blowing-up-groundhogs/font-square-pretrain-20M.image10M<n<100M0 likes1.8k downloads6mo agoHugging Face27Yossh /danbooru2023-webp-4Mpixel-224The data set is just resized to 224*224 https://huggingface.co/datasets/KBlueLeaf/danbooru2023-webp-4Mpixel Pseudo code for processing def resize_image(file_path): with Image.open(file_path) as img: resized_img = img.resize((224, 224)) resized_img.save(file_path) image100K<n<1M2 likes1.7k downloads2y agoHugging Face28kensho /PubTables-v2 PubTables-v2 PubTables-v2 is a new large-scale dataset for full-page and multi-page table extraction. Official dataset evaluation scripts and leaderboard coming soon! In the meantime, you can create your own evaluation using GriTS with our open-source package, pip install grits-metric. Report any issues here: https://github.com/kensho-technologies/grits. See also: Hugging Face Paper Page News 2026 Apr 15: Code for the GriTS metric released… See the full description on the dataset page: https://huggingface.co/datasets/kensho/PubTables-v2.imageimage-to-text1M<n<10M25 likes1.6k downloads5mo agoHugging Face29krishnakalyan3 /emo_webds_2audio10K<n<100K7 likes1.6k downloads2y agoHugging Face30mitermix /audiosnippets_long_2_5Maudio1M<n<10M3 likes1.4k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.