CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LanguageBind /Open-Sora-Plan-v1.1.0 Annotation We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973 Pexels Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.text100K<n<1M46 likes105k downloads2y agoHugging Face02SparkAudio /voxbox VoxBox This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion. Dataset Structure . ├── audios/ │ └── aishell-3/ # Audio files (organised by sub-corpus) │ └── ... └── metadata/ ├── aishell-3.jsonl ├── casia.jsonl ├── commonvoice_cn.jsonl ├── ... └── wenetspeech4tts.jsonl # JSONL metadata files Each JSONL file corresponds to a… See the full description on the dataset page: https://huggingface.co/datasets/SparkAudio/voxbox.audiotext-to-speech10M<n<100M76 likes45k downloads1y agoHugging Face03sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes32k downloads10mo agoHugging Face04nvidia /PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card Dataset Description PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding. Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.image100M<n<1B40 likes27k downloads4mo agoHugging Face05mvp-lab /Sekaitext1M<n<10M0 likes24k downloads10mo agoHugging Face06clip-benchmark /wds_imagenet_sketchimage10K<n<100K1 likes19k downloads4y agoHugging Face07Intelligent-Systems /BEDLAM-depthgated Dataset Mirror of BEDLAM Dataset (Depth Data Subset) Project site: https://bedlam.is.tuebingen.mpg.de/ Please register at project site for additional information and data (Download section) Related Hugging Face dataset mirror: BEDLAM Dataset Information Depth maps (EXR, 32-bit, 3.8TB) Camera ground truth information is not included but can be found in the BEDLAM dataset mirror Image/video data with motion blur is not included but can be found in the BEDLAM dataset… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM-depth.text1M<n<10M0 likes18k downloads7mo agoHugging Face08laion /soundscapesaudio10M<n<100M7 likes16k downloads1y agoHugging Face09adams-story /imagenet1k-256-wdsThis is imagenet1k in webdataset format. Images are stored as jpg files. Every image has been resized to a maximum side length of 256. That means that if an image in the original dataset was 1000 by 500, the new size will be 256 by 128. Images with a maximum side length of under 256 were not resized. The total size of all dataset files is 57.8 GB, there are 1,281,167 rows in the training split and 50,000 rows in the validation split. imageimage-classification100K<n<1M2 likes15k downloads1y agoHugging Face10pixelprose /pixelprose-shards PixelProse Sharding Tars arXiv | public-released version: pixelprose | JSON-only version: pixelprose-jsons summary Each tar file is approximately 500-600 MB, friendly for fast on-the-fly sampling, filtering, and loading in dataloaders. Each tar file contains triplets of images, text, and JSON files. The *.txt files contain the raw original captions, while the *.json files include all the relevant information. Due to Gemini-1.0 internal version changes during the… See the full description on the dataset page: https://huggingface.co/datasets/pixelprose/pixelprose-shards.image1M<n<10M2 likes12k downloads9mo agoHugging Face11Salesforce /3d_optical_flow_droid 3D Optical Flow DROID Dataset Processed DROID robotics dataset with optical flow and scene flow annotations. Dataset Structure Organized by lab, each trajectory in separate tar.gz archive: IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz CLVR/CLVR+2023-...tar.gz ... (15 labs, ~33K trajectories) Each trajectory contains: metadata.json - Trajectory metadata trajectory.h5 - Robot state and actions camera_left/, camera_right/ - Camera data rgb/ - RGB images depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.imagerobotics10M<n<100M0 likes12k downloads8mo agoHugging Face12EasonXiao-888 /SpatialEdit-500K SpatialEdit-500K SpatialEdit-500K is a synthetic training dataset for fine-grained image spatial editing. It is built for learning geometry-aware edits such as object moving, object rotation, and camera viewpoint change. The dataset was introduced in the paper SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing. It is generated with a controllable rendering pipeline to provide structured spatial transformations at scale. Project Resources GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/EasonXiao-888/SpatialEdit-500K.imageimage-to-image100K<n<1M14 likes11k downloads6mo agoHugging Face13Spawning /pd12m-fullThis dataset is the downloaded variant of Spawning/PD12M. More specifically, this dataset is compatible with webdataset. It was made public after obtaining permission from the original authors of the dataset. You can use the following to explore the dataset with webdataset: import webdataset as wds dataset_path = "pipe:curl -s -f -L https://huggingface.co/datasets/sayakpaul/pd12m-full/resolve/main/{00155..02480}.tar" dataset = ( wds.WebDataset(dataset_path… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/pd12m-full.image10M<n<100M21 likes10k downloads2y agoHugging Face14AILab-CVC /obelics_seed2_tokensPart of the OBELISC data set, including 32 Million samples, please refer to dataset.py to use this data text10M<n<100M1 likes9.4k downloads3y agoHugging Face15sarulab-speech /mls_sidon MLS-Sidon Overview This dataset is a cleansed version of Multilingual LibriSpeech (MLS) with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. The dataset is provided in WebDataset format for efficient large-scale training. Source: Multilingual LibriSpeech Languages: English, German, French, Spanish, Italian, Polish, Dutch, Portuguese Format: WebDataset (.tar shards) License: CC-BY-4.0 Dataset Structure Each sample in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/mls_sidon.audiotext-to-speech10M<n<100M11 likes8.9k downloads1y agoHugging Face16speechcolab /gigaspeech2gated Dataset Card for GigaSpeech 2 Dataset Description GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. Repository: https://github.com/SpeechColab/GigaSpeech2 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.audioautomatic-speech-recognition10M<n<100M71 likes6k downloads6mo agoHugging Face17sci-papers /scientific-paperstext10M<n<100M0 likes5.6k downloads1y agoHugging Face18BLIP3o /BLIP3o-Pretrain-Short-Caption BLIP3o Pretrain Short-Caption Dataset This collection contains 5 million images, each paired with a short (~20 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Short-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Short-Caption.image1M<n<10M10 likes4.7k downloads1y agoHugging Face19Smith42 /minty-astro-ph MINT-1T ArXiv Astro-ph An astronomy-focused subset of mlfoundations/MINT-1T-ArXiv, filtered to include only papers from the astro-ph arXiv category (including cross-listed papers). Overview Papers ~845k Total size ~804 GB Format WebDataset tar shards Shards 287 (astro-ph-00000.tar to astro-ph-00286.tar) Shard size ~3 GB each Source MINT-1T (Awadalla et al., 2024) Data Format Each tar shard contains paired files per paper:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/minty-astro-ph.imagetext-generation100K<n<1M1 likes4.6k downloads5mo agoHugging Face20sayakpaul /pickapic_v2_webdatasetwebdataset archive of yuvalkirstain/pickapic_v2. Dataloading code can be found here. image1K<n<10K2 likes3.5k downloads2y agoHugging Face21Intelligent-Systems /BEDLAM2-depthgated Dataset Mirror of BEDLAM2.0 Dataset (Depth Data Subset) Project site: https://bedlam2.is.tuebingen.mpg.de/ Please register at project site for additional information and data in its Download section. Related Hugging Face dataset mirror: BEDLAM2 Dataset Information Depth maps (Multilayer EXR, 16-bit, available for 44% of images, 15TB) Multilayer EXR details 16-bit float depth in red channel (FinalImageMovieRenderQueue_WorldDepth.R) Color image without motion blur Body… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM2-depth.text1M<n<10M0 likes3.2k downloads7mo agoHugging Face22mitermix /audiosnippets_small_with_detailed_annotationaudio100K<n<1M1 likes2.7k downloads2y agoHugging Face23ZhengGuangze /Stereo4D_vlbm Stereo4D (converted to VLBM format) This dataset contains 4,687 sequences from the Stereo4D dataset converted to the VLBM-compatible format using preprocess_stereo4d.py. The sequences have been compressed into .tar.gz archives in chunks of 50 sequences per archive. Scale Metric Value Total sequences 4,687 Image resolution 512 x 512 px Depth type Sparse (projected from tracked 3D points) Dataset Structure Each sequence directory follows this… See the full description on the dataset page: https://huggingface.co/datasets/ZhengGuangze/Stereo4D_vlbm.image1M<n<10M0 likes2.7k downloads6mo agoHugging Face24mitermix /audiosnippets_small_with_detailed_annotation2audio1M<n<10M1 likes2.7k downloads2y agoHugging Face25ScienceOne-AI /S1-MMAlignS1-MMAlign A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers. Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1-MMAlign aims to… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.imageimage-to-text10M<n<100M106 likes2.6k downloads6mo agoHugging Face26datastuff /scientific-stuff-1text10M<n<100M2 likes2.4k downloads1y agoHugging Face27laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M15 likes2.1k downloads11mo agoHugging Face28collabora /hi-stt-preprocessed-webdatasettext100K<n<1M1 likes2k downloads1y agoHugging Face29blowing-up-groundhogs /font-square-pretrain-20M 📚 Citation If you use this dataset in your research, please cite these papers: @article{pippi2023evaluating, title={Evaluating Synthetic Pre-Training for Handwriting Processing Tasks}, author={Pippi, Vittorio and Cascianelli, Silvia and Baraldi, Lorenzo and Cucchiara, Rita}, journal={Pattern Recognition Letters}, year={2023}, publisher={Elsevier} } @InProceedings{pippi2025zeroshot, author = {Pippi, Vittorio and Quattrini, Fabio and Cascianelli, Silvia and Tonioni… See the full description on the dataset page: https://huggingface.co/datasets/blowing-up-groundhogs/font-square-pretrain-20M.image10M<n<100M0 likes1.8k downloads6mo agoHugging Face30semi-truths /Semi-Truths Semi Truths Dataset: A Large-Scale Dataset for Testing Robustness of AI-Generated Image Detectors (NeurIPS 2024 Track Datasets & Benchmarks Track) Recent efforts have developed AI-generated image detectors claiming robustness against various augmentations, but their effectiveness remains unclear. Can these systems detect varying degrees of augmentation? To address these questions, we introduce Semi-Truths, featuring 27, 600 real images, 223, 400 masks, and 1, 472, 700… See the full description on the dataset page: https://huggingface.co/datasets/semi-truths/Semi-Truths.imageimage-classification1M<n<10M9 likes1.8k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.