CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01builddotai /Egocentric-100Kgated Egocentric-100K is the largest dataset of manual labor. You can visualize the dataset here. Egocentric-100K is state-of-the-art in hand visibility and active manipulation density compared to previous in-the-wild egocentric datasets. The complete 30,000 frame evaluation set is available at Egocentric-100K-Evaluation. Dataset Statistics Attribute Value Total Hours 100,405 Total Frames 10.8 billion Video Clips 2,010,759 Median Clip Length 180.0 seconds Mean… See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-100K.text1M<n<10M142 likes125k downloads7mo agoHugging Face02agibot-world /AgiBotWorld-Betagated Key Features 🔑 1 million+ trajectories from 100 robots, with a total duration of 2976.4 hours. 100+ real-world scenarios across 5 target domains. Cutting-edge hardware: visual tactile sensors / 6-DoF dexterous hand / mobile dual-arm robots 200+ types of tasks: Contact-rich manipulation Long-horizon planning Multi-robot collaboration 87 types of Atomic Skills, including Tie, OpenJar, Peel, Sweep etc. Your… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/AgiBotWorld-Beta.textother100M<n<1B77 likes104k downloads11mo agoHugging Face03LanguageBind /Open-Sora-Plan-v1.1.0 Annotation We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973 Pexels Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.text100K<n<1M46 likes104k downloads2y agoHugging Face04InternRobotics /OmniWorld[ICLR 2026] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling         🎉NEWS [2026.3.21] 🔥 OmniWorld-Game with Metric Scale is now released! Check out our latest model Pi3X (an enhanced version of Pi3), which leverages this data to achieve better performance! [2026.1.26] 🎉 OmniWorld was accepted by ICLR 2026! [2026.1.7] Update OmniWorld-Game, release RH20T-Robot, RH20T-Human, Ego-Exo4D, EgoDex, Epic-Kitchens. [2025.11.11] The OmniWorld is… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/OmniWorld.imagetext-to-video1B<n<10B96 likes69k downloads5mo agoHugging Face05bop-benchmark /hot3d HOT3D-Clips This Hugging Face repository hosts HOT3D-Clips, a set of curated sub-sequences of the HOT3D dataset. Download instructions for HOT3D-Clips and the full HOT3D dataset can be found here. See HOT3D Toolkit for documentation of the data format and for Python utilities (for loading, undistorting fisheye images, rendering using fisheye cameras, etc.). More details can be found in the HOT3D paper and BOP 2024 report. image100K<n<1M8 likes60k downloads1y agoHugging Face06Vchitect /Vchitect_T2V_DataVerse Vchitect-T2V-Dataverse Vchitect Team1  1Shanghai Artificial Intelligence Laboratory  Paper | Project Page | Data Overview The Vchitect-T2V-Dataverse is the core dataset used to train our text-to-video diffusion model, Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models. It comprises 14 million high-quality videos collected from the Internet, each paired with detailed textual… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse.texttext-to-video1M<n<10M11 likes58k downloads1y agoHugging Face07clip-benchmark /wds_objectnetimage1K<n<10K4 likes51k downloads4y agoHugging Face08amphion /Emilia-Datasetgated Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.audiotext-to-speech10M<n<100M489 likes48k downloads2y agoHugging Face09SparkAudio /voxbox VoxBox This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion. Dataset Structure . ├── audios/ │ └── aishell-3/ # Audio files (organised by sub-corpus) │ └── ... └── metadata/ ├── aishell-3.jsonl ├── casia.jsonl ├── commonvoice_cn.jsonl ├── ... └── wenetspeech4tts.jsonl # JSONL metadata files Each JSONL file corresponds to a… See the full description on the dataset page: https://huggingface.co/datasets/SparkAudio/voxbox.audiotext-to-speech10M<n<100M76 likes45k downloads1y agoHugging Face10pixparse /cc12m-wds Dataset Card for Conceptual Captions 12M (CC12M) Dataset Summary Conceptual 12M (CC12M) is a dataset with 12 million image-text pairs specifically meant to be used for visionand-language pre-training. Its data collection pipeline is a relaxed version of the one used in Conceptual Captions 3M (CC3M). Usage This instance of Conceptual Captions is in webdataset .tar format. It can be used with webdataset library or upcoming releases of Hugging Face datasets.… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/cc12m-wds.imageimage-to-text10M<n<100M45 likes43k downloads3y agoHugging Face11wendlerc /RenderedTextThis dataset has been created by Stability AI and LAION. This dataset contains 12 million 1024x1024 images of handwritten text written on a digital 3D sheet of paper generated using Blender geometry nodes and rendered using Blender Cycles. The text has varying font size, color, and rotation, and the paper was rendered under random lighting conditions. Note that, the first 10 million examples are in the root folder of this dataset repository and the remaining 2 million are in ./remaining (due… See the full description on the dataset page: https://huggingface.co/datasets/wendlerc/RenderedText.imagetext-to-image10M<n<100M60 likes40k downloads11mo agoHugging Face12xwm /WildGUI WildGUI This repository hosts a personally reprocessed annotation release for WildGUI, the dataset introduced by Video2GUI. The original Video2GUI project builds WildGUI from large-scale Internet tutorial videos for GUI agent pretraining. This repository focuses on the open annotation artifacts: the records were regenerated and cleaned following the full annotation workflow, then reformatted to make the data easier to inspect, reuse, and reproduce. It also ships the screenshot… See the full description on the dataset page: https://huggingface.co/datasets/xwm/WildGUI.image10M<n<100M8 likes40k downloads3mo agoHugging Face13sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes33k downloads10mo agoHugging Face14agibot-world /AgiBotWorld-Alphagated ⚠️Important Notice !!! Dear Users, The Alpha Dataset has been updated as follows: Frame Loss Data Removal: Several episodes with frame loss issues have been removed. For the complete list of removed episode IDs, please refer to this document. Changes in Episode Count: The updated Alpha Dataset retains the original 36 tasks. The new version has been enriched with additional interactive objects, extending the total duration from 474.12… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/AgiBotWorld-Alpha.textrobotics10M<n<100M237 likes33k downloads1y agoHugging Face15nvidia /PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card Dataset Description PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding. Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.image100M<n<1B40 likes25k downloads3mo agoHugging Face16mvp-lab /Sekaitext1M<n<10M0 likes24k downloads10mo agoHugging Face17TencentARC /TimeLens-100K TimeLens-100K 📑 Paper | 💻 Code | 🏠 Project Page | 🤗 Model & Data ✨ Dataset Description TimeLens-100K is a large-scale, diverse, and high-quality training dataset for video temporal grounding. It was proposed in our paper TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs and used for training TimeLens models. The annotation process was conducted using an automated pipeline powered by Gemini-2.5-Pro. 📊 Dataset Statistics Total Videos:… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/TimeLens-100K.textvideo-text-to-text10K<n<100K7 likes21k downloads9mo agoHugging Face18ma-xu /fine-t2i Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning [arxiv] by Xu Ma, Yitian Zhang, Qihua Dong, Yun Fu Northeastern Univeristy Please see our [Dataset Explore] to view detailed samples (loading is slow, be patient). 🆕 What's New [2026.02.20]: Fine-T2I reaches the #1 spot among Hugging Face Datasets Trending list ⭐️⭐️⭐️ [2026.02.16]: Fine-T2I tops the Hugging Face Datasets Trending list, reaching the #2 spot and #1… See the full description on the dataset page: https://huggingface.co/datasets/ma-xu/fine-t2i.imageimage-to-text100K<n<1M120 likes21k downloads7mo agoHugging Face19mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face20Intelligent-Systems /BEDLAM-depthgated Dataset Mirror of BEDLAM Dataset (Depth Data Subset) Project site: https://bedlam.is.tuebingen.mpg.de/ Please register at project site for additional information and data (Download section) Related Hugging Face dataset mirror: BEDLAM Dataset Information Depth maps (EXR, 32-bit, 3.8TB) Camera ground truth information is not included but can be found in the BEDLAM dataset mirror Image/video data with motion blur is not included but can be found in the BEDLAM dataset… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM-depth.text1M<n<10M0 likes19k downloads7mo agoHugging Face21clip-benchmark /wds_imagenet_sketchimage10K<n<100K1 likes18k downloads4y agoHugging Face22laion /LAION-Audio-300Maudio100M<n<1B74 likes18k downloads2y agoHugging Face23mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes16k downloads2y agoHugging Face24laion /soundscapesaudio10M<n<100M7 likes16k downloads1y agoHugging Face25qualialabsAI /DuplexConv DuplexConv DuplexConv is a large-scale Chinese multi-channel conversational speech dataset with LLM-assisted annotations, developed by ASLP@NPU and QualiaLabs as part of the SmoothConv–DuplexConv corpus family. Companion dataset: SmoothConv on HuggingFace (100 hours, expert human annotation). DuplexConv and SmoothConv share the same conversational domains and a unified data design. SmoothConv focuses on high-quality human annotations for benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/qualialabsAI/DuplexConv.audio100K<n<1M13 likes16k downloads3mo agoHugging Face26adams-story /imagenet1k-256-wdsThis is imagenet1k in webdataset format. Images are stored as jpg files. Every image has been resized to a maximum side length of 256. That means that if an image in the original dataset was 1000 by 500, the new size will be 256 by 128. Images with a maximum side length of under 256 were not resized. The total size of all dataset files is 57.8 GB, there are 1,281,167 rows in the training split and 50,000 rows in the validation split. imageimage-classification100K<n<1M2 likes15k downloads1y agoHugging Face27pixparse /cc3m-wds Dataset Card for Conceptual Captions (CC3M) Dataset Summary Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/cc3m-wds.imageimage-to-text1M<n<10M58 likes15k downloads3y agoHugging Face28vaishaal /ImageNetV2image10K<n<100K9 likes13k downloads4y agoHugging Face29Salesforce /3d_optical_flow_droid 3D Optical Flow DROID Dataset Processed DROID robotics dataset with optical flow and scene flow annotations. Dataset Structure Organized by lab, each trajectory in separate tar.gz archive: IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz CLVR/CLVR+2023-...tar.gz ... (15 labs, ~33K trajectories) Each trajectory contains: metadata.json - Trajectory metadata trajectory.h5 - Robot state and actions camera_left/, camera_right/ - Camera data rgb/ - RGB images depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.imagerobotics10M<n<100M0 likes13k downloads8mo agoHugging Face30JoeLeelyf /OVO-Bench OVO-Bench: How Far is Your Video-LLMs from Real-World Online VideO Understanding? 🔥🔥OVO-Bench is accepted by CVPR 2025!🔥🔥 Important Note: Current codebase is modified compared to our initial arXiv paper. We strongly recommend that any use of OVO-Bench should be based on current edition. Introduction 🌟 Three distinct problem-solving modes Backward Tracing: trace back to past events to answer the question.Real-Time Visual Perception:… See the full description on the dataset page: https://huggingface.co/datasets/JoeLeelyf/OVO-Bench.textvideo-text-to-text1K<n<10K9 likes12k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.