CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huggingface /documentation-images This dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng.com/. imagen<1K205 likes2m downloads18h agoHugging Face02world-igr-plum /regions37 likes1.4m downloads1y agoHugging Face03BuLei /imgbedimagen<1K0 likes447k downloads9h agoHugging Face04inclusionAI /OpenAoE-2000h Open-AoE — Egocentric Hand Manipulation Dataset Release Roadmap Tier Duration Status nano ~3 h ✅ Released tiny ~100 h ✅ Released full 2000 h 🚧 Uploading Release notes 2026-07-30: Removed samples flagged in PR #1 for camera-intrinsics vs. video-resolution mismatches. 2026-07-31: Uploaded ~323h of data. 2026-08-12: Uploaded ~694h of data. 2026-09-03: Uploaded ~189h of data. Additional data for the full ~2000h release is still… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/OpenAoE-2000h.37 likes386k downloads20d agoHugging Face05google /IFEval Dataset Card for IFEval Dataset Summary This dataset contains the prompts used in the Instruction-Following Eval (IFEval) benchmark for large language models. It contains around 500 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times" which can be verified by heuristics. To load the dataset, run: from datasets import load_dataset ifeval = load_dataset("google/IFEval") Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/google/IFEval.texttext-generationn<1K167 likes361k downloads2y agoHugging Face06agents-course /course-imagesimagen<1K21 likes328k downloads1y agoHugging Face07huggingface-course /documentation-imagesimagen<1K3 likes314k downloads1y agoHugging Face08hf-internal-testing /transformers_circleci_workflow_runs5 likes300k downloads2y agoHugging Face09ieasybooks-org /prophet-mosque-library Prophet's Mosque Library 📖 Overview Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library.textimage-to-text10K<n<100K6 likes297k downloads1y agoHugging Face10stal-ix /pkgsrchttps://github.com/stal-ix/stal-ix.github.io/blob/main/MIRROR.md 2 likes294k downloads3d agoHugging Face11applied-ai-018 /peacock-data-public-datasets-idc0 likes293k downloads2y agoHugging Face12IPEC-COMMUNITY /FastUMI_100k_lerobot FastUMI-100K: Advancing Data-Driven Robotic Manipulation with a Large-Scale UMI-Style Dataset [paper] [dataset] ## Overview FastUMI-100K is a large-scale, high-quality UMI-style dataset designed for data-driven robotic manipulation learning. Featuring over **100K+ demonstration trajectories** across **54 diverse tasks** and hundreds of object types, the dataset provides multi-view wrist-mounted fisheye images and high-frequency end-effector states. To… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/FastUMI_100k_lerobot.8 likes282k downloads5mo agoHugging Face13IPEC-COMMUNITY /droid_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "franka", "total_episodes": 92233, "total_frames": 27044326, "total_tasks": 31308, "total_videos": 276699, "total_chunks": 93, "chunks_size": 1000, "fps": 15, "splits": { "train": "0:92233" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/droid_lerobot.robotics53 likes258k downloads1y agoHugging Face14fujinchu /imgbedaudion<1K1 likes251k downloads10m agoHugging Face15IPEC-COMMUNITY /language_table_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "xarm", "total_episodes": 442226, "total_frames": 7045476, "total_tasks": 127605, "total_videos": 442226, "total_chunks": 443, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:442226" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/language_table_lerobot.robotics29 likes236k downloads2y agoHugging Face16IPEC-COMMUNITY /kuka_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "kuka_iiwa", "total_episodes": 209880, "total_frames": 2455879, "total_tasks": 1, "total_videos": 209880, "total_chunks": 210, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:209880" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/kuka_lerobot.videorobotics7 likes215k downloads2y agoHugging Face17huggingface /DEH-image-scan-datan<1K22 likes210k downloads9h agoHugging Face18stanfordnlp /imdb Dataset Card for "imdb" Dataset Summary Large Movie Review Dataset. This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/imdb.texttext-classification100K<n<1M1.1k likes209k downloads3y agoHugging Face19wyu1 /Leopard-Instruct Leopard-Instruct Paper | Github | Models-LLaVA | Models-Idefics2 Summaries Leopard-Instruct is a large instruction-tuning dataset, comprising 925K instances, with 739K specifically designed for text-rich, multiimage scenarios. It's been used to train Leopard-LLaVA [checkpoint] and Leopard-Idefics2 [checkpoint]. Loading dataset to load the dataset without automatically downloading and process the images (Please run the following codes with datasets==2.18.0)… See the full description on the dataset page: https://huggingface.co/datasets/wyu1/Leopard-Instruct.image1M<n<10M64 likes194k downloads2y agoHugging Face20AiEDA /iDATA Dataset A dataset of AI + EDA iDATA is a dataset of AI + EDA, which can be used to train AI models for design PPA prediction, PPA-aware physical design, and related tasks. Dataset structure Describe the dataset structure. aes/ ├── iEDA_route_process_data/ # Process data exported by iEDA-iRT 2D routing ├── syn_netlist/ # The synthesized netlist files、sdc files ├── place/ # The place stage def、sdc、vectors └── route/ # The route… See the full description on the dataset page: https://huggingface.co/datasets/AiEDA/iDATA.8 likes184k downloads10mo agoHugging Face21Aak975 /iclr-wm-backup-public ICLR Watermark Benchmark — backup overflow (public part) Companion to the private repo Aak975/iclr-wm-backup, which reached its storage quota. Together the two repos form ONE backup — every file exists in exactly one of them, with the same layout: archives/<sub>/part-0000 ... part-NNNN, MANIFEST.json restore one archive: cat part-* | zstd -d | tar -x MANIFEST.json = {"parts": N, "sha256": <whole-stream>, "total_bytes": M} This public part holds only shareable image data… See the full description on the dataset page: https://huggingface.co/datasets/Aak975/iclr-wm-backup-public.tabularn<1K0 likes179k downloads22d agoHugging Face22IPEC-COMMUNITY /bridge_orig_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "widowx", "total_episodes": 53192, "total_frames": 1893026, "total_tasks": 19974, "total_videos": 212768, "total_chunks": 54, "chunks_size": 1000, "fps": 5, "splits": { "train": "0:53192" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/bridge_orig_lerobot.videorobotics27 likes168k downloads2y agoHugging Face23nasa-impact /WxC-Bench Dataset Card for WxC-Bench WxC-Bench primary goal is to provide a standardized benchmark for evaluating the performance of AI models in Atmospheric and Earth Sciences across various tasks. Dataset Details WxC-Bench contains datasets for six key tasks: Nonlocal Parameterization of Gravity Wave Momentum Flux Prediction of Aviation Turbulence Identifying Weather Analogs Generation of Natural Language Weather Forecasts Long-Term Precipitation Forecasting Hurricane Track and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/WxC-Bench.3 likes167k downloads8mo agoHugging Face24boltzgen /inference-data0 likes166k downloads1y agoHugging Face25Narsil /image_dummy\audion<1K0 likes146k downloads5y agoHugging Face26IPEC-COMMUNITY /fractal20220817_data_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "google_robot", "total_episodes": 87212, "total_frames": 3786400, "total_tasks": 599, "total_videos": 87212, "total_chunks": 88, "chunks_size": 1000, "fps": 3, "splits": { "train": "0:87212" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/fractal20220817_data_lerobot.videorobotics13 likes139k downloads2y agoHugging Face27ieasybooks-org /waqfeya-library Waqfeya Library 📖 Overview Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.imageimage-to-text10K<n<100K12 likes136k downloads1y agoHugging Face28hf-internal-testing /dataset_with_scriptThis is a test dataset.textn<1K0 likes127k downloads2y agoHugging Face29ibrahimhamamci /CT-RATEgated The CT-RATE Team organizes the VLM3D Challenge VLM3D 2026 (2nd Edition) → Challenge Finals at MICCAI 2026 VLM3D 2025 (1st Edition) → Challenge Finals at MICCAI 2025 • Workshop at ICCV 2025 The CT-RATE Team is developing the MR-RATE Dataset A large-scale brain MRI dataset with paired radiology reports for training 3D vision-language models. GitHub   |   Dataset   |   Metadata Dashboard Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimhamamci/CT-RATE.image-to-text10K<n<100K319 likes127k downloads6mo agoHugging Face30Idavidrein /gpqagated Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.tabularquestion-answering1K<n<10K544 likes127k downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.