CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yyyzzzzyyy /envss0 likes166k downloads1y agoHugging Face02ylacombe /cml-tts Dataset Card for CML-TTS Dataset Summary CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/cml-tts.audiotext-to-speech1M<n<10M36 likes119k downloads3y agoHugging Face03ybisk /piqaTo apply eyeshadow without a brush, should I use a cotton swab or a toothpick? Questions requiring this kind of physical commonsense pose a challenge to state-of-the-art natural language understanding systems. The PIQA dataset introduces the task of physical commonsense reasoning and a corresponding benchmark dataset Physical Interaction: Question Answering or PIQA. Physical commonsense knowledge is a major challenge on the road to true AI-completeness, including robots that interact with the world and understand natural language. PIQA focuses on everyday situations with a preference for atypical solutions. The dataset is inspired by instructables.com, which provides users with instructions on how to build, craft, bake, or manipulate objects using everyday materials. The underlying task is formualted as multiple choice question answering: given a question `q` and two possible solutions `s1`, `s2`, a model or a human must choose the most appropriate solution, of which exactly one is correct. The dataset is further cleaned of basic artifacts using the AFLite algorithm which is an improvement of adversarial filtering. The dataset contains 16,000 examples for training, 2,000 for development and 3,000 for testing.question-answering10K<n<100K107 likes118k downloads3y agoHugging Face04defeatbeta /yahoo-finance-data The Financial data from Yahoo! *** Key Points to Note *** All financial data is sourced from Yahoo!Ⓡ Finance, Nasdaq!Ⓡ, and the U.S. Department of the Treasury via publicly available APIs, and is intended for research and educational purposes. I will update the data regularly, and you are welcome to follow this project and use the data. Each time the data is updated, I will record the update time in spec.json. Data Usage Instructions Use DuckDB… See the full description on the dataset page: https://huggingface.co/datasets/defeatbeta/yahoo-finance-data.100M<n<1B127 likes112k downloads3d agoHugging Face05YWZBrandon /webshop-data4 likes95k downloads1y agoHugging Face06YipengGao /3DCode Project page Paper Code 3dcodebench.com arXiv:2606.01057 gaoypeng/3dcodebench News [06/01/2026] Paper released on arXiv: 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code. Note. This is an open-source reproduction of 3DCodeBench. ⚠️ Under final check. The 3DCodeData/ code is still undergoing final quality review and may contain occasional issues (non-executable scripts, mismatched captions/renders, or imperfect geometry). If you run… See the full description on the dataset page: https://huggingface.co/datasets/YipengGao/3DCode.3dtext-to-3d10K<n<100K25 likes88k downloads4d agoHugging Face07ylecun /mnist Dataset Card for MNIST Dataset Summary The MNIST dataset consists of 70,000 28x28 black-and-white images of handwritten digits extracted from two NIST databases. There are 60,000 images in the training dataset and 10,000 images in the validation dataset, one class per digit so a total of 10 classes, with 7,000 images (6,000 train images and 1,000 test images) per class. Half of the image were drawn by Census Bureau employees and the other half by high school students… See the full description on the dataset page: https://huggingface.co/datasets/ylecun/mnist.imageimage-classification10K<n<100K275 likes79k downloads2y agoHugging Face08espnet /yodas-granary Dataset Card for YODAS-Granary Repository: NeMo-speech-data-processor: Granary Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages Shared by: ESPnet Dataset Description YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages. Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.audioautomatic-speech-recognition10M<n<100M33 likes77k downloads1y agoHugging Face09yaak-ai /L2DTL;DR of L2D, the world's largest self-driving dataset! Read more about L2D on the official Huggingface blog: LeRobot goes to driving school 90+ TeraBytes of multimodal data (5000+ hours of driving) from 30 cities in Germany 6x surrounding HD cameras and complete vehicle state: Speed/Heading/GPS/IMU Continuous: Gas/Brake/Steering and discrete actions: Gear/Turn Signals Environment state: Lane count, Road type (highway|residential), Road surface (asphalt, cobbled, sett), Max speed limit.… See the full description on the dataset page: https://huggingface.co/datasets/yaak-ai/L2D.tabularrobotics10M<n<100M52 likes68k downloads4mo agoHugging Face10espnet /yodasUpdates 2024/07/09: we also uploaded a new version of YODAS as YODAS2, it provides unsegmented audios and higher sampling rate (24k) README This is the YODAS manual/automatic subset from our YODAS dataset, it has 369,510 hours of speech. This dataset contains audio utterances and corresponding captions (manual or automatic) from YouTube. Note that manual caption only indicates that it is uploaded by users, but not necessarily transcribed by a human For more details about YODAS… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas.155 likes54k downloads2y agoHugging Face11espnet /yodas2YODAS2 is the long-form dataset from YODAS dataset. It provides the same dataset as espnet/yodas but YODAS2 has the following new features: formatted in the long-form (video-level) where audios are not segmented. audios are encoded using higher sampling rates (i.e. 24k) For detailed information about YODAS dataset, please refer to our paper and the espnet/yodas repo. Usage: Each data point corresponds to an entire video on YouTube, it contains the following fields: video_id:… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas2.56 likes51k downloads1y agoHugging Face12yenle84622 /yenle846220 likes48k downloads13h agoHugging Face13YYF111 /checkpoint Dataset Card for LLaVA-Video-178K Uses This dataset is used for the training of the LLaVA-Video model. We only allow the use of this dataset for academic research and education purpose. For OpenAI GPT-4 generated data, we recommend the users to check the OpenAI Usage Policy. Data Sources For the training of LLaVA-Video, we utilized video-language data from five primary sources: LLaVA-Video-178K: This dataset includes 178,510 caption entries, 960,792 open-ended… See the full description on the dataset page: https://huggingface.co/datasets/YYF111/checkpoint.imagevisual-question-answering100K<n<1M0 likes43k downloads5mo agoHugging Face14Yauca47 /Dimeggvideon<1K0 likes41k downloads4mo agoHugging Face15yifengzhu-hf /LIBERO-datasets LIBERO Datasets This is a repo that stores the LIBERO datasets. The structure of the dataset can be found below: libero_object/ libero_spatial/ libero_goal/ libero_90/ libero_10/ Demonstrations of each task is stored in a hdf5 file. Please refer to download script from the official LIBERO repo for more details. 67 likes40k downloads1y agoHugging Face16yentinglin /aime_2025 AIME 2025 This dataset contains 30 problems from the 2025 AIME tests, including: AIME I: 15 problems AIME II: 15 problems tabularn<1K12 likes37k downloads9mo agoHugging Face17YiboZhang2001 /TexVerse TexVerse: A Universe of 3D Objects with High-Resolution Textures &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Yibo Zhang1,2, Li Zhang1,3, Rui Ma2 *, Nan Cao1,4 1Shanghai Innovation Institute 2Jilin University 3Fudan University 4Tongji University * Corresponding Author TexVerse is a large-scale 3D dataset featuring high-resolution textures. Its key characteristics include: Scale & Source: TexVerse dataset has 858,669 unique 3D models curated from… See the full description on the dataset page: https://huggingface.co/datasets/YiboZhang2001/TexVerse.50 likes35k downloads13d agoHugging Face18yzwwxm /oi-dev0 likes33k downloads5mo agoHugging Face19sarulab-speech /yodas2_sidon YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.audiotext-to-speech1M<n<10M65 likes32k downloads10mo agoHugging Face20yunfanlu /RGB-Event-ISP-Datasetimage1K<n<10K2 likes30k downloads5mo agoHugging Face21YiboZhang2001 /TexVerse-1K7 likes30k downloads3mo agoHugging Face22YWjimmy /PeRFception-v1-2image0 likes27k downloads4y agoHugging Face23Yqy6 /Slides-Align Slides-Align: Human Preference Rankings for AI-Generated Presentations Project Page | Paper | GitHub Overview Slides-Align is a human preference dataset for evaluating AI-generated slide presentations, introduced as part of the SlidesGen-Bench framework. It contains 1,326 human rankings comparing presentations generated by 9 different AI slide generation products across 7 scenario categories and 187 unique topics. This dataset enables: 🎯 Benchmarking AI slide… See the full description on the dataset page: https://huggingface.co/datasets/Yqy6/Slides-Align.imageothern<1K1 likes27k downloads8mo agoHugging Face24Yootta /World-SimReady-Homegated WorldSimReady-Home Dataset description CAD-based SimReady assets Optimized CAD assets with configured collision and physical properties. Manually reviewed scenes Physics configuration reviewed for every household scene. Scalable task generation Batch simulation data across robot embodiments and tasks. WorldSimReady-Home is built from CAD-based object assets, optimized and enriched with physical properties to create simulation-ready assets. The… See the full description on the dataset page: https://huggingface.co/datasets/Yootta/World-SimReady-Home.3drobotics82 likes27k downloads5d agoHugging Face25yahma /alpaca-cleaned Dataset Card for Alpaca-Cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer. "instruction":"Summarize… See the full description on the dataset page: https://huggingface.co/datasets/yahma/alpaca-cleaned.texttext-generation10K<n<100K888 likes25k downloads3y agoHugging Face26NTU-yiwen /code-world-model-project-page-videos Code World Model Project Page Videos Public research-demo video assets used by the Code World Model project page. The gallery/ directory contains aligned RGB and proxy videos for interactive comparison. videon<1K0 likes23k downloads29d agoHugging Face27yizhouz /MASIV MASIV Multi-Sequence Dataset Toward Material-Agnostic System Identification from Videos ICCV 2025 Yizhou Zhao1, Haoyu Chen1, Chunjiang Liu1, Zhenyang Li2, Charles Herrmann3, Junhwa Hur3, Yinxiao Li3, Ming‑Hsuan Yang4, Bhiksha Raj1, Min Xu1* 1Carnegie Mellon University 2University of Alabama at Birmingham 3Google 4UC Merced Introduction The MASIV Multi-Sequence Dataset is a synthetic dataset generated by Genesis to evaluate the generalization of data-driven… See the full description on the dataset page: https://huggingface.co/datasets/yizhouz/MASIV.n<1K1 likes21k downloads1y agoHugging Face28Yelp /yelp_review_full Dataset Card for YelpReviewFull Dataset Summary The Yelp reviews dataset consists of reviews from Yelp. It is extracted from the Yelp Dataset Challenge 2015 data. Supported Tasks and Leaderboards text-classification, sentiment-classification: The dataset is mainly used for text classification: given the text, predict the sentiment. Languages The reviews were mainly written in english. Dataset Structure Data Instances A… See the full description on the dataset page: https://huggingface.co/datasets/Yelp/yelp_review_full.texttext-classification100K<n<1M149 likes21k downloads3y agoHugging Face29yyyzzzzyyy /sd3_5_fine_sixcard DreamBooth training example DreamBooth is a method to personalize text2image models like stable diffusion given just a few(3~5) images of a subject. The train_dreambooth.py script shows how to implement the training procedure and adapt it for stable diffusion. Running locally with PyTorch Installing the dependencies Before running the scripts, make sure to install the library's training dependencies: Important To make sure you can successfully run the latest… See the full description on the dataset page: https://huggingface.co/datasets/yyyzzzzyyy/sd3_5_fine_sixcard.0 likes20k downloads1y agoHugging Face30datasets-maintainers /dataset-with-standalone-yamlThis is a test dataset used in the datasets library CI textn<1K0 likes20k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.