CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google-research-datasets /mbpp Dataset Card for Mostly Basic Python Problems (mbpp) Dataset Summary The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us. Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.text1K<n<10K258 likes488k downloads3y agoHugging Face02jat-project /jat-dataset-tokenized Dataset Card for "jat-dataset-tokenized" More Information needed timeseries10M<n<100M32 likes472k downloads3y agoHugging Face03jat-project /jat-dataset JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.imagereinforcement-learning100M<n<1B78 likes410k downloads3y agoHugging Face04applied-ai-018 /peacock-data-public-datasets-idc0 likes376k downloads2y agoHugging Face05google-research-datasets /paws Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling Dataset Summary PAWS: Paraphrase Adversaries from Word Scrambling This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset. For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.texttext-classification100K<n<1M40 likes220k downloads3y agoHugging Face06GokuScraper /seedance-2-prompts-datasets 🎞️ Seedance-2-prompts-datasets 🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators. This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset. Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.imagetext-to-video1K<n<10K45 likes183k downloads25d agoHugging Face07community-datasets /quarel Dataset Card for "quarel" Dataset Summary QuaRel is a crowdsourced dataset of 2771 multiple-choice story questions, including their logical forms. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of downloaded dataset files: 0.63 MB Size of the generated dataset: 1.53 MB Total amount of disk used: 2.17 MB An example of 'train'… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/quarel.text1K<n<10K2 likes140k downloads2y agoHugging Face08jzr99 /mesh4d_dataset3d100K<n<1M0 likes124k downloads11mo agoHugging Face09hf-internal-testing /dataset_with_scriptThis is a test dataset.textn<1K0 likes119k downloads2y agoHugging Face10google-research-datasets /natural_questions Dataset Card for Natural Questions Dataset Summary The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.textquestion-answering10K<n<100K127 likes80k downloads3y agoHugging Face11klieret /swe-bench-dummy-test-datasettextn<1K0 likes75k downloads1y agoHugging Face12Goku-OpenLab /gpt-image-2-prompts-datasets 🖼️ GPT Image 2 Prompt Dataset 🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset. Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.imagetext-to-image10K<n<100K5 likes69k downloads25d agoHugging Face13albertvillanova /datasets-tests-compressiontextn<1K0 likes64k downloads5y agoHugging Face14legacy-datasets /wikipediaWikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.).text-generationn<1K675 likes59k downloads3y agoHugging Face15hf-internal-testing /multi_dir_datasettextn<1K0 likes55k downloads5y agoHugging Face16ai-for-good-lab /ai4g-flood-dataset Flood Detection Dataset Introduction This dataset accompanies the paper Mapping global floods with 10 years of satellite radar data (Nature Communications, 2025) and contains global flood detections derived from Sentinel-1 Synthetic Aperture Radar (SAR) imagery using a deep learning change detection model. The dataset spans October 2014 – September 2024, offering a longitudinal view of flood-prone areas worldwide. Key features: Cloud-penetrating SAR data for consistent… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/ai4g-flood-dataset.imagen<1K16 likes52k downloads11mo agoHugging Face17hf-internal-testing /dataset_with_data_filestextn<1K0 likes51k downloads2y agoHugging Face18amphion /Emilia-Datasetgated Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.audiotext-to-speech10M<n<100M489 likes48k downloads2y agoHugging Face19tensorshield /reddit_dataset_157 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/tensorshield/reddit_dataset_157.texttext-classification10M<n<100M3 likes47k downloads1y agoHugging Face20zhaocharile66 /My_Hermes_Dataset8 likes46k downloads2m agoHugging Face21llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes43k downloads2y agoHugging Face22leggedrobotics /grand_tour_dataset The GrandTour Dataset A project brought to you by RSL - ETH Zurich. References • Contributing • Citation References Official dataset webpage: grand-tour.leggedrobotics.com Getting started & examples: github.com/leggedrobotics/grand_tour_dataset Boxi used to collect the data: github.com/leggedrobotics/grand_tour_box Contributing We warmly welcome contributions to improve and expand this project. Whether it's new examples, enhancements, or… See the full description on the dataset page: https://huggingface.co/datasets/leggedrobotics/grand_tour_dataset.robotics1M<n<10M19 likes43k downloads8mo agoHugging Face23zahid0 /dataset11 likes43k downloads2mo agoHugging Face24JZSG /synth_dataset0 likes43k downloads11mo agoHugging Face25autogluon /fev_datasets Forecast evaluation datasets This repository contains time series datasets that can be used for evaluation of univariate & multivariate forecasting models. The main focus of this repository is on datasets that reflect real-world forecasting scenarios, such as those involving covariates, missing values, and other practical complexities. The datasets follow a format that is compatible with the fev package. Data format and usage Each dataset satisfies the following… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/fev_datasets.tabulartime-series-forecasting100K<n<1M13 likes42k downloads8mo agoHugging Face26serialexperimentsleon /fish_datasets_real_electrodyn_expertsys_twodim_fourier0 likes42k downloads1y agoHugging Face27yifengzhu-hf /LIBERO-datasets LIBERO Datasets This is a repo that stores the LIBERO datasets. The structure of the dataset can be found below: libero_object/ libero_spatial/ libero_goal/ libero_90/ libero_10/ Demonstrations of each task is stored in a hdf5 file. Please refer to download script from the official LIBERO repo for more details. 65 likes41k downloads1y agoHugging Face28lmarena-ai /leaderboard-dataset Arena Leaderboard Dataset Historical snapshots of the Arena leaderboard. Usage from datasets import load_dataset # Load all historical text style control data ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full") # Load the current text style control leaderboard ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest") # Filter to overall category ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.tabular1M<n<10M25 likes41k downloads6d agoHugging Face29trl-lib /trackio-dataset23 likes39k downloads1m agoHugging Face30ArlingtonCL2 /DogSpeak_Dataset DogSpeak: A Canine Vocalization Classification Dataset Paper: https://dl.acm.org/doi/10.1145/3746027.3758298 Dataset: https://huggingface.co/datasets/ArlingtonCL2/DogSpeak_Dataset Dataset Summary DogSpeak is a large-scale canine vocalization dataset designed to advance research in animal communication and computational bioacoustics. Unlike previous canine vocalization datasets recorded in controlled environments, DogSpeak is sourced from a large collection of… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/DogSpeak_Dataset.10K<n<100K8 likes39k downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.