CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ryanmarten /OpenThoughts-1k-sample [!NOTE] We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles! Inspect the content with rich formatting with Curator Viewer. Available Subsets default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models: ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.text1K<n<10K60 likes1.3m downloads1y agoHugging Face02google-research-datasets /mbpp Dataset Card for Mostly Basic Python Problems (mbpp) Dataset Summary The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us. Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.text1K<n<10K259 likes484k downloads3y agoHugging Face03Rowan /hellaswag Dataset Card for "hellaswag" Dataset Summary HellaSwag: Can a Machine Really Finish Your Sentence? is a new dataset for commonsense NLI. A paper was published at ACL2019. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of downloaded dataset files: 71.49 MB Size of the generated dataset: 65.32 MB Total… See the full description on the dataset page: https://huggingface.co/datasets/Rowan/hellaswag.text10K<n<100K199 likes443k downloads1y agoHugging Face04james-ra-henry /Rosetta-Activations Rosetta Activations Updated: 2026-06-15 02:30 UTC Contrastive activation extractions for 17 semantic concepts across 46 language models, supporting cross-architecture mechanistic interpretability research. Companion concept pair corpus: jamesrahenry/Rosetta_Concept_Pairs Papers: forthcoming Dataset Structure Rosetta-Activations/ ├── rcp_v1/ # Current extraction line — richest data (N≈2000) │ └── {Model_Name}/ │ ├── calibration_{concept}.npy… See the full description on the dataset page: https://huggingface.co/datasets/james-ra-henry/Rosetta-Activations.tabularn<1K0 likes393k downloads1mo agoHugging Face05mteb /resultstext1M<n<10M18 likes314k downloads15h agoHugging Face06rajpurkar /squad Dataset Card for SQuAD Dataset Summary Stanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be unanswerable. SQuAD 1.1 contains 100,000+ question-answer pairs on 500+ articles. Supported Tasks and Leaderboards Question… See the full description on the dataset page: https://huggingface.co/datasets/rajpurkar/squad.textquestion-answering10K<n<100K1k likes289k downloads3y agoHugging Face07google-research-datasets /paws Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling Dataset Summary PAWS: Paraphrase Adversaries from Word Scrambling This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset. For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.texttext-classification100K<n<1M40 likes235k downloads3y agoHugging Face08RekaAI /RekaDaily-10k-raw RekaDaily-10k (raw) Raw, unscripted, first-person daily-life video, collected through Claru, Reka's data collection marketplace — recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Videos are delivered as recorded — no cuts, no trimming, no editing, no filtering beyond basic integrity checks. A processed tier (short clips with machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.imagevideo-classification100K<n<1M22 likes219k downloads13d agoHugging Face09ehovy /race Dataset Card for "race" Dataset Summary RACE is a large-scale reading comprehension dataset with more than 28,000 passages and nearly 100,000 questions. The dataset is collected from English examinations in China, which are designed for middle school and high school students. The dataset can be served as the training and test sets for machine comprehension. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed… See the full description on the dataset page: https://huggingface.co/datasets/ehovy/race.textmultiple-choice100K<n<1M74 likes172k downloads3y agoHugging Face10open-r1 /OpenR1-Math-220k OpenR1-Math-220k Dataset description OpenR1-Math-220k is a large-scale dataset for mathematical reasoning. It consists of 220k math problems with two to four reasoning traces generated by DeepSeek R1 for problems from NuminaMath 1.5. The traces were verified using Math Verify for most samples and Llama-3.3-70B-Instruct as a judge for 12% of the samples, and each problem contains at least one reasoning trace with a correct answer. The dataset consists of two… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/OpenR1-Math-220k.text100K<n<1M802 likes162k downloads2y agoHugging Face11MuteApo /RealCam-Vid RealCam-Vid Dataset News 25/04/08: We provide torch dataset demo code for example usage of our RealCam-Vid. 25/03/26: Release our dataset RealCam-Vid v1 for metric-scale camera-controlled video generation, containing ~100K video clips with dedicated short/long captions and metric-scale camera annotations. 25/02/18: Initial commit of the project, we plan to release the full dataset and data processing code in several… See the full description on the dataset page: https://huggingface.co/datasets/MuteApo/RealCam-Vid.tabularimage-to-video100K<n<1M8 likes121k downloads1y agoHugging Face12jeggers /riddle_sensetext1K<n<10K1 likes120k downloads2y agoHugging Face13feyninc /recipes 🦛 Chonkie Recipes 🍳 Chonkie loves to cook up a storm in the kitchen This repository contains all the recipes that you can use with Chonkie to manage various documents, languages, and more. Usage To use the recipes, you need to install chonkie with the hub feature, with the following command: pip install "chonkie[hub]" This would enable Hubie which is used internally to get the recipes from this repository. So, you can do things like use the from_recipe… See the full description on the dataset page: https://huggingface.co/datasets/feyninc/recipes.textn<1K11 likes114k downloads1y agoHugging Face14nebius /SWE-rebench-V2 SWE-rebench-V2 Dataset Summary SWE-rebench-V2 is a curated dataset of software-engineering tasks derived from real GitHub issues and pull requests. The dataset contains 32,079 samples covering Python, Go, TypeScript, JavaScript, Rust, Java, PHP, Kotlin, Julia, Elixir, Scala, Swift, Dart, C, C++, C#, R, Clojure, OCaml, and Lua. For log parser functions, base Dockerfiles, and the prompts used, please see https://github.com/SWE-rebench/SWE-rebench-V2The detailed technical… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-rebench-V2.texttext-generation10K<n<100K60 likes109k downloads5mo agoHugging Face15rajpurkar /squad_v2 Dataset Card for SQuAD 2.0 Dataset Summary Stanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be unanswerable. SQuAD 2.0 combines the 100,000 questions in SQuAD1.1 with over 50,000 unanswerable questions written adversarially by crowdworkers… See the full description on the dataset page: https://huggingface.co/datasets/rajpurkar/squad_v2.textquestion-answering100K<n<1M263 likes95k downloads3y agoHugging Face16roneneldan /TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/roneneldan/TinyStories.texttext-generation1M<n<10M1.2k likes94k downloads2y agoHugging Face17nebius /SWE-rebench Dataset Summary SWE-rebench is a large-scale dataset designed to support training and evaluation of LLM-based software engineering (SWE) agents, building upon and expanding our earlier release, SWE-bench-extra. It is constructed using a fully automated pipeline that continuously extracts real-world interactive SWE tasks from GitHub repositories at scale, as detailed in our paper SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-rebench.textother10K<n<100K74 likes91k downloads9mo agoHugging Face18tiiuae /falcon-refinedweb 📀 Falcon RefinedWeb Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license. See the 📓 paper on arXiv for more details. RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data. RefinedWeb is also "multimodal-friendly": it contains links and… See the full description on the dataset page: https://huggingface.co/datasets/tiiuae/falcon-refinedweb.texttext-generation100M<n<1B965 likes89k downloads3y agoHugging Face19rtrm /debugtest3 textn<1K10 likes84k downloads3y agoHugging Face20open-llm-leaderboard-old /requests Open LLM Leaderboard Requests This repository contains the request files of models that have been submitted to the Open LLM Leaderboard. You can take a look at the current status of your model by finding its request file in this dataset. If your model failed, feel free to open an issue on the Open LLM Leaderboard! (We don't follow issues in this repository as often) Evaluation Methodology The evaluation process involves running your models against several benchmarks from… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/requests.textn<1K22 likes80k downloads2y agoHugging Face21google-research-datasets /natural_questions Dataset Card for Natural Questions Dataset Summary The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.textquestion-answering10K<n<100K127 likes79k downloads3y agoHugging Face22RevolutionCrossroads /loc_chronicling_america_1770-1810 Dataset Card for Chronicling America: Historic American Newspapers 1770–1810 Dataset Summary A dataset drawn from the Library of Congress Chronicling America digital collection, part of the National Digital Newspaper Program (NDNP). This dataset includes page-level records with images, original Chronicling America OCR, AI-generated OCR, and publication metadata for newspapers published between 1770 and 1810. It provides a foundation for research, machine learning… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810.documentimage-to-text10K<n<100K5 likes75k downloads3mo agoHugging Face23R2E-Gym /R2E-Gym-Litetabular10K<n<100K1 likes68k downloads2y agoHugging Face24racineai /VDR_MEGA_MultiDomain_DocRetrieval Visual Document Retrieval Dataset Overview This dataset is designed for training visual document retrieval models. It combines multiple datasets from the VDR series, Colpali, and LlamaIndex to create the most comprehensive training resource for visual document retrieval tasks. Dataset Structure The dataset contains structured fields including unique identifiers with string lengths ranging from 45 to 50 characters, search query text with variable lengths between… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval.imagevisual-document-retrieval1M<n<10M24 likes68k downloads6mo agoHugging Face25Fsoft-AIC /RobotDesign1M RobotDesign1M: A Large-scale Dataset for Robot Design Understanding RobotDesign1M is a large-scale, multimodal dataset for robot design understanding, built from image–text data curated from scientific literature across a wide range of robotics domains. It is designed to support research on design-aware foundation models, including design image generation, visual question answering about designs, and design image retrieval. 📄 Paper: RobotDesign1M: A Large-scale Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Fsoft-AIC/RobotDesign1M.imageimage-text-to-text1M<n<10M8 likes67k downloads2mo agoHugging Face26HuggingFaceH4 /no_robots Dataset Card for No Robots 🙅‍♂️🤖 Look Ma, an instruction dataset that wasn't generated by GPTs! Dataset Summary No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described in OpenAI's InstructGPT paper, and is comprised mostly of single-turn… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/no_robots.texttext-generation10K<n<100K581 likes63k downloads2y agoHugging Face27cornell-movie-review-data /rotten_tomatoes Dataset Card for "rotten_tomatoes" Dataset Summary Movie Review Dataset. This is a dataset of containing 5,331 positive and 5,331 negative processed sentences from Rotten Tomatoes movie reviews. This data was first used in Bo Pang and Lillian Lee, ``Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.'', Proceedings of the ACL, 2005. Supported Tasks and Leaderboards More Information Needed Languages… See the full description on the dataset page: https://huggingface.co/datasets/cornell-movie-review-data/rotten_tomatoes.texttext-classification10K<n<100K117 likes56k downloads3y agoHugging Face28SWE-Gym /SWE-Gym-RawSWE-Gym Raw contains 64,689 instances sourced from 358 Python repos. Most of the instances there doesn't have associated python environment configured and is not validated with SWE-Bench verification process. If you are working to scale training environments, these instances might be helpful. Otherwise, please take a look at SWE-Gym and SWE-Gym Lite , why are ready to be used for agent training. Get started at project page github.com/SWE-Gym/SWE-Gym Repository Frequency… See the full description on the dataset page: https://huggingface.co/datasets/SWE-Gym/SWE-Gym-Raw.text10K<n<100K1 likes54k downloads2y agoHugging Face29tasksource /reclorhttps://whyu.me/reclor/ @inproceedings{yu2020reclor, author = {Yu, Weihao and Jiang, Zihang and Dong, Yanfei and Feng, Jiashi}, title = {ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning}, booktitle = {International Conference on Learning Representations (ICLR)}, month = {April}, year = {2020} } text1K<n<10K18 likes53k downloads3y agoHugging Face30IFM /Code-Reasoning Code-Reasoning Dataset Description Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2 Web and… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Code-Reasoning.texttext-generation100M<n<1B83 likes52k downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.