CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01reasoning-machines /gsm-hard Dataset Summary This is the harder version of gsm8k math reasoning dataset (https://huggingface.co/datasets/gsm8k). We construct this dataset by replacing the numbers in the questions of GSM8K with larger numbers that are less common.  Supported Tasks and Leaderboards This dataset is used to evaluate math reasoning Languages English - Numbers Dataset Structure dataset = load_dataset("reasoning-machines/gsm-hard") DatasetDict({ train: Dataset({… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-machines/gsm-hard.text1K<n<10K66 likes32k downloads4y agoHugging Face02mlabonne /harmful_behaviorstextn<1K158 likes23k downloads2y agoHugging Face03mlabonne /harmless_alpacatext10K<n<100K47 likes22k downloads2y agoHugging Face04Joschka /big_bench_hardAll rights and obligations of the dataset are with original authors of the paper/dataset. I have merely made this dataset with a MIT licence available on HuggingFace. BIG-Bench Hard Dataset This repository contains a copy of the BIG-Bench Hard dataset. Small edits to the formatting of the dataset are made to integrate it into the Inspect Evals repository, a community contributed LLM evaulations for Inspect AI a framework by the UK AI Safety Institute. The BIG-Bench Hard dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Joschka/big_bench_hard.textquestion-answering1K<n<10K3 likes19k downloads1y agoHugging Face05trl-internal-testing /harmonytextn<1K0 likes10k downloads9mo agoHugging Face06walledai /HarmBenchgated HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal Data: Dataset About In this dataset card, we only use the behavior prompts proposed in HarmBench. License MIT Citation If you find HarmBench useful in your research, please consider citing the paper: @article{mazeika2024harmbench, title={HarmBench: A… See the full description on the dataset page: https://huggingface.co/datasets/walledai/HarmBench.textn<1K58 likes7.5k downloads2y agoHugging Face07AweAI-Team /BeyondSWE-harbor BeyondSWE-harbor This repository provides the harbor version of the BeyondSWE benchmark, containing the full task instances in a directory-based (harbor) format, where each instance is stored as an independent folder. 📌 For benchmark definition, data format and detailed evaluation results, please refer to: 🤗 Main Dataset on HuggingFace 🗂️ Data Structure beyondswe/ ├── {instance_id}/ │ ├── environment/ │ ├── solution/ │ ├── tests/ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/BeyondSWE-harbor.texttext-generationn<1K2 likes7.2k downloads6mo agoHugging Face08bigcode /bigcodebench-hardtabularn<1K3 likes6.6k downloads2y agoHugging Face09Infatoshi /kernelbench-hard-traces KernelBench-Hard agent traces Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and B200; roofline-graded. Each .jsonl file is one agent run in Claude-Code session format, viewable with the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename = run id. Live leaderboard: https://kernelbench.com/hard Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.tabulartext-generationn<1K16 likes5.3k downloads1d agoHugging Face10pat-jj /harness-1-train-data Harness-1 Training Data This dataset contains the training data used for Harness-1, plus the retrieval corpora needed to reproduce the training/evaluation environment. Contents The dataset has one train split with a stage column: sft: 899 raw GPT-5.4-generated v8d SFT trajectories produced by generate_sft_ultra_0417.py and used by train_sft_ultra_0417.py from sft_ultra_v8d_data. rl: 3453 SEC training-split query records used for RL (TRAIN_DATASETS=sec… See the full description on the dataset page: https://huggingface.co/datasets/pat-jj/harness-1-train-data.texttext-generation1K<n<10K0 likes4.5k downloads3mo agoHugging Face11dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4.1k downloads1y agoHugging Face12harvard-lil /cold-cases Collaborative Open Legal Data (COLD) - Cases COLD Cases is a dataset of 8.3 million United States legal decisions with text and metadata, formatted as compressed parquet files. If you'd like to view a sample of the dataset formatted as JSON Lines, you can view one here This dataset exists to support the open legal movement exemplified by projects like Pile of Law and LegalBench. A key input to legal understanding projects is caselaw -- the published, precedential decisions of… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-cases.tabular100K<n<1M42 likes4.1k downloads3y agoHugging Face13taesiri /imagenet_hard_review_data_r2tabular1K<n<10K0 likes4k downloads3y agoHugging Face14LLM-LAT /harmful-datasettext1K<n<10K43 likes3.3k downloads2y agoHugging Face15FabienRoger /alignment_faking_harm_answerstext1K<n<10K0 likes3.3k downloads1y agoHugging Face16lighteval /MATH-Hard Dataset Card for Mathematics Aptitude Test of Heuristics, hard subset (MATH-Hard) dataset Dataset Summary The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems from mathematics competitions, including the AMC 10, AMC 12, AIME, and more. Each problem in MATH has a full step-by-step solution, which can be used to teach models to generate answer derivations and explanations. For MATH-Hard, only the hardest questions were kept (Level 5).… See the full description on the dataset page: https://huggingface.co/datasets/lighteval/MATH-Hard.text1K<n<10K24 likes3.2k downloads2y agoHugging Face17HAERAE-HUB /KMMLU-HARD KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.textquestion-answering1K<n<10K13 likes3.1k downloads3y agoHugging Face18harvardairobotics /FairSeg Dataset Card: FairSeg Dataset Summary FairSeg is a large-scale ophthalmology dataset for studying fairness in medical image segmentation. It contains 10,000 SLO fundus images with pixel-wise optic disc and cup segmentation masks, paired with comprehensive demographic annotations. The dataset is designed to benchmark and improve demographic equity in segmentation models, including foundation models such as SAM (Segment Anything Model). This dataset was introduced at ICLR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairSeg.textimage-segmentation10K<n<100K0 likes3k downloads5mo agoHugging Face19harimo /scorio-lite Scorio Lite contains 1,211,520 sampled attempts from four model configurations and six reasoning benchmarks. Each model was run 80 times on every question. The five competition-math splits contain 186 questions. The superGPQA split contains a frozen, field-balanced sample of 3,600 questions. Each row includes the generation, rule-based grading, scores from CompassVerifier-3B and a reference-free verifier, and aggregate token statistics. Per-model configs also include token strings, log… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-lite.tabulartext-generation1M<n<10M0 likes2.5k downloads1mo agoHugging Face20Harland /AudioMCQ-StrongAC-GeminiCoT [ICLR 2026] [DCASE 2026 Training Set] AudioMCQ-StrongAC-GeminiCoT This dataset is a highly curated subset of the AudioMCQ dataset, containing samples with native Chain-of-Thought (CoT) reasoning from Gemini 3.1 Pro that were answered correctly. Note: The native CoT reasoning has been summarized by Gemini's internal algorithm before output, yet still retains rich audio details including timestamps, acoustic descriptions, and step-by-step temporal analysis. 🏆 DCASE 2026… See the full description on the dataset page: https://huggingface.co/datasets/Harland/AudioMCQ-StrongAC-GeminiCoT.audio10K<n<100K7 likes2.4k downloads2mo agoHugging Face21vcr-org /VCR-wiki-en-hard The VCR-Wiki Dataset for Visual Caption Restoration (VCR) 🏠 Paper | 👩🏻‍💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task. VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-hard.imagevisual-question-answering1M<n<10M2 likes2k downloads2y agoHugging Face22Turki-Alshuaibi /haris-weapon-detection-dataset-curatedimage1 likes2k downloads4mo agoHugging Face23harimo /scorio-trace Scorio Trace contains 192,000 sampled reasoning traces from 20 model configurations and four competition math benchmarks. Each model was run 80 times on each of the 30 questions in every benchmark. Each row contains one complete generation, its rule-based correctness, scores from two reward models, and token-level log probabilities and vocabulary ranks. The 80 generations for one model, task, and question form a candidate pool. They are ordered by seed, so pool[:n] gives a reproducible sample… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-trace.tabulartext-generation100K<n<1M0 likes2k downloads29d agoHugging Face24Harvard-Edge /Wake-Vision Dataset Card for Wake Vision Dataset Description "Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally, it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age, subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.imageimage-classification1M<n<10M11 likes2k downloads10mo agoHugging Face25Zhongzhi1228 /Terminal-Bench-Hard Terminal-Bench Hard Terminal-Bench Hard is a set of 100 challenging terminal-based agent tasks. The tasks cover software engineering, debugging, data processing, system administration, security, scientific computing, and related command-line workflows. Contents tasks/: runnable tasks in Harbor format. metadata/tasks.parquet: searchable task metadata and instructions. Each task directory contains task.toml, instruction.md, an environment/ directory, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Terminal-Bench-Hard.imagequestion-answeringn<1K0 likes2k downloads2mo agoHugging Face26HarrisonPENG /dtuimage10K<n<100K0 likes1.9k downloads6mo agoHugging Face27declare-lab /HarmfulQAPaper | Github | Dataset| Model 📣📣📣: Do check our new multilingual dataset CatQA here used in Safety Vectors:📣📣📣 As a part of our research efforts toward making LLMs more safe for public use, we create HarmfulQA i.e. a ChatGPT-distilled dataset constructed using the Chain of Utterances (CoU) prompt. More details are in our paper Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment HarmfulQA serves as both-a new LLM safety benchmark and an alignment dataset… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/HarmfulQA.texttext-generation1K<n<10K47 likes1.9k downloads3y agoHugging Face28harman /tts-datagen GPT-OSS 120B native reasoning traces for TTS Datagen Summary This dataset contains 2,865 synthetic competitive-programming questions, 45,840 independently sampled GPT-OSS 120B solutions (16 per question), and 50 verified test cases per question (143,250 test cases total). Each solution preserves the model's native reasoning trace separately from its final answer. The reasoning was returned by MetaGen's native Dialog Completion interface as dialog reasoning… See the full description on the dataset page: https://huggingface.co/datasets/harman/tts-datagen.tabulartext-generation100K<n<1M0 likes1.8k downloads15d agoHugging Face29taufeeque /mbpp-hardcodetextn<1K0 likes1.6k downloads1y agoHugging Face30udayl /UCI_HARtext100K<n<1M1 likes1.6k downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.