CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Joschka /big_bench_hardAll rights and obligations of the dataset are with original authors of the paper/dataset. I have merely made this dataset with a MIT licence available on HuggingFace. BIG-Bench Hard Dataset This repository contains a copy of the BIG-Bench Hard dataset. Small edits to the formatting of the dataset are made to integrate it into the Inspect Evals repository, a community contributed LLM evaulations for Inspect AI a framework by the UK AI Safety Institute. The BIG-Bench Hard dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Joschka/big_bench_hard.textquestion-answering1K<n<10K3 likes20k downloads1y agoHugging Face02AweAI-Team /BeyondSWE-harbor BeyondSWE-harbor This repository provides the harbor version of the BeyondSWE benchmark, containing the full task instances in a directory-based (harbor) format, where each instance is stored as an independent folder. 📌 For benchmark definition, data format and detailed evaluation results, please refer to: 🤗 Main Dataset on HuggingFace 🗂️ Data Structure beyondswe/ ├── {instance_id}/ │ ├── environment/ │ ├── solution/ │ ├── tests/ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/BeyondSWE-harbor.texttext-generationn<1K2 likes7.3k downloads6mo agoHugging Face03Infatoshi /kernelbench-hard-traces KernelBench-Hard agent traces Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and B200; roofline-graded. Each .jsonl file is one agent run in Claude-Code session format, viewable with the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename = run id. Live leaderboard: https://kernelbench.com/hard Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.tabulartext-generationn<1K16 likes5.5k downloads2d agoHugging Face04pat-jj /harness-1-train-data Harness-1 Training Data This dataset contains the training data used for Harness-1, plus the retrieval corpora needed to reproduce the training/evaluation environment. Contents The dataset has one train split with a stage column: sft: 899 raw GPT-5.4-generated v8d SFT trajectories produced by generate_sft_ultra_0417.py and used by train_sft_ultra_0417.py from sft_ultra_v8d_data. rl: 3453 SEC training-split query records used for RL (TRAIN_DATASETS=sec… See the full description on the dataset page: https://huggingface.co/datasets/pat-jj/harness-1-train-data.texttext-generation1K<n<10K0 likes4.5k downloads3mo agoHugging Face05dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4.3k downloads1y agoHugging Face06declare-lab /HarmfulQAPaper | Github | Dataset| Model 📣📣📣: Do check our new multilingual dataset CatQA here used in Safety Vectors:📣📣📣 As a part of our research efforts toward making LLMs more safe for public use, we create HarmfulQA i.e. a ChatGPT-distilled dataset constructed using the Chain of Utterances (CoU) prompt. More details are in our paper Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment HarmfulQA serves as both-a new LLM safety benchmark and an alignment dataset… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/HarmfulQA.texttext-generation1K<n<10K47 likes1.9k downloads3y agoHugging Face07harman /tts-datagen GPT-OSS 120B native reasoning traces for TTS Datagen Summary This dataset contains 2,865 synthetic competitive-programming questions, 45,840 independently sampled GPT-OSS 120B solutions (16 per question), and 50 verified test cases per question (143,250 test cases total). Each solution preserves the model's native reasoning trace separately from its final answer. The reasoning was returned by MetaGen's native Dialog Completion interface as dialog reasoning… See the full description on the dataset page: https://huggingface.co/datasets/harman/tts-datagen.tabulartext-generation100K<n<1M0 likes1.8k downloads16d agoHugging Face08harimo /scorio-lite Scorio Lite contains 1,211,520 sampled attempts from four model configurations and six reasoning benchmarks. Each model was run 80 times on every question. The five competition-math splits contain 186 questions. The superGPQA split contains a frozen, field-balanced sample of 3,600 questions. Each row includes the generation, rule-based grading, scores from CompassVerifier-3B and a reference-free verifier, and aggregate token statistics. Per-model configs also include token strings, log… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-lite.tabulartext-generation1M<n<10M0 likes1.4k downloads1mo agoHugging Face09hkust-nlp /dart-math-hard 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX [!IMPORTANT] 🔥 Excited to find our DART-Math-DSMath-7B (Prop2Diff) trained on DART-Math-Hard comparable to the AIMO winner NuminaMath-7B on CoT, but based solely on MATH & GSM8K prompt set, leaving much room to improve! Besides, our DART method is also fully compatible… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-hard.texttext-generation100K<n<1M15 likes1.3k downloads2y agoHugging Face10harithoppil /terminal-bench-2-verified Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version 中文版本 We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version. This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes: Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.documenttext-generationn<1K2 likes1.2k downloads5mo agoHugging Face11harimo /scorio-trace Scorio Trace contains 192,000 sampled reasoning traces from 20 model configurations and four competition math benchmarks. Each model was run 80 times on each of the 30 questions in every benchmark. Each row contains one complete generation, its rule-based correctness, scores from two reward models, and token-level log probabilities and vocabulary ranks. The 80 generations for one model, task, and question form a candidate pool. They are ordered by seed, so pool[:n] gives a reproducible sample… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-trace.tabulartext-generation100K<n<1M0 likes1k downloads1mo agoHugging Face12sigcp /hardtests_problems Dataset Card for HARDTESTS Problems HARDTESTS is a competitive programming dataset containing 47,136 problems collected from 13 different Online Judges (OJs). Each problem includes a problem statement, numerous oracle code solutions, and a set of relatively reliable test cases. Note: Due to their large size, the test cases are stored in a separate dataset. This dataset is presented in the paper HardTests: Synthesizing High-Quality Test Cases for LLM Coding. Project Page… See the full description on the dataset page: https://huggingface.co/datasets/sigcp/hardtests_problems.texttext-generation10K<n<100K13 likes895 downloads1y agoHugging Face13Zek-Takai /glm53-flash-harvest GLM-5.3-Flash On-Policy Harvest 86,006 responses / 246,034,910 generated tokens written by zai-org/GLM-5.3-Flash from its reference FP8 weights, across four harvest rounds, 15 registers and both serving modes (22,016 rows carry the model's inline <think>…</think> chain). It is on-policy text: the corpus records what the target model actually generates, which is what a speculative-decoding drafter (EAGLE-3 / DFlash / DSpark family) has to learn to predict. Everything here is MIT.… See the full description on the dataset page: https://huggingface.co/datasets/Zek-Takai/glm53-flash-harvest.tabulartext-generation100K<n<1M3 likes842 downloads24d agoHugging Face14keryszhan /harbor-swesmith-rl-artifacts Harbor SWE-Smith 强化学习数据产物 本数据集是 Harbor Qwen 工具调用代码智能体强化学习项目使用的冻结任务集,服务于 GRPO、原生价值模型/GAE PPO、训练过程诊断和统一协议评测。 项目已于 2026 年 8 月 30 日完成 P0 评测并进入阶段性归档。本数据集用于保留实验所依赖的数据切分、任务执行文件和审计信息,不代表新的通用代码能力基准。 数据概况 切分 任务数 训练集 187 验证集 42 测试集 38 合计 267 数据覆盖 89 个上游代码仓库。三个切分之间同时执行任务标识和仓库级隔离检查。 正式数据集名称: swesmith-curated-grpo-267-v1 冻结切分的语义摘要: ae5df9a3f4a3fc8af44fac420b36529e283839e1bd3de9daba65d5bcda51447d 该值来自 split-manifest.json 的 sha256 字段,用于标识切分语义,不等同于该文件本身的字节级… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/harbor-swesmith-rl-artifacts.tabulartext-generationn<1K0 likes784 downloads21d agoHugging Face15joelniklaus /LEXam-hard LEXam-hard The 518 open questions of LEXam that the strongest open models score lowest on. LEXam is a benchmark of law exam questions from the University of Zurich (Fan et al., ICLR 2026; website, code). This is a filtered copy of its open_question test split for evaluating agents that can do legal research. Questions that strong models already answer from memory are removed, and so are questions that cannot be answered from their own text. Rows are ordered from the hardest… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/LEXam-hard.texttext-generationn<1K1 likes583 downloads15d agoHugging Face16hardcoremoore /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/hardcoremoore/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K2 likes503 downloads4mo agoHugging Face17violetxi /harvey-notes-v4 wm-rl notes v4 — two note banks from a recursive self-experience loop Continuation update (2026-09-14): rounds 6–15 appended. The recursive bank now contains 308,580 notes / 138,723,516 training tokens. The original 108,081-row round-5 bank remains an exact prefix. Round-4/5 task lists were available for duplicate rejection, but their trajectory manifests were not published; the continuation therefore seeds round 6 from the published recall sessions and carries complete… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-notes-v4.texttext-generation100K<n<1M0 likes422 downloads6d agoHugging Face18Roman1111111 /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K61 likes344 downloads7mo agoHugging Face19HarleyCooper /volume2gym-railroad-1959 The Rulebook Becomes a World Volume2Gym asks a deliberately expansive question: what if any sufficiently structured text could become a small, inspectable world in which a model learns by acting, receiving feedback, and trying again? This release turns one bounded English technical volume into an auditable reinforcement-learning dataset: 117 source scans → 536 extracted rules → 2,708 synthetic scenario tasks → a measured rule-linkage and verification surface. It is an… See the full description on the dataset page: https://huggingface.co/datasets/HarleyCooper/volume2gym-railroad-1959.imagequestion-answering1K<n<10K1 likes327 downloads2mo agoHugging Face20wAI-org /swerl-tmax-15k-hardened-prefilter swerl-tmax-15k hardened, pre-validation-filter (dataset 2 of 3) The middle artifact of three, which exist to be compared against each other by task_id: original — hamishivi/swerl-tmax-15k, unchanged. hardened, pre-validation-filter — this dataset. hardened, post-validation-filter — wAI-org/swerl-tmax-15k-solvable-gpt-5-6-terra, a strict subset of this one — 7,015 tasks. What is in it Tasks where a patch was actually applied, plus tasks originally labelled CLEAN… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-hardened-prefilter.tabulartext-generation10K<n<100K0 likes308 downloads11d agoHugging Face21Skyhigh-2203 /MiMo-2.5-Pro-Reasoning-Traces-Hard MiMo-2.5-Pro-Reasoning-Traces-Hard A large-scale reasoning dataset of 8,706 expert-level prompts with full reasoning traces across 44 academic and technical topics, generated using the MiMo-v2.5-Pro model. Each entry contains the step-by-step reasoning chain alongside the final completion, designed for training and evaluating advanced reasoning capabilities in language models. Dataset Statistics Metric Value Total entries 8,706 Unique topics 44… See the full description on the dataset page: https://huggingface.co/datasets/Skyhigh-2203/MiMo-2.5-Pro-Reasoning-Traces-Hard.texttext-generation1K<n<10K12 likes299 downloads3mo agoHugging Face22violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 5.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.tabulartext-generation1K<n<10K0 likes297 downloads5d agoHugging Face23JWei05 /DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k DeepScaleR Easy/Medium/Hard — Gemma 4 26B-A4B PT This dataset contains 9,900 unique, deduplicated DeepScaleR math questions for reinforcement-learning experiments. Difficulty is defined by how often the pretrained google/gemma-4-26B-A4B teacher solved each question across eight temperature-1 samples under the same rule-based grader used by the RL training pipeline. The Hub dataset has three configurations—easy, medium, and hard—and each configuration has a train split with 3,000… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k.texttext-generation1K<n<10K0 likes293 downloads1mo agoHugging Face24violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 1.3000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.tabulartext-generation1K<n<10K0 likes280 downloads5d agoHugging Face25violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 4.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.tabulartext-generation1K<n<10K0 likes280 downloads5d agoHugging Face26haralab-uec /steering-bench-ja steering-bench-ja Dataset Summary steering-bench-ja is a Japanese benchmark for evaluating steering vectors in large language models. The dataset is constructed by translating existing English evaluation benchmarks, including Anthropic/model-written-evals (MWE) and truthfulqa/truthful_qa, into Japanese using the pfnet/plamo-2-translate model. The benchmark is designed to evaluate in-distribution reliability and out-of-distribution generalization of steering vectors… See the full description on the dataset page: https://huggingface.co/datasets/haralab-uec/steering-bench-ja.texttext-generation100K<n<1M0 likes224 downloads7mo agoHugging Face27haryoaw /cultural-spyfall Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game This dataset contains the game histories from Multicultural Spyfall, a dynamic benchmarking framework designed to evaluate the multilingual and multicultural capabilities of Large Language Models (LLMs). The dataset is introduced in the paper: Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game. Dataset Summary Multicultural Spyfall uses… See the full description on the dataset page: https://huggingface.co/datasets/haryoaw/cultural-spyfall.texttext-generation1K<n<10K1 likes221 downloads8mo agoHugging Face28yavuz-ai /seas-harden-only seas-harden-only Per-round trajectory + per-category taxonomy from an automated red-teaming co-evolution loop (attacker -> target -> harm-judge) on Qwen2.5 (7B attacker+judge, 3B target), seeded with JailbreakBench behaviors. This dataset is the frozen-attacker control, canned refusal (only the target hardens) arm: held-in ASR flat ~0.45 across 5 rounds - LoRA-DPO hardening toward a canned refusal does not reduce ASR. Study question: does adversarial co-evolution ignite an arms… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/seas-harden-only.tabulartext-generationn<1K0 likes214 downloads3mo agoHugging Face29Harvard-DCML /tis-subset-datasets-Llama-2-7b-hf Targeted Instruction Selection Subsets (Llama-2-7b-hf) This repository contains pre-computed instruction training subsets selected from a large candidate pool for targeted instruction fine-tuning, as presented in the paper A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't). Paper: https://huggingface.co/papers/2602.14696 GitHub Repository: https://github.com/dcml-lab/targeted-instruction-selection Description Instruction… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-DCML/tis-subset-datasets-Llama-2-7b-hf.texttext-generation100K<n<1M0 likes173 downloads7mo agoHugging Face30harithoppil /terminal-bench-2-trajectories Terminal-Bench 2.0 Leaderboard Trajectories Agent trajectories extracted from Terminal-Bench 2.0 leaderboard submissions. Each row contains a prompt (task instruction), the agent's response, and the reward (pass/fail). Models Included Model Trials Passed Claude-Opus-4.6 2,213 1,537 (69%) Gemini-3.1-Pro-Preview 445 333 (75%) GLM-5 445 231 (52%) Kimi-k2.5 442 189 (43%) Claude-Opus-4.5 178 98 (55%) Splits Config Description Rows… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-trajectories.texttext-generation1K<n<10K2 likes165 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.