CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01harborframework /terminal-bench-2.0Warning: The leaderboard above is unofficial. The official leaderboard is https://www.tbench.ai/leaderboard/terminal-bench/2.0, in which entires are audited for correct configuration, results show which agent harness is used, and verified trajectories are publicly viewable. Warning: The dataset is a read-only mirror. The primary source for this dataset is on GitHub: https://github.com/harbor-framework/terminal-bench-2. Please open issues and pull requests there. How this mirror was created… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.0.text-generationn<1K51 likes95k downloads5mo agoHugging Face02Joschka /big_bench_hardAll rights and obligations of the dataset are with original authors of the paper/dataset. I have merely made this dataset with a MIT licence available on HuggingFace. BIG-Bench Hard Dataset This repository contains a copy of the BIG-Bench Hard dataset. Small edits to the formatting of the dataset are made to integrate it into the Inspect Evals repository, a community contributed LLM evaulations for Inspect AI a framework by the UK AI Safety Institute. The BIG-Bench Hard dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Joschka/big_bench_hard.textquestion-answering1K<n<10K3 likes19k downloads1y agoHugging Face03dell-research-harvard /AmericanStoriesAmerican Stories offers high-quality structured data from historical newspapers suitable for pre-training large language models to enhance the understanding of historical English and world knowledge. It can also be integrated into external databases of retrieval-augmented language models, enabling broader access to historical information, including interpretations of political events and intricate details about people's ancestors. Additionally, the structured article texts facilitate the application of transformer-based methods for popular tasks like detecting reproduced content, significantly improving accuracy compared to traditional OCR methods. American Stories serves as a substantial and valuable dataset for advancing multimodal layout analysis models and other multimodal applications.text-classification100M<n<1B176 likes11k downloads2y agoHugging Face04AweAI-Team /BeyondSWE-harbor BeyondSWE-harbor This repository provides the harbor version of the BeyondSWE benchmark, containing the full task instances in a directory-based (harbor) format, where each instance is stored as an independent folder. 📌 For benchmark definition, data format and detailed evaluation results, please refer to: 🤗 Main Dataset on HuggingFace 🗂️ Data Structure beyondswe/ ├── {instance_id}/ │ ├── environment/ │ ├── solution/ │ ├── tests/ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/BeyondSWE-harbor.texttext-generationn<1K2 likes7.2k downloads6mo agoHugging Face05Infatoshi /kernelbench-hard-traces KernelBench-Hard agent traces Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and B200; roofline-graded. Each .jsonl file is one agent run in Claude-Code session format, viewable with the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename = run id. Live leaderboard: https://kernelbench.com/hard Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.tabulartext-generationn<1K16 likes5.3k downloads23h agoHugging Face06pat-jj /harness-1-train-data Harness-1 Training Data This dataset contains the training data used for Harness-1, plus the retrieval corpora needed to reproduce the training/evaluation environment. Contents The dataset has one train split with a stage column: sft: 899 raw GPT-5.4-generated v8d SFT trajectories produced by generate_sft_ultra_0417.py and used by train_sft_ultra_0417.py from sft_ultra_v8d_data. rl: 3453 SEC training-split query records used for RL (TRAIN_DATASETS=sec… See the full description on the dataset page: https://huggingface.co/datasets/pat-jj/harness-1-train-data.texttext-generation1K<n<10K0 likes4.5k downloads3mo agoHugging Face07dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4.1k downloads1y agoHugging Face08harimo /scorio-lite Scorio Lite contains 1,211,520 sampled attempts from four model configurations and six reasoning benchmarks. Each model was run 80 times on every question. The five competition-math splits contain 186 questions. The superGPQA split contains a frozen, field-balanced sample of 3,600 questions. Each row includes the generation, rule-based grading, scores from CompassVerifier-3B and a reference-free verifier, and aggregate token statistics. Per-model configs also include token strings, log… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-lite.tabulartext-generation1M<n<10M0 likes2.5k downloads1mo agoHugging Face09harimo /scorio-trace Scorio Trace contains 192,000 sampled reasoning traces from 20 model configurations and four competition math benchmarks. Each model was run 80 times on each of the 30 questions in every benchmark. Each row contains one complete generation, its rule-based correctness, scores from two reward models, and token-level log probabilities and vocabulary ranks. The 80 generations for one model, task, and question form a candidate pool. They are ordered by seed, so pool[:n] gives a reproducible sample… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-trace.tabulartext-generation100K<n<1M0 likes2k downloads29d agoHugging Face10jjjlimaus /chrono-2020-harvestgated ChronoLLM 2020 harvest Public gated (gated=manual) dataset of year-cutoff training text for Bittensor subnet 38. Calendar window 2020-01-01 … 2020-12-31. Repo: jjjlimaus/chrono-2020-harvest Last sync (UTC): 2026-09-04T18:45:47.313133+00:00 Request access on Hugging Face and wait for manual review. Do not treat this as a license to republish every source; some feeds are public-domain, others are news/API text kept for research training only. Streaming layout… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/chrono-2020-harvest.text-generation1 likes1.9k downloads21d agoHugging Face11harborframework /harbor-mixgated Harbor-Mix Harbor-Mix is a curated meta-dataset of 100 difficult, diverse, and high-quality agentic evaluation tasks selected from the Harbor Adapters benchmark pool. It is designed to preserve broad signal from large-scale agent evaluations while being substantially cheaper to run than a full multi-benchmark sweep. What Is Included This repository contains the 100 task directories, flattened at the repository root. Each task directory contains: instruction.md:… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/harbor-mix.text-generationn<1K3 likes1.9k downloads5mo agoHugging Face12declare-lab /HarmfulQAPaper | Github | Dataset| Model 📣📣📣: Do check our new multilingual dataset CatQA here used in Safety Vectors:📣📣📣 As a part of our research efforts toward making LLMs more safe for public use, we create HarmfulQA i.e. a ChatGPT-distilled dataset constructed using the Chain of Utterances (CoU) prompt. More details are in our paper Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment HarmfulQA serves as both-a new LLM safety benchmark and an alignment dataset… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/HarmfulQA.texttext-generation1K<n<10K47 likes1.9k downloads3y agoHugging Face13harman /tts-datagen GPT-OSS 120B native reasoning traces for TTS Datagen Summary This dataset contains 2,865 synthetic competitive-programming questions, 45,840 independently sampled GPT-OSS 120B solutions (16 per question), and 50 verified test cases per question (143,250 test cases total). Each solution preserves the model's native reasoning trace separately from its final answer. The reasoning was returned by MetaGen's native Dialog Completion interface as dialog reasoning… See the full description on the dataset page: https://huggingface.co/datasets/harman/tts-datagen.tabulartext-generation100K<n<1M0 likes1.8k downloads15d agoHugging Face14harithoppil /terminal-bench-2-verified Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version 中文版本 We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version. This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes: Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.documenttext-generationn<1K2 likes1.3k downloads5mo agoHugging Face15hkust-nlp /dart-math-hard 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX [!IMPORTANT] 🔥 Excited to find our DART-Math-DSMath-7B (Prop2Diff) trained on DART-Math-Hard comparable to the AIMO winner NuminaMath-7B on CoT, but based solely on MATH & GSM8K prompt set, leaving much room to improve! Besides, our DART method is also fully compatible… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-hard.texttext-generation100K<n<1M15 likes1.2k downloads2y agoHugging Face16jjjlimaus /chrono-2021-harvestgated ChronoLLM 2021 harvest Public gated (gated=manual) dataset of year-cutoff training text for Bittensor subnet 38. Calendar window 2021-01-01 … 2021-12-31. Repo: jjjlimaus/chrono-2021-harvest Last sync (UTC): 2026-09-12T17:44:33.508539+00:00 Request access on Hugging Face and wait for manual review. Do not treat this as a license to republish every source; some feeds are public-domain, others are news/API text kept for research training only. Streaming layout… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/chrono-2021-harvest.text-generation3 likes928 downloads13d agoHugging Face17sigcp /hardtests_problems Dataset Card for HARDTESTS Problems HARDTESTS is a competitive programming dataset containing 47,136 problems collected from 13 different Online Judges (OJs). Each problem includes a problem statement, numerous oracle code solutions, and a set of relatively reliable test cases. Note: Due to their large size, the test cases are stored in a separate dataset. This dataset is presented in the paper HardTests: Synthesizing High-Quality Test Cases for LLM Coding. Project Page… See the full description on the dataset page: https://huggingface.co/datasets/sigcp/hardtests_problems.texttext-generation10K<n<100K13 likes898 downloads1y agoHugging Face18keryszhan /harbor-swesmith-rl-artifacts Harbor SWE-Smith 强化学习数据产物 本数据集是 Harbor Qwen 工具调用代码智能体强化学习项目使用的冻结任务集,服务于 GRPO、原生价值模型/GAE PPO、训练过程诊断和统一协议评测。 项目已于 2026 年 8 月 30 日完成 P0 评测并进入阶段性归档。本数据集用于保留实验所依赖的数据切分、任务执行文件和审计信息,不代表新的通用代码能力基准。 数据概况 切分 任务数 训练集 187 验证集 42 测试集 38 合计 267 数据覆盖 89 个上游代码仓库。三个切分之间同时执行任务标识和仓库级隔离检查。 正式数据集名称: swesmith-curated-grpo-267-v1 冻结切分的语义摘要: ae5df9a3f4a3fc8af44fac420b36529e283839e1bd3de9daba65d5bcda51447d 该值来自 split-manifest.json 的 sha256 字段,用于标识切分语义,不等同于该文件本身的字节级… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/harbor-swesmith-rl-artifacts.tabulartext-generationn<1K0 likes779 downloads20d agoHugging Face19Zek-Takai /glm53-flash-harvest GLM-5.3-Flash On-Policy Harvest 86,006 responses / 246,034,910 generated tokens written by zai-org/GLM-5.3-Flash from its reference FP8 weights, across four harvest rounds, 15 registers and both serving modes (22,016 rows carry the model's inline <think>…</think> chain). It is on-policy text: the corpus records what the target model actually generates, which is what a speculative-decoding drafter (EAGLE-3 / DFlash / DSpark family) has to learn to predict. Everything here is MIT.… See the full description on the dataset page: https://huggingface.co/datasets/Zek-Takai/glm53-flash-harvest.tabulartext-generation100K<n<1M3 likes776 downloads23d agoHugging Face20harman /tts TTS synthetic programming-question dataset This public dataset contains the frozen set of 2,865 accepted programming questions and their materialized verifier tests. Accepted question bundle accepted_bundle/accepted-questions-2865.tar.zst contains all 42,975 accepted question files: statements, package JSON, three public examples, generators, validators, reference solutions, brute-force solutions, verification records, provenance records, and GPT-OSS hardness… See the full description on the dataset page: https://huggingface.co/datasets/harman/tts.text-generation0 likes768 downloads22d agoHugging Face21sigcp /hardtests_tests Dataset Card for HARDTESTS Tests HARDTESTS Tests is the test suite of HARDTESTS, a competitive programming dataset. The dataset contains multiple .parquet files. Each instance contains structured test case objects for one problem. This dataset is presented in the paper HardTests: Synthesizing High-Quality Test Cases for LLM Coding. Project Page Data Summary The test suite is generated using the HARDTESTSGEN pipeline. The dataset contains the generated test suites of… See the full description on the dataset page: https://huggingface.co/datasets/sigcp/hardtests_tests.text-generation2 likes708 downloads10mo agoHugging Face22Harryis /LifelongSAThis is the two iteration defender of NeurIPS 2025 "Lifelong Safety Alignment for Language Models": https://openreview.net/forum?id=9YkEcAqiIK The defenders are trained on RR and LAT. text-generation0 likes572 downloads8mo agoHugging Face23joelniklaus /LEXam-hard LEXam-hard The 518 open questions of LEXam that the strongest open models score lowest on. LEXam is a benchmark of law exam questions from the University of Zurich (Fan et al., ICLR 2026; website, code). This is a filtered copy of its open_question test split for evaluating agents that can do legal research. Questions that strong models already answer from memory are removed, and so are questions that cannot be answered from their own text. Rows are ordered from the hardest… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/LEXam-hard.texttext-generationn<1K1 likes550 downloads15d agoHugging Face24hardcoremoore /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/hardcoremoore/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K2 likes500 downloads4mo agoHugging Face25jjjlimaus /chrono-2022-harvestgated ChronoLLM 2022 harvest Public gated (gated=manual) dataset of year-cutoff training text for Bittensor subnet 38. Calendar window 2022-01-01 … 2022-12-31. Repo: jjjlimaus/chrono-2022-harvest Last sync (UTC): 2026-09-17T12:52:33.701253+00:00 Request access on Hugging Face and wait for manual review. Do not treat this as a license to republish every source; some feeds are public-domain, others are news/API text kept for research training only. Streaming layout… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/chrono-2022-harvest.text-generation0 likes423 downloads8d agoHugging Face26jjjlimaus /chrono-2021-quality-harvestgated ChronoLLM 2021 quality cutoff corpus Clean, cross-source-deduplicated training text available no later than 2021-12-31. Target: 15B exact anacoluthe89/chrono-2015 tokens. The build rejects any row that would exceed 15B; 20B is an external safety ceiling, not a collection goal. Repository: jjjlimaus/chrono-2021-quality-harvest Initialized: 2026-09-17T13:43:12.278373+00:00 Mix and hard quotas Lane Maximum tokens Cutoff basis English Wikipedia snapshot… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/chrono-2021-quality-harvest.text-generation0 likes404 downloads8d agoHugging Face27violetxi /harvey-notes-v4 wm-rl notes v4 — two note banks from a recursive self-experience loop Continuation update (2026-09-14): rounds 6–15 appended. The recursive bank now contains 308,580 notes / 138,723,516 training tokens. The original 108,081-row round-5 bank remains an exact prefix. Round-4/5 task lists were available for duplicate rejection, but their trajectory manifests were not published; the continuation therefore seeds round 6 from the published recall sessions and carries complete… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-notes-v4.texttext-generation100K<n<1M0 likes385 downloads6d agoHugging Face28HarleyCooper /volume2gym-railroad-1959 The Rulebook Becomes a World Volume2Gym asks a deliberately expansive question: what if any sufficiently structured text could become a small, inspectable world in which a model learns by acting, receiving feedback, and trying again? This release turns one bounded English technical volume into an auditable reinforcement-learning dataset: 117 source scans → 536 extracted rules → 2,708 synthetic scenario tasks → a measured rule-linkage and verification surface. It is an… See the full description on the dataset page: https://huggingface.co/datasets/HarleyCooper/volume2gym-railroad-1959.imagequestion-answering1K<n<10K1 likes341 downloads1mo agoHugging Face29Roman1111111 /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K62 likes336 downloads7mo agoHugging Face30JWei05 /DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k DeepScaleR Easy/Medium/Hard — Gemma 4 26B-A4B PT This dataset contains 9,900 unique, deduplicated DeepScaleR math questions for reinforcement-learning experiments. Difficulty is defined by how often the pretrained google/gemma-4-26B-A4B teacher solved each question across eight temperature-1 samples under the same rule-based grader used by the RL training pipeline. The Hub dataset has three configurations—easy, medium, and hard—and each configuration has a train split with 3,000… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k.texttext-generation1K<n<10K0 likes304 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.