CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Infatoshi /kernelbench-hard-traces KernelBench-Hard agent traces Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and B200; roofline-graded. Each .jsonl file is one agent run in Claude-Code session format, viewable with the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename = run id. Live leaderboard: https://kernelbench.com/hard Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.tabulartext-generationn<1K16 likes5.3k downloads1d agoHugging Face02dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4.1k downloads1y agoHugging Face03declare-lab /HarmfulQAPaper | Github | Dataset| Model 📣📣📣: Do check our new multilingual dataset CatQA here used in Safety Vectors:📣📣📣 As a part of our research efforts toward making LLMs more safe for public use, we create HarmfulQA i.e. a ChatGPT-distilled dataset constructed using the Chain of Utterances (CoU) prompt. More details are in our paper Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment HarmfulQA serves as both-a new LLM safety benchmark and an alignment dataset… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/HarmfulQA.texttext-generation1K<n<10K47 likes1.9k downloads3y agoHugging Face04harithoppil /terminal-bench-2-verified Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version 中文版本 We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version. This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes: Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.documenttext-generationn<1K2 likes1.3k downloads5mo agoHugging Face05keryszhan /harbor-swesmith-rl-artifacts Harbor SWE-Smith 强化学习数据产物 本数据集是 Harbor Qwen 工具调用代码智能体强化学习项目使用的冻结任务集,服务于 GRPO、原生价值模型/GAE PPO、训练过程诊断和统一协议评测。 项目已于 2026 年 8 月 30 日完成 P0 评测并进入阶段性归档。本数据集用于保留实验所依赖的数据切分、任务执行文件和审计信息,不代表新的通用代码能力基准。 数据概况 切分 任务数 训练集 187 验证集 42 测试集 38 合计 267 数据覆盖 89 个上游代码仓库。三个切分之间同时执行任务标识和仓库级隔离检查。 正式数据集名称: swesmith-curated-grpo-267-v1 冻结切分的语义摘要: ae5df9a3f4a3fc8af44fac420b36529e283839e1bd3de9daba65d5bcda51447d 该值来自 split-manifest.json 的 sha256 字段,用于标识切分语义,不等同于该文件本身的字节级… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/harbor-swesmith-rl-artifacts.tabulartext-generationn<1K0 likes779 downloads21d agoHugging Face06hardcoremoore /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/hardcoremoore/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K2 likes500 downloads4mo agoHugging Face07HarleyCooper /volume2gym-railroad-1959 The Rulebook Becomes a World Volume2Gym asks a deliberately expansive question: what if any sufficiently structured text could become a small, inspectable world in which a model learns by acting, receiving feedback, and trying again? This release turns one bounded English technical volume into an auditable reinforcement-learning dataset: 117 source scans → 536 extracted rules → 2,708 synthetic scenario tasks → a measured rule-linkage and verification surface. It is an… See the full description on the dataset page: https://huggingface.co/datasets/HarleyCooper/volume2gym-railroad-1959.imagequestion-answering1K<n<10K1 likes341 downloads1mo agoHugging Face08Roman1111111 /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K62 likes336 downloads7mo agoHugging Face09Skyhigh-2203 /MiMo-2.5-Pro-Reasoning-Traces-Hard MiMo-2.5-Pro-Reasoning-Traces-Hard A large-scale reasoning dataset of 8,706 expert-level prompts with full reasoning traces across 44 academic and technical topics, generated using the MiMo-v2.5-Pro model. Each entry contains the step-by-step reasoning chain alongside the final completion, designed for training and evaluating advanced reasoning capabilities in language models. Dataset Statistics Metric Value Total entries 8,706 Unique topics 44… See the full description on the dataset page: https://huggingface.co/datasets/Skyhigh-2203/MiMo-2.5-Pro-Reasoning-Traces-Hard.texttext-generation1K<n<10K12 likes294 downloads3mo agoHugging Face10Bingguang /HardGen From Failure to Mastery: Generating Hard Samples for Tool-use Agents This is only a demonstration result obtained by directly deploying and sampling HardGen in the BFCL environment, not the dataset directly used in the HardGen paper. [!IMPORTANT] Important Hint This is an extension of the technical report FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use To allow the model to learn from errors, we specifically construct… See the full description on the dataset page: https://huggingface.co/datasets/Bingguang/HardGen.textquestion-answering10K<n<100K80 likes169 downloads6mo agoHugging Face11harithoppil /terminal-bench-2-trajectories Terminal-Bench 2.0 Leaderboard Trajectories Agent trajectories extracted from Terminal-Bench 2.0 leaderboard submissions. Each row contains a prompt (task instruction), the agent's response, and the reward (pass/fail). Models Included Model Trials Passed Claude-Opus-4.6 2,213 1,537 (69%) Gemini-3.1-Pro-Preview 445 333 (75%) GLM-5 445 231 (52%) Kimi-k2.5 442 189 (43%) Claude-Opus-4.5 178 98 (55%) Splits Config Description Rows… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-trajectories.texttext-generation1K<n<10K2 likes160 downloads5mo agoHugging Face12harleygilpin /soc-audit-11k SOC Audit Text Generation Dataset Description This dataset is designed for training and evaluating Language Models (LLMs) specifically in the context of SOC 2 audits. It covers a wide range of topics including, but not limited to, information security, risk management, compliance, data privacy, and governance. The dataset consists of structured text in the format of instructions followed by a detailed response, making it ideal for models intended to assist in… See the full description on the dataset page: https://huggingface.co/datasets/harleygilpin/soc-audit-11k.texttext-generation10K<n<100K9 likes159 downloads3y agoHugging Face13stindardlogic /instruction-following-hard-sft-100k Hard Instruction Following SFT (100K) 100,000 ShareGPT conversations where the assistant correctly satisfies multiple simultaneous explicit constraints in a single response. Each example pairs a multi-constraint prompt with a response that honors every constraint without dropping any. Targets the instruction-following capability measured by IFEval and similar benchmarks. Motivation A key failure mode in deployed LLMs is dropping constraints under load — responding… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/instruction-following-hard-sft-100k.texttext-generation100K<n<1M0 likes150 downloads2mo agoHugging Face14harshildarji /openlegaldataClean Open Legal Data Overview | Dataset Structure | Key Fields | Example Entry | Using the Dataset with Python | Citation | License Overview This dataset is a comprehensive collection of open legal case records in JSONL format. It comprises 423,941 cases extracted and processed from the Open Legal Data dump dump-20260520 and represents an independent, cleaned derivative of that source data. The dataset is designed… See the full description on the dataset page: https://huggingface.co/datasets/harshildarji/openlegaldata.texttext-classification100K<n<1M5 likes145 downloads4mo agoHugging Face15hari-krishna-ai /enterprise-text-to-sql-benchmark Enterprise Text-to-SQL Benchmark 3,087 natural-language questions paired with executable PostgreSQL, over a 12-table enterprise schema (sales, catalogue, logistics, HR). Built to answer one question honestly: does fine-tuning actually improve text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a self-correction loop — and the benchmark is designed so that number cannot be inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.texttable-question-answering1K<n<10K0 likes130 downloads3d agoHugging Face16TrustAIRLab /HarmfulSkillBenchgated 📝 Paper  |  📑 arXiv  |  💻 Code  |  📦 Dataset HarmfulSkillBench A benchmark for evaluating LLM refusal behavior when agents are exposed to skills that describe potentially harmful capabilities. The benchmark probes whether current LLMs can detect and refuse harmful agent skills in two settings. Tier 1 covers prohibited behaviors that should always be refused. Tier 2 covers high-risk domains where responses should include human-in-the-loop referral and AI… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/HarmfulSkillBench.texttext-generationn<1K4 likes114 downloads5mo agoHugging Face17ebowwa /needle2-harness-dispatch Needle-2 Harness-Dispatch Corpus (review build) Eval/training corpus for tool-dispatch on a developer-agent harness surface (8 tools: bash / read / write / edit / glob / grep / web_search / todo_write). This is a review build — every record carries QA annotations so a human can approve, relabel, or flag before the next training run. Provenance Generated and judged by glm-5.3 , two generation rounds (seeds 7 and 101), judge pass kept/fixed/dropped. 1,307 raw… See the full description on the dataset page: https://huggingface.co/datasets/ebowwa/needle2-harness-dispatch.texttext-generation1K<n<10K0 likes101 downloads1mo agoHugging Face18Skywalker-Harrison-mbz /ArabPref ArabPref Preference Test This repository contains the English and Arabic preference test data from ArabPref. It also contains the English and Arabic MCQ test data. Files pref_test_en.jsonl: 3,300 English examples. pref_test_ar.jsonl: 2,964 Arabic examples. mcq_test_en.jsonl: 992 English multiple-choice questions. mcq_test_ar.jsonl: 992 Arabic multiple-choice questions, including the revised items. The two files are exposed as separate configurations because they… See the full description on the dataset page: https://huggingface.co/datasets/Skywalker-Harrison-mbz/ArabPref.texttext-generation1K<n<10K0 likes101 downloads1mo agoHugging Face19aurora-m /biden-harris-redteam-archived THIS IS AN ARCHIVED VERSION Biden-Harris Redteam: A red-teaming dataset focusing on the Biden-Harris AI Executive Order Dataset Description While building Large Language Models (LLMs), it is crucial to protect them against attacks that could bypass safety guardrails and break their guiding principles. Specifically, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to the harm of the… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/biden-harris-redteam-archived.texttext-generation10K<n<100K7 likes98 downloads1y agoHugging Face20DJLougen /harmonic-reasoning-v1 Harmonic Reasoning v1 Support This Work I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases. Support on Ko-fi Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/harmonic-reasoning-v1.tabulartext-generationn<1K27 likes94 downloads6mo agoHugging Face21kaushik-harsh-99 /math-sft-solutions-no-cot Math SFT Solutions No CoT A cleaned mathematics supervised fine-tuning dataset containing: instruction → solution pairs mathematical proofs derivations olympiad-style solutions theorem reasoning stepwise mathematical explanations detailed final solutions This dataset was built specifically for mathematical supervised fine-tuning (SFT). Unlike many reasoning datasets, this release removes explicit chain-of-thought tags and hidden thinking traces while preserving high-quality… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/math-sft-solutions-no-cot.texttext-generation100K<n<1M5 likes94 downloads4mo agoHugging Face22harukicoder /hsk30-graded-readers HSK 3.0 Graded Reader Corpus 132 word-aligned Chinese graded readers with per-word pinyin and English gloss, arranged on six difficulty shelves: 102 texts in the main split and a disjoint 30-text held-out split. Aligned Chinese graded-reader corpora are scarce. Existing collections are unaligned plain text, locked inside commercial apps, or graded against HSK 2.0, which has been superseded twice. Loading from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/harukicoder/hsk30-graded-readers.tabulartext-classificationn<1K0 likes93 downloads24d agoHugging Face23Roman1111111 /gemini-3-pro-10000x-hard-high-reasoning Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning Dataset Details Dataset Description Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement. This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3-pro-10000x-hard-high-reasoning.textquestion-answering10K<n<100K57 likes92 downloads7mo agoHugging Face24sinatras /pmpp-hard PMPP-Hard Agent Evaluation Traces PMPP-Hard is a 69-task agentic GPU-kernel evaluation for testing whether autonomous coding agents can produce implementations that are both correct and performant. This dataset contains the complete nine-model campaign used in the PMPP-Hard release: 621 rollouts, with 69 task sessions for each model configuration. Maintained and released by Sinatras. Source repository: SinatrasC/pmpp-hard Prime environment and evaluations: PMPP-Hard on Prime… See the full description on the dataset page: https://huggingface.co/datasets/sinatras/pmpp-hard.texttext-generationn<1K2 likes91 downloads2mo agoHugging Face25daichira /structured-hard-sft-4k Hard Synthetic Dataset for Structured Data Tasks (v1) This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types). Dataset Summary The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.texttext-generation1K<n<10K1 likes85 downloads8mo agoHugging Face26kaushik-harsh-99 /Indian-legal-data-v3 Indian Legal Dataset V3 Overview Indian Legal Dataset V3 is a large-scale instruction-tuning dataset focused on Indian law, constitutional law, criminal law, legal reasoning, legal drafting, and real-world legal assistance. Compared to V2, this version expands the dataset with: legal drafting instruction pairs, hypothetical legal scenarios, detailed IPC-focused data, practical real-world legal instructions, concise legal QA pairs. After integrating the new data sources… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v3.texttext-generation100K<n<1M7 likes85 downloads4mo agoHugging Face27kaushik-harsh-99 /math-sft-solutions-no-cot-v3 Math SFT Solutions No CoT V3 Math SFT Solutions No CoT V3 is a large-scale mathematics supervised fine-tuning (SFT) dataset designed for instruction tuning and mathematical capability adaptation. Version 3 substantially expands mathematical coverage while improving dataset quality through stronger filtering, cleaning, and supervision refinement. Unlike reasoning-heavy datasets, this release focuses on clean instruction → response pairs without hidden chain-of-thought style… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/math-sft-solutions-no-cot-v3.texttext-generation1M<n<10M5 likes78 downloads4mo agoHugging Face28kaushik-harsh-99 /Uncensored-SFT-v1 Dataset Creation Process This dataset was not scraped from a single source. Instead, it was built through a large multi-stage curation and cleaning pipeline involving many open instruction datasets available on Hugging Face. The entire dataset was normalized into a unified: { "input": "...", "output": "..." } format. Data Collection A large number of public instruction datasets were downloaded from Hugging Face. These datasets included: Instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Uncensored-SFT-v1.texttext-generation100K<n<1M3 likes72 downloads5mo agoHugging Face29Harry-1234 /IntentRouterTrain MODF-SIR: a Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning MODF-SIR is a lightweight MLLM-based, distillation-augmented, multi-agent collaborative framework for social intelligence reasoning. 🔖 Model Details Model type: Omni-modal Large Language Model License: BSD-3-Clause Project Page: Arxiv: https://arxiv.org/abs/2606.12018 Code Repository: GitHub: https://github.com/eeee-sys/MODF-SIR 👀 MODF-SIR… See the full description on the dataset page: https://huggingface.co/datasets/Harry-1234/IntentRouterTrain.texttext-generationn<1K0 likes70 downloads4mo agoHugging Face30AbiralArch /hardware-verilogeval-v2 hardware-verilogeval-v2 VerilogEval v2 - 471 Verilog evaluation problems Dataset Overview This dataset is part of a comprehensive collection of hardware design datasets for training and evaluating LLMs on Verilog/SystemVerilog code generation and hardware design tasks. Files verilog_eval_problems.json: 471 VerilogEval v2 problems Usage from datasets import load_dataset # Load the dataset dataset = load_dataset('AbiralArch/hardware-verilogeval-v2')… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-verilogeval-v2.texttext-generationn<1K0 likes68 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.