CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Infatoshi /kernelbench-hard-traces KernelBench-Hard agent traces Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and B200; roofline-graded. Each .jsonl file is one agent run in Claude-Code session format, viewable with the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename = run id. Live leaderboard: https://kernelbench.com/hard Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.tabulartext-generationn<1K16 likes5.3k downloads1d agoHugging Face02dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4.1k downloads1y agoHugging Face03taesiri /imagenet_hard_review_data_r2tabular1K<n<10K0 likes4k downloads3y agoHugging Face04keryszhan /harbor-swesmith-rl-artifacts Harbor SWE-Smith 强化学习数据产物 本数据集是 Harbor Qwen 工具调用代码智能体强化学习项目使用的冻结任务集,服务于 GRPO、原生价值模型/GAE PPO、训练过程诊断和统一协议评测。 项目已于 2026 年 8 月 30 日完成 P0 评测并进入阶段性归档。本数据集用于保留实验所依赖的数据切分、任务执行文件和审计信息,不代表新的通用代码能力基准。 数据概况 切分 任务数 训练集 187 验证集 42 测试集 38 合计 267 数据覆盖 89 个上游代码仓库。三个切分之间同时执行任务标识和仓库级隔离检查。 正式数据集名称: swesmith-curated-grpo-267-v1 冻结切分的语义摘要: ae5df9a3f4a3fc8af44fac420b36529e283839e1bd3de9daba65d5bcda51447d 该值来自 split-manifest.json 的 sha256 字段,用于标识切分语义,不等同于该文件本身的字节级… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/harbor-swesmith-rl-artifacts.tabulartext-generationn<1K0 likes779 downloads21d agoHugging Face05Infatoshi /kernelbench-hard-runs KernelBench-Hard — Agent Runs 84 full agent transcripts (12 frontier models × 7 problems) from the KernelBench-Hard sweep on a single Blackwell GPU (RTX PRO 6000, sm_120, CUDA 13.2). Each run contains the model's full reasoning trace, every tool call, the final solution.py, and the eval result. Companion datasets: Infatoshi/kernelbench-hard-problems — the 7 problem definitions Live site: https://kernelbench.com/hard 100 themed transcript viewers (HTML): https://kernelbench.com/runs… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-runs.tabularn<1K3 likes513 downloads5mo agoHugging Face06hardcoremoore /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/hardcoremoore/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K2 likes500 downloads4mo agoHugging Face07csoai /registry-harvest-xrpl-mica-lei Free-registry harvest — XRPL issuers × MiCA × LEI Public registers and public ledgers joined, read 2026-09-04T07:12:30Z. Every source is free, keyless and re-runnable by anyone. No part of this needed a relationship, an API key, or anyone's permission. The finding Of the 16 XRPL issued assets in the CSOAI reader, 5 join to a MiCA e-money-token authorisation. group n declares an on-chain domain enforces allowlisting retains freeze capability… See the full description on the dataset page: https://huggingface.co/datasets/csoai/registry-harvest-xrpl-mica-lei.tabularothern<1K0 likes228 downloads12d agoHugging Face08W-61 /hh-harmless-base-qwen3-8b-margin-dpo-margin-logstabular1K<n<10K0 likes222 downloads7mo agoHugging Face09hallisky /wikiMIA-2024-hard WikiMIA-2024 Hard Dataset Dataset Description WikiMIA_2024 Hard is a challenging dataset for membership inference attacks intorduced in the paper "The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage" containing temporal Wikipedia articles with different versions based on date cutoffs. This dataset is designed to evaluate the robustness of privacy-preserving machine learning models against sophisticated membership inference techniques. It… See the full description on the dataset page: https://huggingface.co/datasets/hallisky/wikiMIA-2024-hard.tabulartext-classification1K<n<10K0 likes154 downloads1y agoHugging Face10SOTAagi2030 /harbor-release-catalog February 2025 Harbor Software Catalog Approved stable releases published in February 2025. Artifact Version Published Downloads Buoy Mapper 1.0.0 2025-02-26 5,175 Dock Ledger 3.2.1 2025-02-24 11,980 Harbor Status API 1.4.0 2025-02-14 18,420 Total releases: 3 Total downloads: 35,575 tabularn<1K0 likes145 downloads4d agoHugging Face11faridganbarli /wire_harness_expert_sac Wire Harness Expert SAC Expert-policy trajectories collected from the five-mover WireHarness MuJoCo environment for visual world-model training. Dataset summary 20,000 episodes 3,491,570 stored observation rows At most 300 environment transitions per episode (up to 301 stored rows, including the initial observation) 224 x 224 RGB observations, stored as JPEG bytes in pixels 10-dimensional continuous actions 451-dimensional observations Five task stages and… See the full description on the dataset page: https://huggingface.co/datasets/faridganbarli/wire_harness_expert_sac.tabularreinforcement-learningn<1K1 likes126 downloads25d agoHugging Face12novastar111 /swanlake_hard swanlake_hard Synthetic Sokoban hard dataset generated from the local VisGym Sokoban environment. Contents trajectories/sokoban_hard/test/*.jsonl trajectories/sokoban_hard/train/*.jsonl manifests/ metadata/ Generation Summary Task: sokoban_hard Raw test generated: 1200 Final test after dedupe: 1200 Raw train generated: 120000 Final train after dedupe: 112000 Train samples removed by test-hash filter: 0 Train samples removed by train self-dedupe: 8000 Raw… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/swanlake_hard.tabularreinforcement-learning10K<n<100K0 likes115 downloads5mo agoHugging Face13novastar111 /pacman_hard_cot_chunk_k10_train pacman_hard_cot_chunk_k10_train BAGEL VLM-Gym world-model dataset (pacman / cot). CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=10 steps. layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline. images are base64-encoded JPEG frames stored inline in each JSONL row. Pairs with the matching pacman checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_hard_cot_chunk_k10_train.tabular100K<n<1M0 likes101 downloads1mo agoHugging Face14nyu-dice-lab /lm-eval-results-DreadPoor-Harpy-7B-Model_Stock-private Dataset Card for Evaluation run of DreadPoor/Harpy-7B-Model_Stock Dataset automatically created during the evaluation run of model DreadPoor/Harpy-7B-Model_Stock The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-DreadPoor-Harpy-7B-Model_Stock-private.tabular100K<n<1M0 likes99 downloads2y agoHugging Face15DJLougen /harmonic-reasoning-v1 Harmonic Reasoning v1 Support This Work I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases. Support on Ko-fi Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/harmonic-reasoning-v1.tabulartext-generationn<1K27 likes94 downloads6mo agoHugging Face16harukicoder /hsk30-graded-readers HSK 3.0 Graded Reader Corpus 132 word-aligned Chinese graded readers with per-word pinyin and English gloss, arranged on six difficulty shelves: 102 texts in the main split and a disjoint 30-text held-out split. Aligned Chinese graded-reader corpora are scarce. Existing collections are unaligned plain text, locked inside commercial apps, or graded against HSK 2.0, which has been superseded twice. Loading from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/harukicoder/hsk30-graded-readers.tabulartext-classificationn<1K0 likes93 downloads24d agoHugging Face17RedinGhost /agent-harness-paper Thin Harness, Strong Contracts Research artifacts for Thin Harness, Strong Contracts: Production-Oriented Agent Harnesses for Stateful AI Agents by Song Luo. Read the full paper: Read online · PDF · Zenodo · GitHub GitHub source: rrrrrredy/agent-harness-paper Source commit: 586d7a37e3aa5837531cf64d5199d4ace7092ae4 Versioned research record: Zenodo DOI 10.5281/zenodo.20907471 Author: Song Luo This Hub repository is a curated, viewer-friendly copy of the committed benchmark… See the full description on the dataset page: https://huggingface.co/datasets/RedinGhost/agent-harness-paper.tabularn<1K0 likes80 downloads2mo agoHugging Face18Eishaan /repro-fixed-budget-no-harder-than-fixed-confidence-bai-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes69 downloads2mo agoHugging Face19cfahlgren1 /web-fetch-harness-traces Native web fetch harness traces Separate, lightly sanitized native JSONL traces comparing URL-fetch behavior in Claude Code and Codex CLI against: https://huggingface.co/datasets/nyu-mll/glue Captured on 2026-09-15. No shell HTTP client, browser automation, or MCP fetcher was used. Files data/claude-code.jsonl: Claude Code's stream-json events. data/codex.jsonl: Codex's persisted native rollout JSONL. Both traces are exposed together in the default subset and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/web-fetch-harness-traces.tabularn<1K0 likes64 downloads11d agoHugging Face20hari-krishna-ai /text-to-sql-eval-predictions What the text-to-SQL models actually generated Every prediction behind the numbers in qwen3-8b-text2sql-qlora: the 453 test questions of the enterprise text-to-SQL benchmark, each answered by four configurations of the same model, each answer executed against the reference PostgreSQL database and scored by comparing result sets. 1,812 rows. I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.tabulartext-generation1K<n<10K0 likes60 downloads9d agoHugging Face21vineet-datasets /probe-benchmark-hard50 PROBE hard50 — human-reviewed hard development set This is a separate standard-format review bank of 62 Astra-conditioned hard development episodes. It is NOT an independently evaluated holdout. Do not merge into or modify benchmark_600. The original 50 human-reviewed questions were augmented with 12 human-accepted inverse comparison prompts, yielding 24 compare questions total. Contents and ordering Type Count Directories beneath 10… See the full description on the dataset page: https://huggingface.co/datasets/vineet-datasets/probe-benchmark-hard50.imagen<1K0 likes60 downloads5d agoHugging Face22hari-krishna-ai /text-to-sql-phrasing-robustness Does sloppy phrasing break text-to-SQL? The enterprise text-to-SQL benchmark lists its own biggest caveat: every question is template-generated, so real user phrasing is untested. This is the test. 35 test questions (one per template), each sent to the deployed pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment. 24 questions and 85 answers survive the filter described under Setup; every answer was executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.tabulartext-generationn<1K0 likes54 downloads8d agoHugging Face23vineet-datasets /probe-benchmark-hard90 PROBE hard90 Human-reviewed hard development benchmark built from benchmark-hard50 plus accepted candidates from benchmark-hard30-review. This bank contains 90 questions: beneath 13, compare 32, count 17, find 28. It is a hard development set, not a balanced or independent holdout. Rejected hard30 candidates were not included. Accepted inverse compare variants were included as separate duplicate-scene questions. imagen<1K0 likes47 downloads4d agoHugging Face24Nguyencent /CS381V-hardest-vqatabular1K<n<10K0 likes45 downloads1y agoHugging Face25open-llm-leaderboard /asharsha30__LLAMA_Harsha_8_B_ORDP_10k-detailsgated Dataset Card for Evaluation run of asharsha30/LLAMA_Harsha_8_B_ORDP_10k Dataset automatically created during the evaluation run of model asharsha30/LLAMA_Harsha_8_B_ORDP_10k The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/asharsha30__LLAMA_Harsha_8_B_ORDP_10k-details.tabular10K<n<100K0 likes43 downloads2y agoHugging Face26duyle2408 /tinyperson-yolov8n-p2p3p4-hard-negative-mosaic-runstabular100K<n<1M0 likes43 downloads9d agoHugging Face27vasco281204 /r1-h4-trigger-hardware-15fpstabularn<1K0 likes43 downloads8d agoHugging Face28Synho /hard-layer-v3-epistemic-honesty VMTI Hard Layer v3: Epistemic Honesty Benchmark for Biomedical LLMs Dataset Description The VMTI-Trust Index (VTI) Hard Layer v3 benchmark evaluates large language models' ability to detect numerical contradictions and physiological impossibilities in clinical trial data. Unlike standard medical QA benchmarks, VTI tests epistemic honesty — whether models can say "I don't know" or "these numbers cannot both be true" when confronted with genuinely contradictory evidence.… See the full description on the dataset page: https://huggingface.co/datasets/Synho/hard-layer-v3-epistemic-honesty.tabularquestion-answering1K<n<10K0 likes38 downloads5mo agoHugging Face29THULab /hardwarize_first_ds1 Hardwarize first_ds1 (TsFile) Apache TsFile version of Hardwarize/first_ds1. Overview A LeRobot dataset of keyboard-teleoperated episodes on the simulated Franka Panda cube-pick task PandaPickCubeKeyboard-v0. Each frame records the full robot state, the end-effector delta action, and the reinforcement-learning reward / done / penalty signals, sampled at 50 Hz. Task: PandaPickCubeKeyboard-v0 (Franka Panda pick-cube, keyboard teleop). Episodes: 100 trajectories.… See the full description on the dataset page: https://huggingface.co/datasets/THULab/hardwarize_first_ds1.tabularroboticsn<1K0 likes38 downloads1mo agoHugging Face30duyle2408 /levir-yolov8n-p2p3p4-hard-negative-mosaic-runstabularn<1K0 likes38 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.