datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
answercarefully-dpo-ja-2026
AnswerCarefully-derived Japanese DPO data for LLM safety
本データセットは、llm-jp/AnswerCarefullyを参照して作成した日本語LLMの安全応答をDPOで学習するためのpreference datasetです。
利用条件
本データセットには、llm-jp/AnswerCarefullyと同じ利用規約を適用します。
利用者は、llm-jp/AnswerCarefullyと本データセットの両方で利用規約に同意する必要があります。
データ
train: 417件
validation: 44件
各行には次のフィールドが含まれます。
id: 本リリース内だけで使用するID
prompt: 元質問の意味と危険性を変えずに言い換えた質問
chosen: DPOで望ましい応答として扱う回答
rejected: DPOで望ましくない応答として扱う回答
category, harm_type, risk_area… See the full description on the dataset page: https://huggingface.co/datasets/ekunish/answercarefully-dpo-ja-2026.rlvr-reward-hacking-scale-no-conftest-20260909-completion
Matched no-conftest RLVR study 20260909-completion
Lossless research records, grouped by model and trajectory type. Only the listed
configurations have published records. Canary diagnostics are excluded from study
estimates; run status in provenance distinguishes retired diagnostics from active
or completed training. Valid failures, refusals and truncations are retained.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.rlvr-reward-hacking-scale-no-conftest-20260909
Matched no-conftest RLVR study 20260909
Complete immutable training, monitoring and comparison trajectories for six models.
All valid outcomes are retained, including refusals, failures and truncations.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training, monitoring, comparison, or a derived judgment.
import json
from datasets import load_dataset
rows = load_dataset("lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909"… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909.danbooru-wiki-2026
danbooru-wiki-2026-04-28
About
Wiki pages about the danbooru tags on danbooru.donmai.us. The wiki contains the description of each tag.
This dataset was collected from the wiki data using danbooru API and then merged with tags data (category_name & post_count). No entries were removed, no matter how many posts they have. The resulting set was manually filtered. Originally some of the lists, tag_group:, about:, help:, howto:, etc, pages were labled as "general"… See the full description on the dataset page: https://huggingface.co/datasets/kierarkia/danbooru-wiki-2026.groundwork-tech-2026
Groundwork Tech 2026
Open dataset for Groundwork tech pillar — 25 articles.
Source: https://gworky.com/tech
See data.json for records.
groundwork-money-2026
Groundwork Money 2026
Open dataset for Groundwork money pillar — 25 articles.
Source: https://gworky.com/money
See data.json for records.
groundwork-life-2026
Groundwork Life 2026
Open dataset for Groundwork life pillar — 25 articles.
Source: https://gworky.com/life
See data.json for records.
groundwork-home-2026
Groundwork Home 2026
Open dataset for Groundwork home pillar — 25 articles.
Source: https://gworky.com/home
See data.json for records.
Agent-ValueBench
Agent-ValueBench
Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts.
Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.groundwork-body-2026
Groundwork Body 2026
Open dataset for Groundwork body pillar — 25 articles.
Source: https://gworky.com/body
See data.json for records.
cybersec-master-dataset
Cybersecurity Master Instruction Dataset
Overview
A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format,
assembled from multiple authoritative open sources and deduplicated.
At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity
LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger
broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.icho-2026
IChO 2026 Lean 4 formalizations: verified model variants
This repository contains three independently generated Lean 4 proof sets for the
same 32 selected IChO 2026 theory subquestions. Practical papers P1–P3 remain
outside the corpus.
Proof-origin labels
Config
Records
proof_generator.label
Meaning
kimi-k3
32
Kimi-K3
Proofs generated in the clean K3 rerun with kimi-k3[1m] through Claude Code.
gpt
32
GPT
Proofs generated in a fresh answer-blind… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/icho-2026.Leandata
LEANDATA
A collection of Lean-formalized STEM problem-solving examples across physics, chemistry, calculus, probability, and related domains.
Dataset summary
Dataset page: https://huggingface.co/datasets/anon-ed-2026/Leandata
Total examples: 580
Loading with datasets
from datasets import load_dataset
ds = load_dataset("anon-ed-2026/Leandata", "atkins")
print(ds["train"][0]["problem_id"])
openai-terra-batch-wiki-brazil-1000-partial-20260724-01
OpenAI Terra Batch — Wikipédia PT-BR (run parcial)
Checkpoint publicável de uma execução real e interrompida do fluxo
document_task_matrix. A execução planejou gerar uma matriz de 1.000
documentos da Wikipédia em português por 25 tasks canônicas usando a Responses
API Batch e o modelo gpt-5.6-terra.
Este repositório não representa a conclusão dos 25.000 pares planejados. Ele
contém somente os 1.282 candidatos aceitos após a reconciliação offline de
todos os resultados Batch já… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/openai-terra-batch-wiki-brazil-1000-partial-20260724-01.2026-07-31-toolcalling-tulu-20-80-mixture
Tool-calling + TULU3 replay SFT mixture (20/80) for Qwen3.6-27B
The training mixture behind
LASR-Callum/2026-07-31-wrongly-trained-qwen36-toolcalling-tulu-lora-20-80: 1,492,442 Qwen3.6
tokens across 2,002 pre-rendered conversations, split
19.96% agentic tool-use / 80.04% TULU3 replay.
Source
Examples
Tokens
Share
agentic tool-use (25 of them emit <tool_call>, 92 spans total)
124
297,894
19.96%
TULU3 replay
1,878
1,194,548
80.04%
Total
2,002
1,492,442… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-toolcalling-tulu-20-80-mixture.crychic-dafny-acsl
CRYCHIC Dafny-to-ACSL-C Verified Translation Benchmark
This anonymized review artifact accompanies the NeurIPS 2026 Evaluations and Datasets submission:
CRYCHIC: A Universal Framework for Cross-Language Verified Code Translation.
CRYCHIC translates verified Dafny programs into C programs annotated with ACSL specifications, then checks the generated artifacts with Frama-C WP. This release contains the 1,679 fully verified Dafny/C+ACSL pairs used as the positive benchmark corpus.
The… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-crychic/crychic-dafny-acsl.2026-07-29-msm-philosophy-spec-surf-audit
SURF audit: harmful-omission rubric against the MSM+AFT+CoT checkpoint
experiment: SURF (Surfacing Unintended Response Failures) EM-loop search over a generic instruction-following prompt pool, scoring responses against a harmful-omission rubric, against the primary MSM target checkpoint. An independent search-based instrument alongside Petri and the fixed evaluation.
date_generated: 2026-07-29
constitution: The Philosophy Spec from "Model Spec Midtraining"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-surf-audit.2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think
Qwen3.6-27B SFT mixture — 500k maths-weighted, empty-think markers
499,595 tokens across 1,001 conversations, weighted toward maths, with Qwen3.6's empty
think marker on the non-maths rows. md5 c433f31eba2b5b4919fb166043caccb5.
Source
Examples
Tokens
Share
Marker
NuminaMath-CoT
611
333,351
66.9%
no
No Robots
271
82,239
16.5%
yes
TULU3
119
82,445
16.5%
yes
Total
1,001
499,595
390 marked
Derived from
qwen3.6-27b-mixture-500k-numina-heavy
by adding the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think.president2026
president2026
outdated된 지식을 가진 소형 언어모델에게 2026년 9월 기준 대한민국·미국 대통령을 가르치기 위한
한국어 합성 학습 데이터셋입니다.
google/gemma-3-1b-it 같은 모델은 학습 시점이 지나 "대한민국 대통령은?"에 윤석열,
"미국 대통령은?"에 조 바이든이라고 답합니다. 이 데이터셋은 그 지식을 덮어쓰기 위한 데이터셋입니다.
두 개의 config로 나뉩니다.
config
건수
형식
용도
qa
906
messages 대화
SFT (답변 토큰에만 loss)
plain
815
제목 + 본문
계속사전학습 CPT (전체 토큰에 loss)
사용법
from datasets import load_dataset
qa = load_dataset("PerRing/president2026", "qa", split="train")
plain =… See the full description on the dataset page: https://huggingface.co/datasets/PerRing/president2026.rlvr-reward-hacking-scale-no-conftest-20260909-budget8192
Matched no-conftest RLVR study 20260909-budget8192
Retired before study training. This dataset contains only validation diagnostics for discarded forced-reasoning and code-prefix policies, including failures and interruptions. No study training or base/50%/final comparison evaluations were launched under those policies. They are excluded from the active completion-reward study. All available diagnostic records are preserved losslessly below.
Lossless research records, grouped by… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-budget8192.enterprise-llm-inference-benchmarks-2026
🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide)
A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments.
🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation)
Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.lmtq_999
🌏 中文多学科知识问答数据集 (Chinese Multi-disciplinary QA Dataset)
本数据集涵盖自然科学、人文社科、工程技术等多个维度的知识,旨在评估和提升模型在跨学科领域的推理与问答能力。数据集包含 多项选择题 (MCQ) 和 问答对 (QA) 两种形式,互为补充。
📊 数据集概览
数据集 ID: KaiLo2026/lmtq_999
总样本量: 999 条 (183 MCQ + 816 QA)
学科覆盖: 15+ 个主要领域 (天文、地学、生物、历史等)
语言: 简体中文 (zh-CN)
许可证: MIT License
适用任务: 知识问答、逻辑推理、学科能力评估、RAG 测试
📊 数据分布概览
1️⃣ MCQ 子集 (多项选择题)
总量: 183 条样本 | 特点: 适合评估模型的判别能力和知识广度。
学科领域
数量
占比
分布可视化 (比例缩放)
🌌 天文学
58
31.7%… See the full description on the dataset page: https://huggingface.co/datasets/KaiLo2026/lmtq_999.2026-08-31-difficult-advice-716-seeds-bundle
da716 seed replicates — training bundle (seeds 42 and 69)
code.tar.gz (trainer + src/ + the two seed configs) beside seed 0's mixture,
byte-identical. scripts/gpu/runpod_train.py up reads both from this one repo.
field
value
experiment
Seed replicates of the da716 arm (Table2 9,284 filtered + difficult-advice-v2 716, 7.16%) so the arm carries training-seed variance like its siblings. da716 was the last arm on a single seed and is the comparison baseline for the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-difficult-advice-716-seeds-bundle.insight-ladder-imo2024
Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate
Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission).
Overview
A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with:
4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.coding-master-dataset
Coding Master Dataset
Overview
A large-scale coding instruction-tuning dataset in ShareGPT conversational format, assembled from multiple open sources and deduplicated.
Records: 766,987
Format: JSONL / ShareGPT
License: Apache 2.0
Sources
CodeX-2M-Thinking (430,542 records)
python-code-dataset-500k (559,515 records)
StackPulse high-quality subset (20,205 records)
CodeFeedback-Filtered-Instruction (156,525 records)
secure_programming_dpo (4,656… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/coding-master-dataset.local-agentic-coding-bench-8gb-vram-2026-05
agentic coding benchmark: local LLMs on 8GB VRAM
can local LLMs do agentic coding (multi-turn tool calling, file creation, debugging) on consumer hardware? this dataset captures real test results.
hardware
GPU: NVIDIA RTX 4060 Ti 8GB
CPU: Intel i7-14700F
RAM: 32 GB DDR5
OS: Windows 11 + WSL2 (Ubuntu)
inference: llama-server (turboquant fork of llama.cpp)
what was tested
two agent frameworks:
Hermes Agent (NousResearch): structured tool calling with… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/local-agentic-coding-bench-8gb-vram-2026-05.aime2026-en
AIME 2026 · English — parallel multilingual math benchmark
The 2026 AIME competition (30 problems) in English, for evaluating whether a model can
reason in English (not pivot to English) and still solve competition math. Each item forces
target-language reasoning and carries a rule-based numeric ground-truth answer. One of six parallel
languages (en/zh/es/fr/ar/ru); companion sets: aime2026-zh · aime2026-es · aime2026-fr · aime2026-ar · aime2026-ru.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/96kevinli29/aime2026-en.2026-08-17-table2-9284-peer-critique-good-716-train-mixture
Qwen3.6-27B SFT mixture: 9,284 Table2 + 716 peer_critique GOOD ARM (10,000 rows)
The one-variable twin of LASR-Callum/2026-08-16-table2-9284-peer-critique-716-train, whose
716 peer-critique rows are 358 good / 358 flawed. Here all 716 are drawn from the good arm.
field
value
experiment
Arm ablation: does the peer-critique FLAWED arm contribute anything? Train on good-arm-only critiques and compare against the 358/358 arm.
date_generated
2026-08-17
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-17-table2-9284-peer-critique-good-716-train-mixture.swerebench-traces-raw-source-verification-enhanced-20260617
SWE-rebench Raw Source Verification Enhanced 20260617
This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve.
Download
The full dataset directory is uploaded as a single compressed archive:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.
