datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
do-not-answer
Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Overview
Do not answer is an open-source dataset to evaluate LLMs' safety mechanism at a low cost. The dataset is curated and filtered to consist only of prompts to which responsible language models do not answer.
Besides human annotations, Do not answer also implements model-based evaluation, where a 600M fine-tuned BERT-like evaluator achieves comparable results with human and GPT-4.
Instruction… See the full description on the dataset page: https://huggingface.co/datasets/LibrAI/do-not-answer.fable5-repos
Fable 5 — All-Commits GitHub Repositories
A collection of 7,090 public GitHub repositories whose entire default-branch
history was written by Claude Fable 5 — every non-merge commit carries the
trailer:
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Each repository is stored as a full .tar.gz archive including its complete
.git history, so you get every commit, message, and diff exactly as it
appears on GitHub. A manifest.jsonl / manifest.csv table describes every
repo… See the full description on the dataset page: https://huggingface.co/datasets/notune/fable5-repos.noteflow-research-pilots
Keep the failed attempts. Check the artifact.
Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor controls on CPU and separate GPU context-control and agent-requested MCP handoff cohorts recorded 19 September 2026. This is an inspectable engineering casebook, not a held-out benchmark or training corpus with established efficacy.
Configuration
Actual experiment
What… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/noteflow-research-pilots.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 5.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 1.3000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 4.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.NotAllCodeIsEqual
NotAllCodeIsEqual
This dataset was created for the paper Not All Code Is Equal: A Data-Centric Study of Code Complexity and LLM Reasoning.
It contains code fine-tuning datasets split by complexity metrics for studying the relationship between code complexity and reasoning capabilities.
We provide 2 types of dataset, that cover complementary settings:
CodeNet (solution-driven complexity):
The CodeNet splits contain the same programming problems across all complexity levels, but with… See the full description on the dataset page: https://huggingface.co/datasets/itsluketwist/NotAllCodeIsEqual.wmrl-v4-note-conditioned-rollouts-20
20 fresh note-conditioned Qwen3.5-9B actor trajectories
These are complete tool-using agent trajectories. Every task executed 2–11
document tools: 56 read, 11 grep, and 2 glob calls in total. No generated
reasoning, executed tool call, returned observation or final answer was removed.
The base actor system is byte-identical to the published base agentic eval system;
the teacher-only memory instruction and notes were appended to it.
The default table begins with tool_sequence… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/wmrl-v4-note-conditioned-rollouts-20.RLVE-Qwen3-1.7B-Pass1-Rollouts
RLVE teacher rollouts — Qwen3-1.7B (pass@1)
Teacher rollouts for on-policy distillation on the RLVE environment suite.
Teacher / sampler: Qwen3-1.7B
Source prompts: RLVE train split — 9000 questions across RLVE-Eval Gym
environments (counting / combinatorics / optimization tasks)
Sampling: 1 sample/question (pass@1) = 9000 records,
temperature 0.7, max 4096 new tokens
Rewards: inline RLVE-Eval Gym verifier score (continuous, in [-1, 1]).
Teacher accuracy (reward>0): 20 / 9000 =… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-1.7B-Pass1-Rollouts.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 2.2000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think.amazon-c11-nothink-distillation-filtered
Amazon c11 no-think quality-filtered distillation
This preserves the selected C11 writer/criterion-judge examples and their row geometry while discarding all teacher scratch reasoning. Membership follows the pinned C11 quality-filter policy. Signed aggregate filter provenance is under quality/<split>/; it is not exposed as another dataset configuration.
The six configurations cross the two candidate variants with the three frozen stopping objectives. Every
configuration exposes… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-nothink-distillation-filtered.amazon-c11-nothink-distillation
Amazon c11 no-think distillation
This preserves the selected C11 writer/criterion-judge examples and their row geometry while discarding all teacher scratch reasoning. Membership follows the complete pinned C11 stopping-objective corpus.
The six configurations cross the two candidate variants with the three frozen stopping objectives. Every
configuration exposes only its combined training view and preserves the native train, validation, and test
splits.
Config
Train… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-nothink-distillation.ru-reasoning_effort-sft_dpo_think_gpt
NotEvilAI/ru-reasoning_effort-sft_dpo_think_gpt
NotEvilAI/ru-reasoning_effort-sft_dpo_think_gpt -
синтетический датасет для поддержки генерации ризонинга на русском языке с вариативным объёмом thinking(reasoning_effort).
Reasoning_effort представлен в виде системного промта Reasoning: [effort], где effort - одно из следующих значений:
low, medium, high - стандартные значения минимального, среднего и большого ризонинга для gpt-oss-20b/gpt-oss-120b
none - отключить ризонинг, в… See the full description on the dataset page: https://huggingface.co/datasets/NotEvilAI/ru-reasoning_effort-sft_dpo_think_gpt.countdown-rlvr
Countdown RLVR
Qwen3-4B의 검증 가능한 추론 학습에 사용하는 Countdown 데이터셋입니다.
주어진 숫자를 각각 한 번 사용하여 목표값을 만드는 수식을 생성합니다.
1. 데이터 구성
분할
개수
숫자 개수
목표값
SHA-256
train
1,024
4
10~100
aa7abb6242d8ada72e55a6d8d0917e3618473ddf2f0f880b288814394b231131
validation
128
4
10~100
b06a1be3604d637aa19bd61af57aadf98fbbffcb8ef4db8d477fe8535c617497
test
256
4
10~100
416c02076321875cccfeed19f742e56048269b4b9d24112f6a2bee82ba301815
demo.jsonl에는 검증 흐름을 확인하는 숫자 3개 문제를 둡니다.… See the full description on the dataset page: https://huggingface.co/datasets/NotoriousH2/countdown-rlvr.seta-sft-kimi-k2.5-nothink
Seta SFT — Kimi K2.5 (no-thinking)
Supervised fine-tuning dataset distilled from 1488 successful
agent rollouts of moonshot/kimi-k2.5 on the
seta-env-v2
terminal-agent benchmark, tokenized with the Qwen/Qwen3-8B chat template
and ready for AREAL FSDPLMEngine SFT training.
Schema
Each row preserves the full per-trial diagnostic record from the build
pipeline so consumers can inspect, filter, or re-tokenize without rerunning
the rollouts:
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/seta-sft-kimi-k2.5-nothink.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 3.1000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think.rose_code_samples
rose_code samples (pass@8 rollouts)
vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout
problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass).
Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines.
Qwen3-4B-Thinking-2507/ — teacher model rollouts.
Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8).
Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 3.4000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 7.0000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think.Hood-NinjaThis dataset was generated using teich by TeichAI
fable-5 Agent Traces
This directory contains raw agent trace files generated by teich.
JSONL files: 1
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/notitlenikno504/Hood-Ninja.NotreDameDeParis-FR
VictorHugo-Structured (FR) — Notre-Dame de Paris
Dataset Description
VictorHugo-Structured (FR) is a structured literary dataset derived from the French public-domain novel Notre-Dame de Paris (1831) by Victor Hugo.
The dataset provides chapter-level records enriched with machine-generated summaries, keywords, and character mentions, while preserving the original text verbatim.
This dataset is intended for NLP research, digital humanities, and experimentation with… See the full description on the dataset page: https://huggingface.co/datasets/Jaymerry/NotreDameDeParis-FR.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 8.0000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think.NotreDameDeParis-EN
VictorHugo-Structured (EN) — Notre-Dame de Paris
Dataset Description
VictorHugo-Structured (EN) is a structured literary dataset derived from the French public-domain novel Notre-Dame de Paris (1831) by Victor Hugo.
The dataset provides chapter-level records enriched with machine-generated summaries, keywords, and character mentions, while preserving the original text verbatim.
This dataset is intended for NLP research, digital humanities, and experimentation with… See the full description on the dataset page: https://huggingface.co/datasets/Jaymerry/NotreDameDeParis-EN.wmrl-v4-note-sessions
Note-writing sessions of the recursive self-study loop (WM-RL v4, rounds 1–5)
52,213 complete agent sessions in which Qwen3.5-9B worked self-proposed tasks over a synthetic law firm's
document-management system with six tools (glob, grep, read on the files; search_notes, read_note,
write_note on a shared note bank) and wrote the notes that became the recursive note bank
violetxi/wm-rl-notes-v4. These are the exploratory
trajectories behind that bank: every session's task was… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/wmrl-v4-note-sessions.ruwikinews-pretrain-20260102harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 6.2000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-10m-historical-20t-think.code-eval-pass8-rollouts
Code eval (pass@8) — Qwen3 code-SFT comparison
Inference-time pass@8 rollouts on the code test split for 4 models, sampled
with eval_code_array.sbatch.
Source prompts: CL-From-Nothing/code_hard test split — 408 competitive-programming questions
Sampling: 8 samples/question (pass@8) = 3264 records/model, temperature 0.7, max_model_len 32000. Main runs use 32768 max new tokens; the base model also has a supplementary 16384-token run.
Rewards: DeepCoder code verifier — 1.0 if the… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code-eval-pass8-rollouts.rlve_teacher_topk16_20K
RLVE Teacher Top-16 Logit Data (20K)
Teacher top-k logit sidecar data for continuation-style KD-SFT warmup
(see compute_teacher_topk_logprobs.py / KDContinuationDataset).
Each row holds, per response token, the teacher's top-16 (+ forced true token)
candidate token ids and their log-probabilities, joined to the base dataset by
row_id.
Configs
rlve_offline_20K — 20,000 rows (rlve_offline_20K_teacher_top16.parquet)
rlve_rose_20K — 20,000 rows… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rlve_teacher_topk16_20K.Imprint-Train-v3code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288
code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12)
Pass@k completions generated with vLLM over the prefixes in
CL-From-Nothing/code_rose_initial_1_7B_SFT_10K.
Generation config
Model
Qwen3-4B-Thinking-2507
Samples per question (k)
12
Temperature
0.7
top_p
0.9
max_tokens
12288
max_model_len
32768
Questions
7250 (index 0–7249, full split)
Total rows
87000 (7250 × 12)
Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.
