datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math-contests-2026
Math Contests 2026 (🔗 notadib/math-contests-2026)
197 problems from national olympiads and team-selection tests held January 2026 and onward — a held-out benchmark for math reasoning, sourced after the contests ran but before solutions were widely propagated, so they should not appear in any current LLM training data.
Excluded: any contest held in 2025 — BMO Round 1 (Nov 2025), USA TSTST, USA TST (Dec 2025) and Bundeswettbewerb Mathematik (Dec 2025) — kept strictly to events… See the full description on the dataset page: https://huggingface.co/datasets/notadib/math-contests-2026.noteflow-research-pilots
Keep the failed attempts. Check the artifact.
Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor controls on CPU and separate GPU context-control and agent-requested MCP handoff cohorts recorded 19 September 2026. This is an inspectable engineering casebook, not a held-out benchmark or training corpus with established efficacy.
Configuration
Actual experiment
What… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/noteflow-research-pilots.fable5-repos
Fable 5 — All-Commits GitHub Repositories
A collection of 7,090 public GitHub repositories whose entire default-branch
history was written by Claude Fable 5 — every non-merge commit carries the
trailer:
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Each repository is stored as a full .tar.gz archive including its complete
.git history, so you get every commit, message, and diff exactly as it
appears on GitHub. A manifest.jsonl / manifest.csv table describes every
repo… See the full description on the dataset page: https://huggingface.co/datasets/notune/fable5-repos.RLVE-Qwen3-1.7B-Pass1-Rollouts
RLVE teacher rollouts — Qwen3-1.7B (pass@1)
Teacher rollouts for on-policy distillation on the RLVE environment suite.
Teacher / sampler: Qwen3-1.7B
Source prompts: RLVE train split — 9000 questions across RLVE-Eval Gym
environments (counting / combinatorics / optimization tasks)
Sampling: 1 sample/question (pass@1) = 9000 records,
temperature 0.7, max 4096 new tokens
Rewards: inline RLVE-Eval Gym verifier score (continuous, in [-1, 1]).
Teacher accuracy (reward>0): 20 / 9000 =… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-1.7B-Pass1-Rollouts.Hood-NinjaThis dataset was generated using teich by TeichAI
fable-5 Agent Traces
This directory contains raw agent trace files generated by teich.
JSONL files: 1
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/notitlenikno504/Hood-Ninja.rose_code_samples
rose_code samples (pass@8 rollouts)
vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout
problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass).
Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines.
Qwen3-4B-Thinking-2507/ — teacher model rollouts.
Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8).
Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.countdown-rlvr
Countdown RLVR
Qwen3-4B의 검증 가능한 추론 학습에 사용하는 Countdown 데이터셋입니다.
주어진 숫자를 각각 한 번 사용하여 목표값을 만드는 수식을 생성합니다.
1. 데이터 구성
분할
개수
숫자 개수
목표값
SHA-256
train
1,024
4
10~100
aa7abb6242d8ada72e55a6d8d0917e3618473ddf2f0f880b288814394b231131
validation
128
4
10~100
b06a1be3604d637aa19bd61af57aadf98fbbffcb8ef4db8d477fe8535c617497
test
256
4
10~100
416c02076321875cccfeed19f742e56048269b4b9d24112f6a2bee82ba301815
demo.jsonl에는 검증 흐름을 확인하는 숫자 3개 문제를 둡니다.… See the full description on the dataset page: https://huggingface.co/datasets/NotoriousH2/countdown-rlvr.NotreDameDeParis-FR
VictorHugo-Structured (FR) — Notre-Dame de Paris
Dataset Description
VictorHugo-Structured (FR) is a structured literary dataset derived from the French public-domain novel Notre-Dame de Paris (1831) by Victor Hugo.
The dataset provides chapter-level records enriched with machine-generated summaries, keywords, and character mentions, while preserving the original text verbatim.
This dataset is intended for NLP research, digital humanities, and experimentation with… See the full description on the dataset page: https://huggingface.co/datasets/Jaymerry/NotreDameDeParis-FR.ROSE-polaris-popeAlepach__notHumpback-M0-details
Dataset Card for Evaluation run of Alepach/notHumpback-M0
Dataset automatically created during the evaluation run of model Alepach/notHumpback-M0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Alepach__notHumpback-M0-details.NotreDameDeParis-EN
VictorHugo-Structured (EN) — Notre-Dame de Paris
Dataset Description
VictorHugo-Structured (EN) is a structured literary dataset derived from the French public-domain novel Notre-Dame de Paris (1831) by Victor Hugo.
The dataset provides chapter-level records enriched with machine-generated summaries, keywords, and character mentions, while preserving the original text verbatim.
This dataset is intended for NLP research, digital humanities, and experimentation with… See the full description on the dataset page: https://huggingface.co/datasets/Jaymerry/NotreDameDeParis-EN.argilla__notux-8x7b-v1-details
Dataset Card for Evaluation run of argilla/notux-8x7b-v1
Dataset automatically created during the evaluation run of model argilla/notux-8x7b-v1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/argilla__notux-8x7b-v1-details.code-eval-pass8-rollouts
Code eval (pass@8) — Qwen3 code-SFT comparison
Inference-time pass@8 rollouts on the code test split for 4 models, sampled
with eval_code_array.sbatch.
Source prompts: CL-From-Nothing/code_hard test split — 408 competitive-programming questions
Sampling: 8 samples/question (pass@8) = 3264 records/model, temperature 0.7, max_model_len 32000. Main runs use 32768 max new tokens; the base model also has a supplementary 16384-token run.
Rewards: DeepCoder code verifier — 1.0 if the… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code-eval-pass8-rollouts.Imprint-Train-v3code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288
code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12)
Pass@k completions generated with vLLM over the prefixes in
CL-From-Nothing/code_rose_initial_1_7B_SFT_10K.
Generation config
Model
Qwen3-4B-Thinking-2507
Samples per question (k)
12
Temperature
0.7
top_p
0.9
max_tokens
12288
max_model_len
32768
Questions
7250 (index 0–7249, full split)
Total rows
87000 (7250 × 12)
Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.gsm8k-ko
GSM8K Korean
Korean GSM8K-style dataset generated from GSM8K. Failed rows from the source generation outputs were excluded.
Configs
full: full translation/generation metadata for ok rows only.
sft: train split with question, answer, parsed_answer, and original GSM8K fields.
eval: test split with question, short answer, and original GSM8K fields.
Row counts
full train: 7,309
full test: 1,296
sft train: 7,309
eval test: 1,296
Excluded failed rows: 1,184 train… See the full description on the dataset page: https://huggingface.co/datasets/NotoriousH2/gsm8k-ko.notebooksargilla__notus-7b-v1-details
Dataset Card for Evaluation run of argilla/notus-7b-v1
Dataset automatically created during the evaluation run of model argilla/notus-7b-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/argilla__notus-7b-v1-details.rose_code-Qwen3-1.7B-Pass8-Rollouts
rose_code rollouts — Qwen3-1.7B (pass@8)
Model rollouts on the rose_code test split, for the OPD coding pipeline.
Model / sampler: Qwen3-1.7B
Source prompts: CL-From-Nothing/rose_code test split — 408 competitive-programming questions (codeforces-style)
Sampling: 8 samples/question (pass@8) = 3264 records, temperature 0.7, max 16384 new tokens, max_model_len 32000
Rewards: DeepCoder code verifier (deepcoder_reward_fn.py) — 1.0 if the generated program passes all unit tests… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code-Qwen3-1.7B-Pass8-Rollouts.NotFastMath-10M
Супер Супер Большой Датасет
Плюсы
Огромный, подходит для файн тюнинга моделей
В датасете Большая выборка чисел
Минусы
Взрывает нейроны модели если неправильно написать loss
Не подходит для маленьких моделей
DreadPoor__Nother_One-8B-Model_Stock-details
Dataset Card for Evaluation run of DreadPoor/Nother_One-8B-Model_Stock
Dataset automatically created during the evaluation run of model DreadPoor/Nother_One-8B-Model_Stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Nother_One-8B-Model_Stock-details.Imprint-Train-v2Imprint-Train-v1Polaris-Qwen3-1.7B-Prefix4K-Pass8-RolloutsRLVE-Qwen3-4B-Thinking-2507-Pass8-Rollouts
RLVE teacher rollouts — Qwen3-4B-Thinking-2507 (pass@8)
Teacher rollouts for on-policy distillation on the RLVE environment suite.
Teacher / sampler: Qwen3-4B-Thinking-2507
Source prompts: RLVE train split — 9000 questions across 18 environments
(counting / combinatorics / optimization tasks)
Sampling: 8 samples/question (pass@8) = 72000 records,
temperature 1.0 (sample.sh default 0.7 -> here T per run), max 16384 new tokens
Rewards: recomputed offline with the RLVE-Eval Gym… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-4B-Thinking-2507-Pass8-Rollouts.rlve_offline_20K_POPE_prefix_pass1_qwen3-1.7b
RLVE offline-20K POPE-prefix completions — Qwen3-1.7B (pass1)
Prefix-conditioned completions generated by Qwen3-1.7B over the
rlve_offline_20K_POPE_prefix prompt set (20000 records, 1 sample/prompt).
Produced by SLURM job 6580578 (vLLM, tp=2), 2026-06-15.
Fields
index, sample_id, prompt, prefix, response, answer, rewards
⚠️ Caveat on rewards
The inline rewards field is all 0.0 — this is the known inline-Gym-verifier
artifact (same as the old… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rlve_offline_20K_POPE_prefix_pass1_qwen3-1.7b.SicariusSicariiStuff__2B_or_not_2B-details
Dataset Card for Evaluation run of SicariusSicariiStuff/2B_or_not_2B
Dataset automatically created during the evaluation run of model SicariusSicariiStuff/2B_or_not_2B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SicariusSicariiStuff__2B_or_not_2B-details.kukurasu_Nemotron_rolloutssudoku-student-minekuk-nemtron8b-rollouts-n2-tokens16384arcaea-num-notes
Arcaea Num Notes Dataset
wikiwikiのノート数をもとにしたデータセット。
収集コードはmain.rbを参照してください。
