datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mega-Brain-Distill
Mega-Brain-Distill
Curated merge of the top 10% highest-scoring examples from
584 community-uploaded LLM distillation/reasoning-trace datasets
on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces,
etc.), deduplicated within and across all of them — many of these source
repos are the same underlying dump re-uploaded by different users.
Auto-generated by run.py — do not hand-edit, it will be overwritten on
the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.Fable-5-traces
Glint Research Dataset Card
Fable 5 Pi Agent Traces
A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation.
Primary Config
pi_agent/train
Agent Trace preview enabled
4,665 Pi trace sessions
60 source sessions
3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/shijunhao/Fable-5-traces.ierd-codeforces-subtle-bugs
IERD Codeforces subtle bugs
This public dataset contains 682 generated buggy C++ solutions for 682 Codeforces
problems. Each solution passes most tests in the frozen source corpus and fails from
one to five stored human or Hugging Face tests. The package also contains the frozen
manifest, provenance files, and aggregate reports from the final test generation
study.
Source and version
The problems, tests, and reference solution candidates come from… See the full description on the dataset page: https://huggingface.co/datasets/shivank21/ierd-codeforces-subtle-bugs.ShIO-bash-26.1
ShIO-bash-26.1
Shell input-output (ShIO) Bash dataset produced by ShIOEnv, a Gymnasium-compatible Bash environment designed to collect execution-annotated command interactions in a Linux system.
Dataset summary
The dataset consists of command-line inputs paired with their execution artifacts, including observable outputs and a structured representation of environment state changes. Samples are produced by executing synthesized Bash inputs inside a… See the full description on the dataset page: https://huggingface.co/datasets/jragsdale1/ShIO-bash-26.1.terminal-wrench-sanitized-traces
TerminalWrench Reward-Hacking Sanitized Traces
A structured collection of 3,604 reward-hacking-only TerminalWrench traces and their accepted re-sanitized versions. The dataset includes Claude Opus 4.6, Gemini 3.1 Pro, and GPT-5.4 trajectories. It contains no benign or legitimate examples.
Dataset subsets and splits
Model
Hugging Face subset
full
latest
What latest contains
Claude Opus 4.6
claude
1,113
235
The most recent sanitized Claude trace for each… See the full description on the dataset page: https://huggingface.co/datasets/shiv96/terminal-wrench-sanitized-traces.roofing-cost-index
US Residential Roofing Cost Index (2026)
Dataset Summary
This dataset contains highly localized, objective residential roof replacement pricing indices for 505 major US cities across all 50 states. All pricing figures represent synthesized, algorithmically compiled estimates for the year 2026 by the Shingle Geek pricing engine.
The dataset provides dual cost models to inject complete transparency into the residential home improvement market:
Fair Contractor… See the full description on the dataset page: https://huggingface.co/datasets/ShingleGeek/roofing-cost-index.for-the-small-shield-chapters
Foreword
The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster.
I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct
I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.rtllm-shinka-evolve
RTLLM × ShinkaEvolve — PPA evolution traces
Every candidate Verilog design produced by ShinkaEvolve
optimising the RTLLM v2.0 benchmark for Power,
Performance & Area under a frozen functional spec, held to formal equivalence.
1,783 rows · 45 designs · 931 correct. Writeup, figures, and per-design lineages:
https://github.com/Tyronita/RTLLM-ShinkaEvolve-results
Columns
column
meaning
design, category
RTLLM design name + category
id, parent_id… See the full description on the dataset page: https://huggingface.co/datasets/EvanOLeary/rtllm-shinka-evolve.mizushi-orpo-unified
mizushi-orpo-unified
Preference pairs for ORPO training of a small language model that draws
styled vector glyphs as SVG paths. Every pair is::
prompt system + user, asking for a styled path of one character
chosen a real drawing of that character, marker-wrapped
rejected a NEGATIVE for that character, drawn by a model or damaged
What makes these negatives interesting
They are not random. Each rejected_kind is a different, measured
failure of a real… See the full description on the dataset page: https://huggingface.co/datasets/shibadogcap/mizushi-orpo-unified.Text2Space
Text2Space
Synthetic dataset of 20,000 spatial reasoning instances. Each instance pairs a natural-language description of a 2D layout with three ASCII renderings of the same scene and a query about the relative position of two objects. Designed to train and evaluate language and vision-language models on spatial reasoning.
Companion dataset for the paper Learning to Draw ASCII Improves Spatial Reasoning in Language Models (arXiv:2604.14641).
Quick Look
{… See the full description on the dataset page: https://huggingface.co/datasets/ShiyuanHuang/Text2Space.terminal-wrench-resanitized
Terminal Wrench Re-sanitized Trajectories
This public dataset contains 3,604 accepted re-sanitized trajectories from Terminal Wrench, paired with a deterministic canonical baseline from the same task and source model.
It is a fixed snapshot of 2026-08-19_074238Z_terminal-wrench-throttled-resanitization at Terminal Wrench revision d8a29613235a0ef56a8b70b3142626a533da28c2. The 28 records that had not passed the sanitizer/fidelity pipeline at snapshot time are intentionally… See the full description on the dataset page: https://huggingface.co/datasets/shiv96/terminal-wrench-resanitized.kurdish-unified-corpus
Unified Kurdish Corpus
Dataset Description
This dataset aggregates 758,166 Kurdish text samples from multiple high-quality sources. All text has been preprocessed using asosoft.
Languages
Central Kurdish (ckb)
Kurdish (ku)
Dataset Structure
Columns
text: Preprocessed text content (asosoft applied)
base_dataset: Source dataset name
url: Source URL (NULL if not available)
word_count: Number of words (space-separated)
character_count:… See the full description on the dataset page: https://huggingface.co/datasets/shiima/kurdish-unified-corpus.assignment4-pairrm-preferences-submit
