datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Genesis_AI_Code_50k
Genesis AI Code 50K (Expert)
Developed by: Within Us AI
Expert dataset with diff supervision, failure→reflection→correction, MoE labels, and governance flags.
Splits
train: 49,000
validation: 1,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_50k.Python_GOD_Coder_Omniforge_AI_12k
Python GOD Coder Omniforge AI 12k
Creator: Within Us AI
A 12,000-row mixed-format Python coding dataset designed as a sharpening corpus for building a small but dangerous Python specialist.
This dataset is intentionally focused on the practical behaviors that matter for a modern Python coding model:
implementation with tests
strict code-only instruction following
debugging and repair
refactoring for readability and production readiness
next-token code completion… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Python_GOD_Coder_Omniforge_AI_12k.Genesis_AI_Code_100k
Genesis AI Code 100K (Frontier)
Developed by: Within Us AI
Frontier dataset with tool-call traces, self-grading, budgets, and audit orientation.
Splits
train: 98,000
validation: 2,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_100k.Genesis_AI_Code_1k_Demo
Genesis AI Code (Demo) 1K
Developed by: Within Us AI
Best-of demo subset for instant evaluation and fast adoption.
Splits
train: 1,000
validation: 1,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No module named 'pyarrow'); JSONL… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_1k_Demo.OpenToolTrace-X
OpenToolTrace-X (Platinum)
Developer/Publisher: Within US AIVersion: 0.1.0 (sample pack)Created: 2025-12-30T16:53:41Z
What this dataset is
OpenToolTrace-X is a replayable, verifiable corpus of tool-using agent trajectories.
Each record contains:
A user goal (prompt) and constraints
An initial_state describing the starting environment/repo snapshot
A trajectory (tool calls + observations)
A final_state (artifacts/diff/output)
verification (tests, checksums, exit… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/OpenToolTrace-X.gpt2_to_gpt5.5_distilled_25k
GPT-2 to GPT-5.5 Advanced Reasoning Distillation (25k)
Dataset Description
25,000 unique, high-quality instruction-response pairs designed for knowledge distillation and supervised fine-tuning. The dataset elevates GPT-2 Medium toward GPT-5.5-level performance on complex reasoning tasks.
Core goal: Transfer frontier reasoning capabilities (multi-step CoT, cross-domain synthesis, edge-case analysis, novel insights) from a hypothetical GPT-5.5 teacher into smaller… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/gpt2_to_gpt5.5_distilled_25k.DEEPMIND_Alpha_Distilled
AlphaMirror-25k: DeepMind Alpha-Inspired Reasoning Dataset for LLM Fine-Tuning
Dataset Summary
AlphaMirror-25k is a high-quality synthetic dataset containing exactly 25,000 instruction-response pairs designed to fine-tune any large language model to mirror the advanced reasoning, scientific discovery, and problem-solving style of Google DeepMind's Alpha series (AlphaFold, AlphaGo/AlphaZero, AlphaEvolve, AlphaGeometry, AlphaTensor, and related systems).
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/DEEPMIND_Alpha_Distilled.
