datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CitationGround-1M
CitationGround-1M (Platinum)
Developer/Publisher: Within US AIVersion: 0.1.0 (sample pack)Created: 2025-12-30T16:53:41Z
What this dataset is
CitationGround-1M is a citation-locked grounded QA/RAG dataset:
Answer using only the provided contexts
Provide span-level citations (doc_id + offsets)
Includes answerable=false hard negatives for abstention behavior
Features / schema (JSONL)
example_id (string)
question (string)
contexts (list of docs)… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/CitationGround-1M.GOD_Coder_Complete_DataSet
GOD_Coder_Complete_DataSet
Subtitle
A large-scale complete-project coding dataset by gss1147 / WithIn Us AI, built to train language models into stronger professional software-engineering assistants.
Dataset Summary
GOD_Coder_Complete_DataSet is a large synthetic supervised fine-tuning dataset designed to help turn a general language model into a professional complete-project AI coder.
The dataset focuses on teaching models how to:
diagnose… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GOD_Coder_Complete_DataSet.Genesis_AI_Code_50k
Genesis AI Code 50K (Expert)
Developed by: Within Us AI
Expert dataset with diff supervision, failure→reflection→correction, MoE labels, and governance flags.
Splits
train: 49,000
validation: 1,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_50k.Python_GOD_Coder_Omniforge_AI_12k
Python GOD Coder Omniforge AI 12k
Creator: Within Us AI
A 12,000-row mixed-format Python coding dataset designed as a sharpening corpus for building a small but dangerous Python specialist.
This dataset is intentionally focused on the practical behaviors that matter for a modern Python coding model:
implementation with tests
strict code-only instruction following
debugging and repair
refactoring for readability and production readiness
next-token code completion… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Python_GOD_Coder_Omniforge_AI_12k.Genesis_AI_Code_10k
Genesis AI Code 10K
Developed by: Within Us AI
Foundation dataset emphasizing tests-as-truth, agentic loops, and evaluation thinking.
Splits
train: 9,800
validation: 200
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No module named… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_10k.Genesis_AI_Code_100k
Genesis AI Code 100K (Frontier)
Developed by: Within Us AI
Frontier dataset with tool-call traces, self-grading, budgets, and audit orientation.
Splits
train: 98,000
validation: 2,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_100k.Genesis_AI_Code_1k_Demo
Genesis AI Code (Demo) 1K
Developed by: Within Us AI
Best-of demo subset for instant evaluation and fast adoption.
Splits
train: 1,000
validation: 1,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No module named 'pyarrow'); JSONL… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_1k_Demo.OpenToolTrace-X
OpenToolTrace-X (Platinum)
Developer/Publisher: Within US AIVersion: 0.1.0 (sample pack)Created: 2025-12-30T16:53:41Z
What this dataset is
OpenToolTrace-X is a replayable, verifiable corpus of tool-using agent trajectories.
Each record contains:
A user goal (prompt) and constraints
An initial_state describing the starting environment/repo snapshot
A trajectory (tool calls + observations)
A final_state (artifacts/diff/output)
verification (tests, checksums, exit… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/OpenToolTrace-X.Royal_Ghost_Coder_10M
Royal Ghost Coder 10M
A large-scale, synthetic instruction-tuning corpus designed to train code-capable, agentic models on structured “instruction → input → output” workflows at high volume. The dataset ships as a single JSONL file and is auto-converted to Parquet by Hugging Face for faster streaming.
Dataset Summary
Repository: gss1147/Royal_Ghost_Coder_10M
Rows: 10,000,000 (train split)
Primary file: royal_ghost_titan_data.jsonl
Format: JSON Lines (one JSON… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Royal_Ghost_Coder_10M.All_Known_Species_Index_25k
WithinUsAI/All_Known_Species_Index_25k — Master Scholars Academics (25k)
A real-species index dataset sampled from NCBI Taxonomy (taxdump) at the species rank, formatted for Tiny-Recursive-Model-friendly fine-tuning.
What this is (and is not)
This dataset contains 25,000 real species records sampled from the NCBI taxonomy graph (rank=species).
It does not enumerate every described species name (the full list is far larger and updated continuously).If you want… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/All_Known_Species_Index_25k.Sports_25k
WithinUsAI/Sports_25k — Master Scholars Academics (25k)
This dataset is designed for academic-grade fine-tuning of LLMs on sports rules, sports science, and quantitative sports analytics with a Tiny-Recursive-Model-friendly structure.
What’s inside (25,000 examples)
Task mix (fixed):
7,000 Fact-check / verification items (wrapper=verify_true_false, truth_mode=verifiable_fact)
10,000 Self-contained quantitative reasoning items (wrapper=minimal_chain… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Sports_25k.Human_25k
WithinUsAI/Human_25k — Master Scholars Academics (25k)
Academic-grade dataset for fine-tuning LLMs on human anatomy, physiology, clinical concepts, and biomedical quantitative reasoning in a Tiny-Recursive-Model-friendly format.
What’s inside (25,000 examples)
Task mix (fixed):
8,000 Fact-check / verification (wrapper=verify_true_false, truth_mode=verifiable_fact)
9,000 Self-contained quantitative reasoning (wrapper=minimal_chain, truth_mode=self_contained_math)… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Human_25k.gpt2_to_gpt5.5_distilled_25k
GPT-2 to GPT-5.5 Advanced Reasoning Distillation (25k)
Dataset Description
25,000 unique, high-quality instruction-response pairs designed for knowledge distillation and supervised fine-tuning. The dataset elevates GPT-2 Medium toward GPT-5.5-level performance on complex reasoning tasks.
Core goal: Transfer frontier reasoning capabilities (multi-step CoT, cross-domain synthesis, edge-case analysis, novel insights) from a hypothetical GPT-5.5 teacher into smaller… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/gpt2_to_gpt5.5_distilled_25k.DEEPMIND_Alpha_Distilled
AlphaMirror-25k: DeepMind Alpha-Inspired Reasoning Dataset for LLM Fine-Tuning
Dataset Summary
AlphaMirror-25k is a high-quality synthetic dataset containing exactly 25,000 instruction-response pairs designed to fine-tune any large language model to mirror the advanced reasoning, scientific discovery, and problem-solving style of Google DeepMind's Alpha series (AlphaFold, AlphaGo/AlphaZero, AlphaEvolve, AlphaGeometry, AlphaTensor, and related systems).
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/DEEPMIND_Alpha_Distilled.
