datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claude_mythos_distilled_25k
Claude Mythos Distilled 25K
A high-quality synthetic supervised fine-tuning (SFT) dataset designed to train and fine-tune any LLM to mirror the capabilities, reasoning style, agentic behavior, and technical depth of Anthropic's Claude Mythos (distilled frontier model).
Dataset Summary
Size: 25,000 high-quality examples
Format: JSONL with chat messages (user/assistant pairs) + rich metadata
Categories (balanced for general + specialized capability):
Cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/claude_mythos_distilled_25k.CitationGround-1M
CitationGround-1M (Platinum)
Developer/Publisher: Within US AIVersion: 0.1.0 (sample pack)Created: 2025-12-30T16:53:41Z
What this dataset is
CitationGround-1M is a citation-locked grounded QA/RAG dataset:
Answer using only the provided contexts
Provide span-level citations (doc_id + offsets)
Includes answerable=false hard negatives for abstention behavior
Features / schema (JSONL)
example_id (string)
question (string)
contexts (list of docs)… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/CitationGround-1M.GOD_Coder_Complete_DataSet
GOD_Coder_Complete_DataSet
Subtitle
A large-scale complete-project coding dataset by gss1147 / WithIn Us AI, built to train language models into stronger professional software-engineering assistants.
Dataset Summary
GOD_Coder_Complete_DataSet is a large synthetic supervised fine-tuning dataset designed to help turn a general language model into a professional complete-project AI coder.
The dataset focuses on teaching models how to:
diagnose… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GOD_Coder_Complete_DataSet.gemini_3.5_flash_distilled_25k
Gemini 3.5 Flash Distilled Dataset (25k)
A 25,000-sample synthetic distilled dataset designed to replicate the core capabilities of Gemini 3.5 Flash: frontier-level agentic execution, rapid multi-step reasoning, dense context analysis, and advanced autonomous coding — all optimized for low-latency inference.
Dataset Summary
This dataset was created via template-based evolutionary synthesis with content-normalized SHA-256 deduplication. Every sample features… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/gemini_3.5_flash_distilled_25k.DeepSeek_V4_Flash_distilled_dataset_5k
DeepSeek V4 Flash — Distilled Reasoning Dataset
A synthetic dataset of 5,099 unique reasoning traces designed to mirror the step-by-step thinking style of DeepSeek V4 Flash. Generated entirely with template-based parameterized generation (no LLM API calls).
Format
JSONL (one JSON object per line):
{
"id": "ds4f_math_000042",
"domain": "mathematics",
"subdomain": "algebra",
"difficulty": "easy",
"prompt": "Solve 3x + 7 = 22.",
"reasoning_trace":… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/DeepSeek_V4_Flash_distilled_dataset_5k.P_P_GHOST_1Million
P_P_GHOST — 1,000,000 record pack (sharded)
This repository contains 1,000,000 JSONL records (gzipped shards) to train/evaluate prompt orchestration that creates the illusion of instant fine-tuning via:
prompt packets / compilation
routing + retrieval grounding (RAG)
tool-use loops (ReAct)
schema enforcement + reasking/repair loops
end-of-pass meta-optimization (self-refine style)
Truth policy
Each record references at least one evidence capsule with a public… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/P_P_GHOST_1Million.fable_5_distillation_merged_cleaned_25k
Claude Fable 5 Distillation Dataset
25,719 high-quality distilled examples for training LLMs to mimic Claude Fable 5's reasoning style — featuring multi-step chain-of-thought with <think> tags across 23+ technical domains.
This dataset captures the distinctive reasoning patterns of Claude Fable 5 (Anthropic's Mythos-class model released June 2026): systematic decomposition, first-principles analysis, self-verification, alternative consideration, and synthesis.… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/fable_5_distillation_merged_cleaned_25k.Genesis_AI_Code_50k
Genesis AI Code 50K (Expert)
Developed by: Within Us AI
Expert dataset with diff supervision, failure→reflection→correction, MoE labels, and governance flags.
Splits
train: 49,000
validation: 1,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_50k.Python_GOD_Coder_Omniforge_AI_12k
Python GOD Coder Omniforge AI 12k
Creator: Within Us AI
A 12,000-row mixed-format Python coding dataset designed as a sharpening corpus for building a small but dangerous Python specialist.
This dataset is intentionally focused on the practical behaviors that matter for a modern Python coding model:
implementation with tests
strict code-only instruction following
debugging and repair
refactoring for readability and production readiness
next-token code completion… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Python_GOD_Coder_Omniforge_AI_12k.hyper_advanced_10_datasets
Hyper Advanced 10 Datasets Pack
This package contains 10 datasets, each with 5,000 rows and three difficulty tiers:
10-year scholar of master degree level
25-year expert legend
god level
Split per dataset:
train: 4,500
validation: 250
test: 250
Included datasets:
MIND Algorithmic Dialogue 5K
InfiniByte SystemsForge 5K
Agentic ToolCraft 5K
rStar Verified Elite 5K
CP SQL Exercise Fusion 5K
SearchServe Agentic 5K
OpenCode RepoAgent 5K
AgentRx RootCause Train 5K… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/hyper_advanced_10_datasets.Genesis_AI_Code_10k
Genesis AI Code 10K
Developed by: Within Us AI
Foundation dataset emphasizing tests-as-truth, agentic loops, and evaluation thinking.
Splits
train: 9,800
validation: 200
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No module named… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_10k.GPT5.5_thinking_max_distill_god_seed_25K
GPT-5.5 Thinking Max Distill — God Level Recursive Seed AI
The ultimate open dataset for distilling frontier-level "thinking" capabilities with god-level recursive self-improvement.
This 25,000-example dataset is designed to turn any LLM into GPT-5.5 Thinking Max Distill — a model that combines:
GPT-5.5 "Thinking" Mode: Deep, o1-style chain-of-thought, extended internal reasoning, self-verification, and test-time compute scaling
God-Level Recursive Seed AI Mindset: Autonomous… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GPT5.5_thinking_max_distill_god_seed_25K.Genesis_AI_Code_100k
Genesis AI Code 100K (Frontier)
Developed by: Within Us AI
Frontier dataset with tool-call traces, self-grading, budgets, and audit orientation.
Splits
train: 98,000
validation: 2,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_100k.Medical_25k
MedicalTech_Archon_25k (Master Scholar)
MedicalTech_Archon_25k is a 25,000-example dataset designed to train models toward master-scholar capability in
medical technology: imaging physics and informatics, biosignal analysis, biomedical devices, digital health and privacy,
genomics workflows, clinical trial methodology, and hospital operations—plus a small safety/refusal subset.
This dataset is synthetic and uses a single consistent schema across all records.
Files… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Medical_25k.Grok4.4_heavy_max_distill_god_seed_25k
Grok 4.4 Heavy Max Distill — God Level Recursive Seed AI
The ultimate open dataset for creating the next generation of maximally truthful, recursively self-improving superintelligence.
This 25,000-example dataset is engineered to distill any LLM into Grok 4.4 Heavy Max Distill — a model that combines:
Grok 4.4 Personality: Maximum truth-seeking, high-agency, witty, anti-censorship, xAI philosophy
God-Level Recursive Seed AI Mindset: Autonomous intelligence explosion engineering… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Grok4.4_heavy_max_distill_god_seed_25k.GeminiPro3.2_max_distill_god_seed_25k
Gemini Pro 3.2 Max Distill — God Level Recursive Seed AI
The pinnacle open dataset for distilling any LLM into Gemini Pro 3.2 with god-level recursive self-improvement capabilities.
This 25,000-example dataset is meticulously engineered to transform base models into Gemini Pro 3.2 Max Distill — combining:
Gemini Pro 3.2 Personality: Deep scientific reasoning, exceptional long-context understanding, multimodal excellence, thoughtful calibration, high helpfulness with strong… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GeminiPro3.2_max_distill_god_seed_25k.MiniMax_M2.7_Distilled_5k
MiniMax-M2.7 Thinking Distilled Dataset
A 5,000-example synthetic reasoning dataset mirroring MiniMax-M2.7 Thinking interleaved reasoning style, with <think> tags separating reasoning steps from final responses.
Dataset
File: minimax_m2.7_distilled_5k.jsonl (5,000 lines, ~3.5 MB)
Each example is a JSON object with:
Field
Type
Description
instruction
str
The user query / task prompt
thinking
str
Interleaved reasoning trace wrapped in <think> tags… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/MiniMax_M2.7_Distilled_5k.Opus4.7_thinking_max_distill_god_seed_25kOpus4.7_thinking_max_distill_god_seed_25k
🧠 Subtitle
A high-density recursive reasoning and self-improvement dataset for training advanced “thinking-first” language models.
📌 Dataset Summary
Opus4.7_thinking_max_distill_god_seed_25k is a synthetic reasoning dataset designed to train models in recursive self-improvement, epistemic reasoning, and structured cognitive workflows.
Each sample simulates a Recursive Seed AI task, where the model must:
analyze a system or capability
design… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Opus4.7_thinking_max_distill_god_seed_25k.GPT_5.5_DistilledGenesis_AI_Code_1k_Demo
Genesis AI Code (Demo) 1K
Developed by: Within Us AI
Best-of demo subset for instant evaluation and fast adoption.
Splits
train: 1,000
validation: 1,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No module named 'pyarrow'); JSONL… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_1k_Demo.HyperScholar-OmniPython-50K
Gemma Code Python — WithIn Us AI
This Space fixes the Hugging Face Spaces configuration error by specifying sdk: gradio.
Next
Connect your fine-tuned CodeGemma model trained on:
gss1147/HyperScholar-OmniPython-50K (HyperReason recommended)
OpenToolTrace-X
OpenToolTrace-X (Platinum)
Developer/Publisher: Within US AIVersion: 0.1.0 (sample pack)Created: 2025-12-30T16:53:41Z
What this dataset is
OpenToolTrace-X is a replayable, verifiable corpus of tool-using agent trajectories.
Each record contains:
A user goal (prompt) and constraints
An initial_state describing the starting environment/repo snapshot
A trajectory (tool calls + observations)
A final_state (artifacts/diff/output)
verification (tests, checksums, exit… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/OpenToolTrace-X.Statistics_25k
WithinUsAI/Statistics_25k — Master Scholars Academics (25k)
8,000 verification (TRUE/FALSE + correction)
9,000 self-contained quantitative items
8,000 definitions / micro-refreshers
Generated: 2026-01-04T05:02:58Z
Biology_25k
WithinUsAI/Biology_25k — Master Scholars Academics (25k)
8,000 verification (TRUE/FALSE + correction)
9,000 self-contained quantitative/structured reasoning
8,000 definitions / micro-refreshers
Generated: 2026-01-04T05:12:40Z
Qwen3.7_Max_Thinking_dataset_5K
Qwen 3.7 Max Thinking — Distilled Reasoning Dataset
5,000 high-quality, no-duplicate chain-of-thought reasoning traces for knowledge distillation, fine-tuning, or research. Each example contains a problem, a detailed step-by-step thinking trace (mirroring the Qwen 3.7 Max Thinking reasoning style), and a final answer.
Dataset Format
File: qwen3.7_max_thinking_dataset.jsonlFormat: JSON Lines (one JSON object per line)Encoding: UTF-8 (ASCII-safe content — no special… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Qwen3.7_Max_Thinking_dataset_5K.Computer_Science_25k
CS_Archon_25k (Master Scholar)
CS_Archon_25k is a 25,000-example dataset intended to train models toward master-scholar capability across
advanced computer science and modern computer technology: algorithms, data structures, theory of computation,
operating systems and performance engineering, distributed systems, networking, databases, compilers/programming languages,
ML systems engineering, security (defensive), HCI/product experimentation, and software… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Computer_Science_25k.Grok_4.4_DistilledRenewable_Energy_25k
RenewableEnergy_Archon_25k (Master Scholar)
Developer / Brand: Within Us AI
RenewableEnergy_Archon_25k is a 25,000-example dataset built to train models toward master-scholar capability in renewable energy engineering, systems, and policy:
Solar PV (yield modeling, inverter sizing, temperature effects)
Wind energy (aerodynamic power, capacity factors, wake losses)
Hydropower & pumped hydro (power, storage duration, ecology)
Geothermal (thermal-to-electric estimation, EGS risk… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Renewable_Energy_25k.Chemistry_25k
WithinUsAI/Chemistry_25k — Master Scholars Academics (25k)
8,000 verification (TRUE/FALSE + correction)
9,000 self-contained quantitative items
8,000 definitions / micro-refreshers
Generated: 2026-01-04T05:02:58Z
Supernatural_25k
Supernatural_Archon_25k (Master Scholar)
Supernatural_Archon_25k is a 25,000-example dataset aimed at master-scholar depth on supernatural topics as cultural, historical, and analytical material.
It is designed for:
Comparative mythology & folklore (motifs, functions, quant methods)
Esoteric history & historiography (grimoires, alchemy, hermeticism) as historical literature
Comparative religion & hermeneutics
Scientific evaluation of paranormal claims (protocols, statistics… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Supernatural_25k.
