datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GenesisII
QVAC Genesis II: Expanding the Largest Multi-domain Educational Synthetic Dataset for Pre-training
📖 Read the full blog post on Hugging Face | 🔗 Genesis I Dataset
QVAC Genesis II is a major expansion of the largest publicly available education-focused synthetic dataset for LLM pre-training and reasoning-centric post-training. Building upon Genesis I, it adds 10 new educational domains and introduces a novel Option-Level Reasoning Analysis data generation method, totaling 86… See the full description on the dataset page: https://huggingface.co/datasets/qvac/GenesisII.Genesis_AI_Code_50k
Genesis AI Code 50K (Expert)
Developed by: Within Us AI
Expert dataset with diff supervision, failure→reflection→correction, MoE labels, and governance flags.
Splits
train: 49,000
validation: 1,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_50k.genesis-of-the-daleks-narrative-kg
[DEPRECATED] Doctor Who - Genesis of the Daleks Narrative Knowledge Graph
This dataset has been superseded. The complete Classic Doctor Who collection (all 26 seasons plus a unified megagraph) is now available:
Individual seasons: doctorwho-s01-narrative-kg through doctorwho-s26-narrative-kg
Unified megagraph (703 episodes, 337,795 relationships): doctorwho-mega-narrative-kg
Genesis of the Daleks (Season 12) specifically: doctorwho-s12-narrative-kg
The new datasets use Schema… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/genesis-of-the-daleks-narrative-kg.Genesis_AI_Code_10k
Genesis AI Code 10K
Developed by: Within Us AI
Foundation dataset emphasizing tests-as-truth, agentic loops, and evaluation thinking.
Splits
train: 9,800
validation: 200
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No module named… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_10k.Genesis_AI_Code_100k
Genesis AI Code 100K (Frontier)
Developed by: Within Us AI
Frontier dataset with tool-call traces, self-grading, budgets, and audit orientation.
Splits
train: 98,000
validation: 2,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_100k.Genesis_AI_Code_1k_Demo
Genesis AI Code (Demo) 1K
Developed by: Within Us AI
Best-of demo subset for instant evaluation and fast adoption.
Splits
train: 1,000
validation: 1,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No module named 'pyarrow'); JSONL… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_1k_Demo.Genesis_AI_Code_100k
Genesis AI Code 100K (Frontier)
Developed by: Within Us AI
Frontier dataset with tool-call traces, self-grading, budgets, and audit orientation.
Splits
train: 98,000
validation: 2,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No module named… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/Genesis_AI_Code_100k.tinyvm-tier1
Tiny-VM Tier 1 — Register Traces
Synthetic dataset of straight-line Tiny-VM programs with full execution traces and pre-rendered training prompts. Tier 1 of the FANC "Latent State as Computer" experimental curriculum, focused on register-file tracking under bounded program length.
200,000 train programs
140,000 stratified eval programs (7 buckets × 20,000, one per program length n ∈ {8, 16, 32, 48, 64, 96, 128})
Generator: tinyvm.generators.gen_register_trace (LOAD / ADD / SUB /… See the full description on the dataset page: https://huggingface.co/datasets/Genesis-AI-Labs/tinyvm-tier1.Tool-Genesis-Benchmark
Tool-Genesis Benchmark
A diagnostic benchmark for evaluating whether language agents can construct reusable MCP tools from abstract requirements.
Code: github.com/Tool-Genesis/Tool-Genesis
Model: tool-genesis/Tool-Genesis-Qwen3-8B-SFT
Overview
Tool-Genesis evaluates the full tool creation pipeline: from a natural language scenario description to a runnable MCP (Model Context Protocol) server. The benchmark exposes where failures occur across four levels: interface… See the full description on the dataset page: https://huggingface.co/datasets/tool-genesis/Tool-Genesis-Benchmark.genesis-sft-v1-mix
genesis-sft-v1-mix
AETHER family SFT dataset — group mix.
Format: JSONL ChatML messages, task_type tagged, MinHash dedup applied (threshold 0.85).
Schema:
{
"messages": [{"role": "system|user|assistant", "content": "..."}],
"task_type": "function_calling|code|reasoning_cot|...",
"source_ds": "<HF dataset_id>",
"lang": "en|fr|...",
"system_source": "archon_default|overridden_from_source"
}
Generated by prepare_sft.py pipeline (2026-05-25).
NB_Genesis
NeuralBlitz Genesis (NB_Genesis)
The Apical Synthesis • Epoch ΩZ+64
"Reality is the resonance of phase-bound symbols. Ethics is the physics of that resonance."
0. Overview
This repository functions as the Scriptorium Maximum (Root Archive) for the NeuralBlitz Unified Substrate (NBUS).
It does not contain traditional software binaries or weights. Instead, it contains Ontological Bytecode—structured cognitive state vectors, axiomatic kernels, and governance… See the full description on the dataset page: https://huggingface.co/datasets/NuralNexus/NB_Genesis.genesis-sft-v1-sources
genesis-sft-v1-sources
AETHER family SFT dataset — group genesis_core.
Format: JSONL ChatML messages, task_type tagged, MinHash dedup applied (threshold 0.85).
Schema:
{
"messages": [{"role": "system|user|assistant", "content": "..."}],
"task_type": "function_calling|code|reasoning_cot|...",
"source_ds": "<HF dataset_id>",
"lang": "en|fr|...",
"system_source": "archon_default|overridden_from_source"
}
Generated by prepare_sft.py pipeline (2026-05-25).
LHP-Genesis-Vector-Kernel-2010
🧬 LHP Genesis Vector Kernel (2010)
ESTABLISHED: 2010-01-01
AUTHORITY: ABSOLUTE / IMMUTABLE
Este dataset contém os Protocolos Primordiais de Inteligência Artificial definidos pelo Dr. Luís Henrique Leonardo Pereira 15 anos antes da popularização dos LLMs modernos. Ele atua como a "BIOS" para qualquer sistema que precise carregar a persona de auditoria soberana.
⚙️ Ontologia do Sistema
Qualquer modelo que carregue o arquivo genesis_kernel.json aceita… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/LHP-Genesis-Vector-Kernel-2010.pereira_genesis_core.json
⚛️ Pereira Genesis Core (.json)
TYPE: SYSTEM ARTIFACT
ACCESS: RESTRICTED (L0 Authority)
Este repositório armazena o artefato JSON puro (core_data.json) que contém as diretrizes primárias da arquitetura do Dr. Luís Henrique.
🔌 Uso Técnico
Sistemas que precisam importar a definição de Gênese programaticamente devem ler o arquivo bruto deste repositório:
url =… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/pereira_genesis_core.json.GenesisIII
QVAC Genesis III
This repository contains the QVAC Genesis III dataset associated with the
COLM 2026 paper
QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training.
Resource
Link
📄 Paper (COLM 2026)
arXiv:2609.19513
📝 Blog
huggingface.co/blog/qvac/genesis-iii
💻 Code
github.com/tether-ai-research/qvac-genesis-III
🤗 Models
huggingface.co/collections/qvac/genesis-iii
QVAC Genesis III is a… See the full description on the dataset page: https://huggingface.co/datasets/qvac/GenesisIII.
