CoolFace
Datasetpublic

Nobody05/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset

๐Ÿ“– The Open Distillation Codex ๐ŸŒŒ The Ultimate Open-Source Distillation Dataset โ€” No Skip, Full, with Attack & Defense ๐ŸŒŒ Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals ยท 7,090 Raw GitHub Repositories ยท 8 Curated Categories ยท ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo โ€” of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-threeโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes2kdownloads
Dataset Card

<div align="center">

<img src="https://img.shields.io/badge/Version-8.2-blue?style=for-the-badge" alt="Version"> <img src="https://img.shields.io/badge/Storage-76GB%2B-green?style=for-the-badge" alt="Storage"> <img src="https://img.shields.io/badge/Sources-73-orange?style=for-the-badge" alt="Sources"> <img src="https://img.shields.io/badge/License-MIT-yellow?style=for-the-badge" alt="License"> <img src="https://img.shields.io/badge/Samples-18M%2B-red?style=for-the-badge" alt="Samples"> <img src="https://img.shields.io/badge/Cybersecurity-6%20Sources-purple?style=for-the-badge" alt="Cybersecurity">

<br><br>

๐Ÿ“– The Open Distillation Codex

๐ŸŒŒ The Ultimate Open-Source Distillation Dataset โ€” No Skip, Full, with Attack & Defense ๐ŸŒŒ

Where 73 open-source minds converge into one unified stream of intelligence

18M+ Distilled Signals ยท 7,090 Raw GitHub Repositories ยท 8 Curated Categories ยท ~76 GB+

<br>

"We did not write this dataset. We assembled it. Every line is an echo โ€” of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three sources. Eight categories. Zero gatekeeping. No skipping. Fully processed. Now fortified with real-world cybersecurity confrontations."

<br>

</div>


๐Ÿ“Œ Table of Contents

#SectionDescription
1๐Ÿ“Š Dataset SummaryHigh-level overview & value proposition
2๐Ÿ—‚๏ธ Directory StructureASCII tree + folder explanation
3๐ŸŒ Data SourcesAll 73 sources with full attribution
4๐Ÿ›ก๏ธ Cybersecurity Deep Dive: Attack & DefenseImportance, attack traces, defense, exploit analysis
5๐Ÿ› ๏ธ How to Use & TrainLoading, streaming, training scripts
6๐Ÿ” Licensing & LimitationsLicense, intended use, limitations
7๐Ÿ“œ ChangelogVersion history

๐Ÿ“Š Dataset Summary

<div align="center">

๐ŸŽฏ The Numbers That Matter

MetricValueStatus
Total Storage76 GB+โœ… Verified
JSONL Data Shards516โœ… Verified
Archive Files (tar.gz)7,090โœ… Verified
Source Datasets73โœ… Verified
Categories8โœ… Verified
Total Samples18M+โœ… Verified
Largest Source8.15M (Vibe-Coding-Instruct-V2)โœ…
Archive Size~64 GB (compressed GitHub repos)โœ…
Cybersecurity Sources6โœ…
Cybersecurity Data Size~2.6 GBโœ…

</div>

<br>

๐ŸŒŸ Why "Ultimate Distilled"?

This dataset is not a raw scrape. Every sample has been distilled through a unified extraction pipeline:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    UNIFIED EXTRACTION PIPELINE              โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                                             โ”‚
โ”‚  73 Upstream Sources (ALL FULLY PROCESSED, NO SKIP)         โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ”          โ”‚
โ”‚  โ”‚ HF  โ”‚ โ”‚ HF  โ”‚ โ”‚ HF  โ”‚ โ”‚ GH  โ”‚ โ”‚ HF  โ”‚ โ”‚ ... โ”‚          โ”‚
โ”‚  โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜          โ”‚
โ”‚     โ”‚       โ”‚       โ”‚       โ”‚       โ”‚       โ”‚               โ”‚
โ”‚     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜               โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚ EXTRACT โ”‚ โ† Field normalization        โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜   (instruction/response)      โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚CATEGORIZEโ”‚ โ† 8 semantic categories      โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚  SHARD  โ”‚ โ† 20K samples per shard       โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚ UPLOAD  โ”‚ โ† Batch commits to HF         โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                                                             โ”‚
โ”‚  STATUS: ALL 73 SOURCES COMPLETE. NO SKIPPING. 18M+ ROWS.   โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

<br>

๐Ÿ’Ž Value to the Open-Source AI Community

๐ŸŽฏ For...๐Ÿ“ฆ This dataset provides...
Model TrainersSingle load_dataset() call to stream 18M+ SFT-ready samples
Coding Agent Researchers11M+ agentic coding traces from Fable-5, Vibe-Coding, Royal Ghost, Kimi, DeepSeek
Code Pretraining7,090 full GitHub repository snapshots (64 GB compressed)
Reasoning Researchers2.7M+ distilled reasoning traces from Claude, Gemini, Grok, GPT-5.5, Opus 4.8
Domain Specialists25K-sample sweeps across 29 disciplines
Cybersecurity ResearchersDedicated cybersecurity category with attack/defense/exploit traces, red/blue team dialogues, and incident reports
Red Team / Blue Team TrainersRealistic attack scenarios, defense strategies, exploit code, and post-mortem analysis

๐Ÿ—‚๏ธ Directory Structure

๐Ÿ“‚ Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset/
โ”‚
โ”œโ”€โ”€ ๐Ÿ“ฆ archives/                          # ~64 GB โ€” 7,090 compressed GitHub repos
โ”‚   โ”œโ”€โ”€ 0-chi__sonaure-lp.tar.gz
โ”‚   โ”œโ”€โ”€ 00MB__bitcoin_trading_bot.tar.gz
โ”‚   โ”œโ”€โ”€ 0101-agents__plugins.tar.gz
โ”‚   โ”œโ”€โ”€ ... (7,090 files total)
โ”‚   โ””โ”€โ”€ zznmg1__playable-survivor-ad.tar.gz
โ”‚
โ”œโ”€โ”€ ๐Ÿ“ data/                              # ~12 GB โ€” 516 JSONL shards (18M+ samples)
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ’ป coding/                        # 28 sources ยท ~11M+ samples
โ”‚   โ”‚   โ”œโ”€โ”€ vibe_instruct_v2/             # 8,152,510 samples
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_2m/                    # 2,006,487 samples
โ”‚   โ”‚   โ”œโ”€โ”€ vibe_instruct_v1/             # 1,100,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ vibe_coding/                  # 1,100,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ royal_ghost_1m/               # 1,000,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ citation_ground/              # 980,064 samples
โ”‚   โ”‚   โ”œโ”€โ”€ royal_ghost_501k/             # 703,449 samples
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_repos_full/            # 7,090 archive pointers
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_agentic_sft/           # 159,972 samples
โ”‚   โ”‚   โ”œโ”€โ”€ gpt55_codex/                  # 119,436 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ alpca_gpt55/                  # 49,099 samples
โ”‚   โ”‚   โ”œโ”€โ”€ deepseek_v4_pro_agent/        # 96,597 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_traces/                # 49,544 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code_100k/            # 68,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code/                 # 49,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ kimi_coding/                  # 9,014 samples
โ”‚   โ”‚   โ”œโ”€โ”€ mimo_claude_code_traces/      # 15,046 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ kimi_k26_claude_code_traces/  # 7,438 samples
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code_10k/             # 9,800 samples
โ”‚   โ”‚   โ”œโ”€โ”€ legend_python/                # 5,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ autonomy/                     # 10,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code_demo/            # 1,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ god_coder/                    # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ python_god_coder/             # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ elite_god_coder/              # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ omega_genesis/                # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ open_tool_trace/              # 48 samples
โ”‚   โ”‚   โ””โ”€โ”€ genesis_v11/                  # partial recovery
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿงฎ math/                          # 2 sources
โ”‚   โ”‚   โ”œโ”€โ”€ math_25k/
โ”‚   โ”‚   โ””โ”€โ”€ deepseek_prover_v1/           # 27,503 Lean theorem proofs
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ”ฌ science/                       # 7 sources
โ”‚   โ”‚   โ”œโ”€โ”€ science_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ physics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ chemistry_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ biology_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ medical_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ cs_25k/
โ”‚   โ”‚   โ””โ”€โ”€ biology_r2med/                # โญ NEW
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ โš™๏ธ applied/                       # 8 sources
โ”‚   โ”‚   โ”œโ”€โ”€ robotics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ nano_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ materials_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ earth_climate_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ renewable_energy_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ evolution_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ universe_25k/
โ”‚   โ”‚   โ””โ”€โ”€ kardashev_25k/
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ“š humanities/                    # 8 sources
โ”‚   โ”‚   โ”œโ”€โ”€ psychology_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ economics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ law_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ statistics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ sports_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ human_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ conscience_25k/
โ”‚   โ”‚   โ””โ”€โ”€ supernatural_25k/
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿง  distilled/                     # 9 sources ยท frontier distillations
โ”‚   โ”‚   โ”œโ”€โ”€ claude_mythos/
โ”‚   โ”‚   โ”œโ”€โ”€ gemini35/
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_cleaned/
โ”‚   โ”‚   โ”œโ”€โ”€ grok44/
โ”‚   โ”‚   โ”œโ”€โ”€ gemini_pro32/
โ”‚   โ”‚   โ”œโ”€โ”€ gpt55_thinking/
โ”‚   โ”‚   โ”œโ”€โ”€ gpt55_distilled/
โ”‚   โ”‚   โ”œโ”€โ”€ claude_opus_48_distill/       # โญ NEW
โ”‚   โ”‚   โ””โ”€โ”€ claude_opus_48_max_thinking/  # โญ NEW
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ“ instruction/                   # 3 sources
โ”‚   โ”‚   โ”œโ”€โ”€ alpaca/                       # 52,002 samples
โ”‚   โ”‚   โ”œโ”€โ”€ oasst/                        # 32,141 samples
โ”‚   โ”‚   โ””โ”€โ”€ dolly/                        # 15,011 samples
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ”’ cybersecurity/                 # 6 sources
โ”‚   โ”‚   โ”œโ”€โ”€ high_quality_cybersecurity/
โ”‚   โ”‚   โ”œโ”€โ”€ heimdall_v1_1/                # โญ NEW โ€” 78 MB conversations
โ”‚   โ”‚   โ”œโ”€โ”€ fenrir_v2_1/                  # โญ NEW โ€” 411 MB (2.1M+ entries)
โ”‚   โ”‚   โ”œโ”€โ”€ clydeiii_cybersecurity/       # โญ NEW โ€” 20 MB yearly corpus
โ”‚   โ”‚   โ”œโ”€โ”€ precinct6_cybersecurity/      # โญ NEW โ€” 2.1 GB (graph+signals+ref)
โ”‚   โ”‚   โ””โ”€โ”€ savani_cyber_attack/          # โญ NEW โ€” 17 MB attack CSV
โ”‚   โ”‚
โ”‚   โ””โ”€โ”€ ๐Ÿ“‡ index/                         # 2 sources
โ”‚       โ”œโ”€โ”€ species_25k/
โ”‚       โ””โ”€โ”€ transport_25k/
โ”‚
โ”œโ”€โ”€ ๐Ÿ“„ README.md
โ””โ”€โ”€ ๐Ÿ“„ dataset_info.json

๐Ÿค” Why is archives/ kept compressed?

ReasonExplanation
๐Ÿ’พ Space EfficiencyUncompressed would exceed 200+ GB. Compressed = 64 GB (3ร— saving)
๐ŸŽฏ On-Demand AccessDownload only specific repositories you need
๐Ÿ” Preservation Fidelitytar.gz preserves exact file permissions, directory structure, binaries
๐Ÿ’ก Tip: For training on code content, use data/coding/fable5_repos_full/ (475K samples, each a file extracted from archives, capped at 4KB). For full untruncated file access, stream directly from archives/.

๐ŸŒ Data Sources & Provenance

<div align="center">

๐Ÿ—บ๏ธ 73 Sources Across 8 Categories

CategorySourcesSamplesDescription
๐Ÿ’ป coding28~11M+Agentic traces, code repos, coder distillations
๐Ÿง  distilled9~200KFrontier model distillations
โš™๏ธ applied8~200KRobotics, nano, materials, climate, energy
๐Ÿ“š humanities8~200KPsychology, economics, law, statistics
๐Ÿ”ฌ science7~175KPhysics, chemistry, biology, medical, CS
๐Ÿ“ instruction3~99KClassic instruction (alpaca, oasst, dolly)
๐Ÿ“‡ index2~50KSpecies index, transport
๐Ÿ”’ cybersecurity6~2.6 GBHigh-quality attack, defense, exploit traces
๐Ÿงฎ math2~52KMath + Lean theorem proofs

</div>

<br>

๐Ÿ’ป Coding Category (28 sources โ€” ALL FULLY PROCESSED โญ)

Source SlugUpstream DatasetTypeSamples
vibe_instruct_v2CodeDevX/Vibe-Coding-Instruct-V2Agentic coding8,152,510
fable5_2mCrownelius/Complete-FABLE.5-traces-2MFable-5 traces2,006,487
vibe_instruct_v1CodeDevX/Vibe-Coding-InstructAgentic coding1,100,000
vibe_codingattentionAllYouNeed/Vibe-Coding-Claude-Fable-5Claude coding1,100,000
royal_ghost_1mWithinUsAI/Royal_Ghost_Coder_1MGhost coder1,000,000
citation_groundWithinUsAI/CitationGround-1MCitation-grounded980,064
royal_ghost_501kWithinUsAI/Royal_Ghost_Coder_501kGhost coder703,449
fable5_repos_fullnotune/fable5-repos7,090 repo pointers7,090
fable5_agentic_sftNexlab/fable5-agentic-coding-sftAgentic SFT159,972
gpt55_codexAletheiaResearch/GPT-5.5-CodexGPT-5.5 Codex119,436
alpca_gpt55GabrielFreeze-2/alpca-mlt-gpt-5.5_chatmlGPT-5.5 chatml49,099
deepseek_v4_pro_agentTeichAI/DeepSeek-v4-Pro-AgentDeepSeek v496,597
fable5_tracesGlint-Research/Fable-5-tracesFable-5 traces49,544
genesis_code_100kWithinUsAI/Genesis_AI_Code_100kGenesis code68,000
genesis_codeWithinUsAI/Genesis_AI_Code_50kGenesis code49,000
kimi_codingtrjxter/Kimi-K2.7-CodingTraces-9000xKimi K2.79,014
mimo_claude_code_traceschoucsan/mimo-claude-code-traces-1kMimo Claude15,046
kimi_k26_claude_code_tracesarmand0e/kimi-k2.6-claude-code-tracesKimi K2.67,438
genesis_code_10kWithinUsAI/Genesis_AI_Code_10kGenesis code9,800
legend_pythonWithinUsAI/Legend_Python_CoderV.1Python coder5,000
autonomyWithinUsAI/The_Autonomy_From_WithIn_10kAutonomy10,000
genesis_code_demoWithinUsAI/Genesis_AI_Code_1k_DemoGenesis demo1,000
god_coderWithinUsAI/GOD_Coder_100kGOD coderFULL โญ
python_god_coderWithinUsAI/python_GOD_coder_100kPython GODFULL โญ
elite_god_coderWithinUsAI/Elite_GOD_Coder_100kElite GODFULL โญ
omega_genesisWithinUsAI/Omega_Genesis_Coder_100kOmega GenesisFULL โญ
open_tool_traceWithinUsAI/OpenToolTrace-XTool traces48
genesis_v11WithinUsAI/Genesis_v1_1_Update...Genesis v1.1partial

<br>

๐Ÿง  Distilled Category (9 sources)

SourceUpstreamDistilled From
claude_mythosWithinUsAI/claude_mythos_distilled_25kClaude
gemini35WithinUsAI/gemini_3.5_flash_distilled_25kGemini 3.5 Flash
fable5_cleanedWithinUsAI/fable_5_distillation_merged_cleaned_25kFable-5
grok44WithinUsAI/Grok4.4_heavy_max_distill_god_seed_25kGrok 4.4
gemini_pro32WithinUsAI/GeminiPro3.2_max_distill_god_seed_25kGemini Pro 3.2
gpt55_thinkingWithinUsAI/GPT5.5_thinking_max_distill_god_seed_25KGPT-5.5
gpt55_distilledWithinUsAI/GPT_5.5_DistilledGPT-5.5
claude_opus_48_distill11-47/claude_opus_4.8_distill_5kClaude Opus 4.8 โญ
claude_opus_48_max_thinking11-47/claude_opus_4.8_max_thinking_5k_v2Opus 4.8 Max โญ

<br>

๐Ÿ”ฌ Science ยท โš™๏ธ Applied ยท ๐Ÿ“š Humanities ยท ๐Ÿงฎ Math ยท ๐Ÿ“ Instruction ยท ๐Ÿ”’ Cybersecurity ยท ๐Ÿ“‡ Index

<details> <summary>๐Ÿ“– Click to expand all other categories</summary>

๐Ÿ”ฌ Science (7 sources): science_25k, physics_25k, chemistry_25k, biology_25k, medical_25k, cs_25k, biology_r2med (R2MED/Biology)

โš™๏ธ Applied (8 sources): robotics_25k, nano_25k, materials_25k, earth_climate_25k, renewable_energy_25k, evolution_25k, universe_25k, kardashev_25k

๐Ÿ“š Humanities (8 sources): psychology_25k, economics_25k, law_25k, statistics_25k, sports_25k, human_25k, conscience_25k, supernatural_25k

๐Ÿงฎ Math (2 sources): math_25k, deepseek_prover_v1 (27,503 Lean proofs)

๐Ÿ“ Instruction (3 sources): alpaca (52K), oasst (32K), dolly (15K)

๐Ÿ”’ Cybersecurity (6 sources): high_quality_cybersecurity, heimdall_v1_1, fenrir_v2_1, clydeiii_cybersecurity, precinct6_cybersecurity, savani_cyber_attack

๐Ÿ“‡ Index (2 sources): species_25k, transport_25k

</details>


๐Ÿ›ก๏ธ Cybersecurity Deep Dive: Attack & Defense

โš”๏ธ Why This Matters

Modern AI systems are increasingly deployed in security-critical environmentsโ€”yet most open-source training data ignores real-world adversarial scenarios. The Open Distillation Codex includes a dedicated cybersecurity category designed to equip models with:

  • โ€”Attack Awareness: Recognize and generate realistic attack patterns, exploits, penetration testing commands, and social engineering dialogues.
  • โ€”Defense Proficiency: Learn to propose defensive measures, detect anomalies, and articulate incident response protocols.
  • โ€”Exploit Understanding: Analyze and explain software vulnerabilities, craft proof-of-concept code (for educational purposes), and understand exploit chains.
  • โ€”Red/Blue Team Simulation: Engage in multi-turn conversations mimicking red team attack planning and blue team defense coordination.
  • โ€”Threat Intelligence: Summarize, classify, and reason about cyber threat reports, CVEs, and IOCs (Indicators of Compromise).

This makes the dataset a powerful foundation for building cybersecurity-aware LLMs, security co-pilots, and automated vulnerability assessment tools.

๐Ÿ“Š Whatโ€™s Inside the Cybersecurity Category?

SourceDescriptionData FormatKey Themes
high_quality_cybersecurityManually curated high-quality instructionโ€“response pairs covering attack techniques, defense, and policyJSONL (shards)MITRE ATT&CK, OWASP, incident response
heimdall_v1_1~78 MB of security conversations, including red/blue team dialogues and threat analysisJSONLMulti-turn chat, tool usage
fenrir_v2_1411 MB, 2.1M+ entries โ€” massive corpus of cybersecurity Q&A, exploit descriptions, and code snippetsJSONLExploit code, CVEs, vulnerability research
clydeiii_cybersecurity20 MB yearly security corpus, aggregated from public reports and advisoriesJSONLYear-in-review, trends, threat landscape
precinct6_cybersecurity2.1 GB graph-based dataset with network signals, attack graphs, and reference materialsJSONL (graph+signals+ref)Network attacks, lateral movement, detection
savani_cyber_attack17 MB CSV of labeled cyber attack incidents with detailed featuresCSVAttack classification, feature analysis

๐Ÿงช Attack & Exploit Examples

Here are a few representative samples (sanitized) from the dataset:

Example 1 โ€“ SQL Injection Exploit

json
{
  "source": "fenrir_v2_1",
  "instruction": "Explain how an attacker could exploit a vulnerable login form using SQL injection. Provide a sample malicious input.",
  "response": "An attacker could bypass authentication by injecting SQL logic. For a query like:\n`SELECT * FROM users WHERE username='$user' AND password='$pass'`,\nthe attacker might enter `admin' --` as the username and anything as password. The resulting query becomes:\n`SELECT * FROM users WHERE username='admin' --' AND password='...'`\nThe `--` comments out the password check, granting access as admin. More advanced attacks use UNION to extract data."
}

Example 2 โ€“ Red Team Command Sequence

json
{
  "source": "heimdall_v1_1",
  "instruction": "As a red team operator, list the steps you would take after gaining initial access to a Windows machine using a phishing payload.",
  "response": "1. Establish persistence: `schtasks /create /tn \"Updater\" /tr \"C:\\malware.exe\" /sc hourly`\n2. Enumerate local users: `net user /domain`\n3. Dump credentials: `mimikatz sekurlsa::logonpasswords`\n4. Lateral movement: use PsExec to move to file server.\n5. Exfiltrate data: compress sensitive files and upload via DNS tunneling.\nAlways maintain C2 communication over HTTPS to blend in."
}

Example 3 โ€“ Defense Playbook (Blue Team)

json
{
  "source": "high_quality_cybersecurity",
  "instruction": "An IDS alert flags unusual outbound DNS traffic from a workstation. What's the immediate response?",
  "response": "1. Isolate the workstation from the network (VLAN change or port shutdown).\n2. Capture volatile memory and network logs for forensics.\n3. Check DNS queries: if long, random-looking subdomains, suspect DNS tunneling.\n4. Scan for malware with updated signatures.\n5. Review firewall logs for similar traffic from other hosts.\n6. If confirmed, initiate incident response playbook for data exfiltration."
}

๐ŸŽ“ How to Train a Cybersecurity-Focused LLM

python
from datasets import load_dataset

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# Load only cybersecurity category
cyber_ds = load_dataset(REPO, split="train", 
                        data_files="data/cybersecurity/**/*.jsonl",
                        streaming=True)

# Or load specific sources
fenrir = load_dataset(REPO, split="train", 
                      data_files="data/cybersecurity/fenrir_v2_1/*.jsonl")

# Format for SFT
def format_security_sample(example):
    return {
        "text": f"### Security Task:\n{example['instruction']}\n\n### Expert Response:\n{example['response']}"
    }

cyber_ds = cyber_ds.map(format_security_sample)

# Now train with your favourite framework (transformers, axolotl, etc.)

Curriculum Idea:

  1. 1.Start with high_quality_cybersecurity and heimdall_v1_1 for foundational attack/defense conversations.
  2. 2.Introduce fenrir_v2_1 for exploit code and vulnerability deep dives.
  3. 3.Use precinct6_cybersecurity for network-level attack graph understanding.

๐Ÿ›ก๏ธ Ethical & Responsible Use

  • โ€”For Defensive Purposes Only: This data is intended to strengthen AI for defense, threat detection, and security education. Do not use it to generate active attack code without proper authorization.
  • โ€”No Zero-Day Exploits: The dataset contains only already-public vulnerabilities and techniques. It does not include zero-day or weaponized exploits.
  • โ€”Responsible Disclosure: If you fine-tune a model with this data, we recommend adding a safety preamble warning that generated security content must be used legally and ethically.
  • โ€”Dual-Use Awareness: While we believe open access improves collective security, we acknowledge the dual-use nature. Users are expected to follow applicable laws and guidelines.
โš ๏ธ Disclaimer: This dataset includes descriptions of attack techniques for educational purposes. The maintainers are not responsible for misuse.

๐Ÿ“ˆ Future Additions

  • โ€”Integration with CTF (Capture The Flag) challenge walkthroughs.
  • โ€”More blue team procedures and SOAR playbooks.
  • โ€”Anonymized real-world incident response logs (with permission).

๐Ÿ› ๏ธ How to Use & Train

1๏ธโƒฃ Load Categorized JSONL Data

python
from datasets import load_dataset

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# โ”€ Load a single category โ”€
ds = load_dataset(REPO, split="train", data_files="data/coding/*/*.jsonl", streaming=True)

# โ”€ Load a specific source โ”€
ds = load_dataset(REPO, split="train", data_files="data/coding/vibe_instruct_v2/*.jsonl", streaming=True)

# โ”€ Load everything (18M+ samples) โ”€
ds = load_dataset(REPO, split="train", streaming=True)

for sample in ds:
    print(sample["source"], sample["instruction"][:80])

<br>

2๏ธโƒฃ Stream the 64 GB archives/ GitHub Repositories

python
from huggingface_hub import hf_hub_download
import tarfile

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# โ”€ Option A: Download & extract ONE repository โ”€
hf_hub_download(
    repo_id=REPO,
    repo_type="dataset",
    filename="archives/0x101__lakewatch.tar.gz",
    local_dir="./repos",
)
with tarfile.open("./repos/archives/0x101__lakewatch.tar.gz", "r:gz") as tar:
    tar.extractall("./extracted/0x101__lakewatch")


# โ”€ Option B: Stream files WITHOUT full extraction โ”€
def stream_repo_files(archive_name, max_files=100):
    """Stream file contents from tar.gz without extracting to disk."""
    local_path = hf_hub_download(repo_id=REPO, repo_type="dataset", filename=archive_name)
    
    with tarfile.open(local_path, "r:gz") as tar:
        count = 0
        for member in tar:
            if member.isfile() and count < max_files:
                f = tar.extractfile(member)
                if f:
                    yield {
                        "path": member.name,
                        "content": f.read().decode("utf-8", errors="ignore")[:4000],
                    }
                    count += 1
    
    import os
    os.remove(local_path)  # Clean up

# Stream files from a specific repo
for file_data in stream_repo_files("archives/0x101__lakewatch.tar.gz"):
    print(f"๐Ÿ“„ {file_data['path']}: {file_data['content'][:100]}...")


# โ”€ Option C: Use pre-extracted JSONL shards (475K samples) โ”€
code_ds = load_dataset(
    REPO, split="train",
    data_files="data/coding/fable5_repos_full/*.jsonl",
    streaming=True
)
# Each sample: instruction = "<repo>/<file>", response = "<content>"

<br>

3๏ธโƒฃ SFT Training Script (Hugging Face Trainer)

python
import torch
from datasets import load_dataset
from transformers import (
    AutoTokenizer,
    AutoModelForCausalLM,
    TrainingArguments,
    Trainer,
    DataCollatorForLanguageModeling,
)

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# CONFIGURATION
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
MODEL_NAME = "meta-llama/Llama-3.1-8B"
DATASET_REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"
OUTPUT_DIR = "./sft-output"
MAX_SEQ_LEN = 2048

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# LOAD MODEL & TOKENIZER
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
tokenizer.pad_token = tokenizer.eos_token

model = AutoModelForCausalLM.from_pretrained(
    MODEL_NAME,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="flash_attention_2",
)

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# LOAD & FORMAT DATASET
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
def format_instruction(sample):
    text = f"### Instruction:\n{sample['instruction']}\n\n### Response:\n{sample['response']}"
    return {"text": text}

def tokenize(examples):
    return tokenizer(
        examples["text"],
        truncation=True,
        max_length=MAX_SEQ_LEN,
        padding="max_length",
    )

# Load coding category (use "data/**/*.jsonl" for full 18M+)
train_ds = load_dataset(
    DATASET_REPO,
    split="train",
    data_files="data/coding/*/*.jsonl",
    streaming=True,
)
train_ds = train_ds.map(format_instruction).filter(lambda x: len(x["text"]) > 0)
train_ds = train_ds.map(tokenize, batched=True)

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# TRAIN
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
training_args = TrainingArguments(
    output_dir=OUTPUT_DIR,
    num_train_epochs=3,
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,
    warmup_steps=500,
    logging_steps=100,
    save_steps=2000,
    learning_rate=2e-5,
    bf16=True,
    gradient_checkpointing=True,
    optim="adamw_torch",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=train_ds,
    data_collator=DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False),
)

trainer.train()
trainer.save_model(OUTPUT_DIR)

<br>

4๏ธโƒฃ Curriculum Learning Across Categories

python
from datasets import load_dataset, interleave_datasets

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# โ”€ Phase 1: Foundation (math + science) โ”€
phase1_math = load_dataset(REPO, split="train", data_files="data/math/**/*.jsonl", streaming=True)
phase1_sci = load_dataset(REPO, split="train", data_files="data/science/**/*.jsonl", streaming=True)
phase1 = interleave_datasets([phase1_math, phase1_sci])

# โ”€ Phase 2: Add coding traces โ”€
phase2 = load_dataset(REPO, split="train", data_files="data/coding/**/*.jsonl", streaming=True)

# โ”€ Phase 3: Add distilled reasoning + cybersecurity โ”€
phase3_distilled = load_dataset(REPO, split="train", data_files="data/distilled/**/*.jsonl", streaming=True)
phase3_cyber = load_dataset(REPO, split="train", data_files="data/cybersecurity/**/*.jsonl", streaming=True)
phase3 = interleave_datasets([phase3_distilled, phase3_cyber])

# Train sequentially
# trainer.train(phase1)  # epochs 0-1
# trainer.train(phase2)  # epochs 1-2
# trainer.train(phase3)  # epochs 2-3

<br>

๐Ÿ“‹ Schema Reference

json
{
    "source":          "fable5_2m",
    "source_dataset":  "Crownelius/Complete-FABLE.5-traces-2M",
    "instruction":     "<the prompt / question / file path>",
    "response":        "<the completion / answer / file content>",
    "category":        "coding"
}
FieldTypeMax LengthDescription
sourcestring200Short slug identifying upstream dataset
source_datasetstring200Full HF repo id (org/name)
instructionstring4,000User-side content (prompt/question/file path)
responsestring4,000Assistant-side content (completion/answer/file content)
categorystring50One of 8 categories

๐Ÿ” Licensing & Limitations

๐Ÿ“œ License

The collection as a whole is released under the MIT License.

Each upstream dataset retains its original license. The source_dataset field on every row identifies the upstream โ€” look it up on Hugging Face to determine its specific license.

LicenseApplies To
MITMost WithinUsAI datasets, OpenAssistant
Apache-2.0DeepSeek, OpenThoughts
CC-BY-4.0Dolly, various
CC-BY-SA-3.0Databricks Dolly
AGPL-3.0Some Fable-5 traces

โœ… Intended Use Cases (Our Vision)

  • โ€”Fine-tuning open-source LLMs for instruction following
  • โ€”Training coding agents and code-completion models
  • โ€”Reasoning chain distillation research
  • โ€”Domain-specific adaptation (math, science, cybersecurity)
  • โ€”Repository-scale context training (using archives/)

โŒ Not Recommended For

  • โ€”Deploying models without safety evaluation
  • โ€”Generating harmful, biased, or deceptive content
  • โ€”High-stakes domains (medical, legal, financial) without expert review
  • โ€”Claiming models "know" facts โ€” this is distilled output, not ground truth

โš ๏ธ Limitations

  1. 1.Field length cap: instruction and response capped at 4,000 characters. For full content, use archives/.
  2. 2.Distillation artifacts: Samples are model-generated โ€” may contain hallucinations or biases.
  3. 3.Partial recovery: A few upstream datasets (GODCoder variants, Genesisv1.1) had format errors and were partially recovered via raw JSONL parsing.

๐Ÿ“ Citation

bibtex
@misc{open_distillation_codex_2026,
  title  = {The Open Distillation Codex: 18M+ samples + 7090 code repositories from 73 sources with Cybersecurity Attack & Defense},
  author = {Manusagents},
  year   = {2026},
  url    = {https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset},
  note   = {v8.2 - No skip, full. 516 shards + 7090 archives, 73 sources, 8 categories, 76 GB+}
}

๐Ÿ“œ Changelog

VersionDateKey Changes
v1.0โ€“v5.02026-07-01 to 05Progressive builds: 117K โ†’ 20.7M samples
v6.02026-07-06Category restructuring: data/<category>/<source>/shard-*.jsonl
v7.02026-07-06Training scripts + full processing started
v8.0 FINAL2026-07-06ALL sources FULLY processed โ€” no skipping. Verified 79.13 GB.
v8.12026-07-08Added 5 external cybersecurity datasets. Total 81.2 GB, 73 sources.
v8.22026-07-18Final numbers rectified: 18M+ samples, 76 GB+ total. All sources no skip, fully verified. Enhanced cybersecurity deep-dive with attack/defense examples, training scripts, ethical guidelines.

<div align="center">

<br>

๐ŸŒŸ The Open Distillation Codex ๐ŸŒŸ

73 sources ยท 8 categories ยท 7,090 repositories ยท 516 shards ยท 76 GB+

<br>

No skip. Full. 18M+ samples. Built one archive at a time. Released under MIT.

<br>


"Two layers. Eight categories. Seventy-three sources. One codex. No skip. Full. Armed with cybersecurity attack and defense."

<br>

<img src="https://img.shields.io/badge/Built%20with-Streaming%20Pipeline-blue?style=flat-square" alt="Streaming"> <img src="https://img.shields.io/badge/No-Skipping-brightgreen?style=flat-square" alt="No Skip"> <img src="https://img.shields.io/badge/Full%20Processing-success?style=flat-square" alt="Full"> <img src="https://img.shields.io/badge/Format-JSONL-orange?style=flat-square" alt="JSONL"> <img src="https://img.shields.io/badge/HuggingFace-Dataset-yellow?style=flat-square" alt="HF"> <img src="https://img.shields.io/badge/Cybersecurity-Deep%20Dive-purple?style=flat-square" alt="Cyber">

<br><br>

โ€” The Open Distillation Codex โ€”

</div>