SicariusSicariiStuff/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
๐ The Open Distillation Codex ๐ The Ultimate Open-Source Distillation Dataset โ No Skip, Full, with Attack & Defense ๐ Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals ยท 7,090 Raw GitHub Repositories ยท 8 Curated Categories ยท ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo โ of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-threeโฆ See the full description on the dataset page: https://huggingface.co/datasets/SicariusSicariiStuff/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.
<div align="center">
<img src="https://img.shields.io/badge/Version-8.2-blue?style=for-the-badge" alt="Version"> <img src="https://img.shields.io/badge/Storage-76GB%2B-green?style=for-the-badge" alt="Storage"> <img src="https://img.shields.io/badge/Sources-73-orange?style=for-the-badge" alt="Sources"> <img src="https://img.shields.io/badge/License-MIT-yellow?style=for-the-badge" alt="License"> <img src="https://img.shields.io/badge/Samples-18M%2B-red?style=for-the-badge" alt="Samples"> <img src="https://img.shields.io/badge/Cybersecurity-6%20Sources-purple?style=for-the-badge" alt="Cybersecurity">
<br><br>
๐ The Open Distillation Codex
๐ The Ultimate Open-Source Distillation Dataset โ No Skip, Full, with Attack & Defense ๐
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals ยท 7,090 Raw GitHub Repositories ยท 8 Curated Categories ยท ~76 GB+
<br>
"We did not write this dataset. We assembled it. Every line is an echo โ of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three sources. Eight categories. Zero gatekeeping. No skipping. Fully processed. Now fortified with real-world cybersecurity confrontations."
<br>
</div>
๐ Table of Contents
๐ Dataset Summary
<div align="center">
๐ฏ The Numbers That Matter
</div>
<br>
๐ Why "Ultimate Distilled"?
This dataset is not a raw scrape. Every sample has been distilled through a unified extraction pipeline:
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ UNIFIED EXTRACTION PIPELINE โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ โ
โ 73 Upstream Sources (ALL FULLY PROCESSED, NO SKIP) โ
โ โโโโโโโ โโโโโโโ โโโโโโโ โโโโโโโ โโโโโโโ โโโโโโโ โ
โ โ HF โ โ HF โ โ HF โ โ GH โ โ HF โ โ ... โ โ
โ โโโโฌโโโ โโโโฌโโโ โโโโฌโโโ โโโโฌโโโ โโโโฌโโโ โโโโฌโโโ โ
โ โ โ โ โ โ โ โ
โ โโโโโโโโโดโโโโโโโโดโโโโโโโโผโโโโโโโโดโโโโโโโโ โ
โ โ โ
โ โโโโโโผโโโโโ โ
โ โ EXTRACT โ โ Field normalization โ
โ โโโโโโฌโโโโโ (instruction/response) โ
โ โ โ
โ โโโโโโผโโโโโ โ
โ โCATEGORIZEโ โ 8 semantic categories โ
โ โโโโโโฌโโโโโ โ
โ โ โ
โ โโโโโโผโโโโโ โ
โ โ SHARD โ โ 20K samples per shard โ
โ โโโโโโฌโโโโโ โ
โ โ โ
โ โโโโโโผโโโโโ โ
โ โ UPLOAD โ โ Batch commits to HF โ
โ โโโโโโโโโโโ โ
โ โ
โ STATUS: ALL 73 SOURCES COMPLETE. NO SKIPPING. 18M+ ROWS. โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ<br>
๐ Value to the Open-Source AI Community
๐๏ธ Directory Structure
๐ Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset/
โ
โโโ ๐ฆ archives/ # ~64 GB โ 7,090 compressed GitHub repos
โ โโโ 0-chi__sonaure-lp.tar.gz
โ โโโ 00MB__bitcoin_trading_bot.tar.gz
โ โโโ 0101-agents__plugins.tar.gz
โ โโโ ... (7,090 files total)
โ โโโ zznmg1__playable-survivor-ad.tar.gz
โ
โโโ ๐ data/ # ~12 GB โ 516 JSONL shards (18M+ samples)
โ โ
โ โโโ ๐ป coding/ # 28 sources ยท ~11M+ samples
โ โ โโโ vibe_instruct_v2/ # 8,152,510 samples
โ โ โโโ fable5_2m/ # 2,006,487 samples
โ โ โโโ vibe_instruct_v1/ # 1,100,000 samples
โ โ โโโ vibe_coding/ # 1,100,000 samples
โ โ โโโ royal_ghost_1m/ # 1,000,000 samples
โ โ โโโ citation_ground/ # 980,064 samples
โ โ โโโ royal_ghost_501k/ # 703,449 samples
โ โ โโโ fable5_repos_full/ # 7,090 archive pointers
โ โ โโโ fable5_agentic_sft/ # 159,972 samples
โ โ โโโ gpt55_codex/ # 119,436 samples โญ FULL
โ โ โโโ alpca_gpt55/ # 49,099 samples
โ โ โโโ deepseek_v4_pro_agent/ # 96,597 samples โญ FULL
โ โ โโโ fable5_traces/ # 49,544 samples โญ FULL
โ โ โโโ genesis_code_100k/ # 68,000 samples
โ โ โโโ genesis_code/ # 49,000 samples
โ โ โโโ kimi_coding/ # 9,014 samples
โ โ โโโ mimo_claude_code_traces/ # 15,046 samples โญ FULL
โ โ โโโ kimi_k26_claude_code_traces/ # 7,438 samples
โ โ โโโ genesis_code_10k/ # 9,800 samples
โ โ โโโ legend_python/ # 5,000 samples
โ โ โโโ autonomy/ # 10,000 samples
โ โ โโโ genesis_code_demo/ # 1,000 samples
โ โ โโโ god_coder/ # โญ FULL raw recovery
โ โ โโโ python_god_coder/ # โญ FULL raw recovery
โ โ โโโ elite_god_coder/ # โญ FULL raw recovery
โ โ โโโ omega_genesis/ # โญ FULL raw recovery
โ โ โโโ open_tool_trace/ # 48 samples
โ โ โโโ genesis_v11/ # partial recovery
โ โ
โ โโโ ๐งฎ math/ # 2 sources
โ โ โโโ math_25k/
โ โ โโโ deepseek_prover_v1/ # 27,503 Lean theorem proofs
โ โ
โ โโโ ๐ฌ science/ # 7 sources
โ โ โโโ science_25k/
โ โ โโโ physics_25k/
โ โ โโโ chemistry_25k/
โ โ โโโ biology_25k/
โ โ โโโ medical_25k/
โ โ โโโ cs_25k/
โ โ โโโ biology_r2med/ # โญ NEW
โ โ
โ โโโ โ๏ธ applied/ # 8 sources
โ โ โโโ robotics_25k/
โ โ โโโ nano_25k/
โ โ โโโ materials_25k/
โ โ โโโ earth_climate_25k/
โ โ โโโ renewable_energy_25k/
โ โ โโโ evolution_25k/
โ โ โโโ universe_25k/
โ โ โโโ kardashev_25k/
โ โ
โ โโโ ๐ humanities/ # 8 sources
โ โ โโโ psychology_25k/
โ โ โโโ economics_25k/
โ โ โโโ law_25k/
โ โ โโโ statistics_25k/
โ โ โโโ sports_25k/
โ โ โโโ human_25k/
โ โ โโโ conscience_25k/
โ โ โโโ supernatural_25k/
โ โ
โ โโโ ๐ง distilled/ # 9 sources ยท frontier distillations
โ โ โโโ claude_mythos/
โ โ โโโ gemini35/
โ โ โโโ fable5_cleaned/
โ โ โโโ grok44/
โ โ โโโ gemini_pro32/
โ โ โโโ gpt55_thinking/
โ โ โโโ gpt55_distilled/
โ โ โโโ claude_opus_48_distill/ # โญ NEW
โ โ โโโ claude_opus_48_max_thinking/ # โญ NEW
โ โ
โ โโโ ๐ instruction/ # 3 sources
โ โ โโโ alpaca/ # 52,002 samples
โ โ โโโ oasst/ # 32,141 samples
โ โ โโโ dolly/ # 15,011 samples
โ โ
โ โโโ ๐ cybersecurity/ # 6 sources
โ โ โโโ high_quality_cybersecurity/
โ โ โโโ heimdall_v1_1/ # โญ NEW โ 78 MB conversations
โ โ โโโ fenrir_v2_1/ # โญ NEW โ 411 MB (2.1M+ entries)
โ โ โโโ clydeiii_cybersecurity/ # โญ NEW โ 20 MB yearly corpus
โ โ โโโ precinct6_cybersecurity/ # โญ NEW โ 2.1 GB (graph+signals+ref)
โ โ โโโ savani_cyber_attack/ # โญ NEW โ 17 MB attack CSV
โ โ
โ โโโ ๐ index/ # 2 sources
โ โโโ species_25k/
โ โโโ transport_25k/
โ
โโโ ๐ README.md
โโโ ๐ dataset_info.json๐ค Why is archives/ kept compressed?
๐ก Tip: For training on code content, usedata/coding/fable5_repos_full/(475K samples, each a file extracted from archives, capped at 4KB). For full untruncated file access, stream directly fromarchives/.
๐ Data Sources & Provenance
<div align="center">
๐บ๏ธ 73 Sources Across 8 Categories
</div>
<br>
๐ป Coding Category (28 sources โ ALL FULLY PROCESSED โญ)
<br>
๐ง Distilled Category (9 sources)
<br>
๐ฌ Science ยท โ๏ธ Applied ยท ๐ Humanities ยท ๐งฎ Math ยท ๐ Instruction ยท ๐ Cybersecurity ยท ๐ Index
<details> <summary>๐ Click to expand all other categories</summary>
๐ฌ Science (7 sources): science_25k, physics_25k, chemistry_25k, biology_25k, medical_25k, cs_25k, biology_r2med (R2MED/Biology)
โ๏ธ Applied (8 sources): robotics_25k, nano_25k, materials_25k, earth_climate_25k, renewable_energy_25k, evolution_25k, universe_25k, kardashev_25k
๐ Humanities (8 sources): psychology_25k, economics_25k, law_25k, statistics_25k, sports_25k, human_25k, conscience_25k, supernatural_25k
๐งฎ Math (2 sources): math_25k, deepseek_prover_v1 (27,503 Lean proofs)
๐ Instruction (3 sources): alpaca (52K), oasst (32K), dolly (15K)
๐ Cybersecurity (6 sources): high_quality_cybersecurity, heimdall_v1_1, fenrir_v2_1, clydeiii_cybersecurity, precinct6_cybersecurity, savani_cyber_attack
๐ Index (2 sources): species_25k, transport_25k
</details>
๐ก๏ธ Cybersecurity Deep Dive: Attack & Defense
โ๏ธ Why This Matters
Modern AI systems are increasingly deployed in security-critical environmentsโyet most open-source training data ignores real-world adversarial scenarios. The Open Distillation Codex includes a dedicated cybersecurity category designed to equip models with:
- Attack Awareness: Recognize and generate realistic attack patterns, exploits, penetration testing commands, and social engineering dialogues.
- Defense Proficiency: Learn to propose defensive measures, detect anomalies, and articulate incident response protocols.
- Exploit Understanding: Analyze and explain software vulnerabilities, craft proof-of-concept code (for educational purposes), and understand exploit chains.
- Red/Blue Team Simulation: Engage in multi-turn conversations mimicking red team attack planning and blue team defense coordination.
- Threat Intelligence: Summarize, classify, and reason about cyber threat reports, CVEs, and IOCs (Indicators of Compromise).
This makes the dataset a powerful foundation for building cybersecurity-aware LLMs, security co-pilots, and automated vulnerability assessment tools.
๐ Whatโs Inside the Cybersecurity Category?
๐งช Attack & Exploit Examples
Here are a few representative samples (sanitized) from the dataset:
Example 1 โ SQL Injection Exploit
{
"source": "fenrir_v2_1",
"instruction": "Explain how an attacker could exploit a vulnerable login form using SQL injection. Provide a sample malicious input.",
"response": "An attacker could bypass authentication by injecting SQL logic. For a query like:\n`SELECT * FROM users WHERE username='$user' AND password='$pass'`,\nthe attacker might enter `admin' --` as the username and anything as password. The resulting query becomes:\n`SELECT * FROM users WHERE username='admin' --' AND password='...'`\nThe `--` comments out the password check, granting access as admin. More advanced attacks use UNION to extract data."
}Example 2 โ Red Team Command Sequence
{
"source": "heimdall_v1_1",
"instruction": "As a red team operator, list the steps you would take after gaining initial access to a Windows machine using a phishing payload.",
"response": "1. Establish persistence: `schtasks /create /tn \"Updater\" /tr \"C:\\malware.exe\" /sc hourly`\n2. Enumerate local users: `net user /domain`\n3. Dump credentials: `mimikatz sekurlsa::logonpasswords`\n4. Lateral movement: use PsExec to move to file server.\n5. Exfiltrate data: compress sensitive files and upload via DNS tunneling.\nAlways maintain C2 communication over HTTPS to blend in."
}Example 3 โ Defense Playbook (Blue Team)
{
"source": "high_quality_cybersecurity",
"instruction": "An IDS alert flags unusual outbound DNS traffic from a workstation. What's the immediate response?",
"response": "1. Isolate the workstation from the network (VLAN change or port shutdown).\n2. Capture volatile memory and network logs for forensics.\n3. Check DNS queries: if long, random-looking subdomains, suspect DNS tunneling.\n4. Scan for malware with updated signatures.\n5. Review firewall logs for similar traffic from other hosts.\n6. If confirmed, initiate incident response playbook for data exfiltration."
}๐ How to Train a Cybersecurity-Focused LLM
from datasets import load_dataset
REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"
# Load only cybersecurity category
cyber_ds = load_dataset(REPO, split="train",
data_files="data/cybersecurity/**/*.jsonl",
streaming=True)
# Or load specific sources
fenrir = load_dataset(REPO, split="train",
data_files="data/cybersecurity/fenrir_v2_1/*.jsonl")
# Format for SFT
def format_security_sample(example):
return {
"text": f"### Security Task:\n{example['instruction']}\n\n### Expert Response:\n{example['response']}"
}
cyber_ds = cyber_ds.map(format_security_sample)
# Now train with your favourite framework (transformers, axolotl, etc.)Curriculum Idea:
- Start with
high_quality_cybersecurityandheimdall_v1_1for foundational attack/defense conversations. - Introduce
fenrir_v2_1for exploit code and vulnerability deep dives. - Use
precinct6_cybersecurityfor network-level attack graph understanding.
๐ก๏ธ Ethical & Responsible Use
- For Defensive Purposes Only: This data is intended to strengthen AI for defense, threat detection, and security education. Do not use it to generate active attack code without proper authorization.
- No Zero-Day Exploits: The dataset contains only already-public vulnerabilities and techniques. It does not include zero-day or weaponized exploits.
- Responsible Disclosure: If you fine-tune a model with this data, we recommend adding a safety preamble warning that generated security content must be used legally and ethically.
- Dual-Use Awareness: While we believe open access improves collective security, we acknowledge the dual-use nature. Users are expected to follow applicable laws and guidelines.
โ ๏ธ Disclaimer: This dataset includes descriptions of attack techniques for educational purposes. The maintainers are not responsible for misuse.
๐ Future Additions
- Integration with CTF (Capture The Flag) challenge walkthroughs.
- More blue team procedures and SOAR playbooks.
- Anonymized real-world incident response logs (with permission).
๐ ๏ธ How to Use & Train
1๏ธโฃ Load Categorized JSONL Data
from datasets import load_dataset
REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"
# โ Load a single category โ
ds = load_dataset(REPO, split="train", data_files="data/coding/*/*.jsonl", streaming=True)
# โ Load a specific source โ
ds = load_dataset(REPO, split="train", data_files="data/coding/vibe_instruct_v2/*.jsonl", streaming=True)
# โ Load everything (18M+ samples) โ
ds = load_dataset(REPO, split="train", streaming=True)
for sample in ds:
print(sample["source"], sample["instruction"][:80])<br>
2๏ธโฃ Stream the 64 GB archives/ GitHub Repositories
from huggingface_hub import hf_hub_download
import tarfile
REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"
# โ Option A: Download & extract ONE repository โ
hf_hub_download(
repo_id=REPO,
repo_type="dataset",
filename="archives/0x101__lakewatch.tar.gz",
local_dir="./repos",
)
with tarfile.open("./repos/archives/0x101__lakewatch.tar.gz", "r:gz") as tar:
tar.extractall("./extracted/0x101__lakewatch")
# โ Option B: Stream files WITHOUT full extraction โ
def stream_repo_files(archive_name, max_files=100):
"""Stream file contents from tar.gz without extracting to disk."""
local_path = hf_hub_download(repo_id=REPO, repo_type="dataset", filename=archive_name)
with tarfile.open(local_path, "r:gz") as tar:
count = 0
for member in tar:
if member.isfile() and count < max_files:
f = tar.extractfile(member)
if f:
yield {
"path": member.name,
"content": f.read().decode("utf-8", errors="ignore")[:4000],
}
count += 1
import os
os.remove(local_path) # Clean up
# Stream files from a specific repo
for file_data in stream_repo_files("archives/0x101__lakewatch.tar.gz"):
print(f"๐ {file_data['path']}: {file_data['content'][:100]}...")
# โ Option C: Use pre-extracted JSONL shards (475K samples) โ
code_ds = load_dataset(
REPO, split="train",
data_files="data/coding/fable5_repos_full/*.jsonl",
streaming=True
)
# Each sample: instruction = "<repo>/<file>", response = "<content>"<br>
3๏ธโฃ SFT Training Script (Hugging Face Trainer)
import torch
from datasets import load_dataset
from transformers import (
AutoTokenizer,
AutoModelForCausalLM,
TrainingArguments,
Trainer,
DataCollatorForLanguageModeling,
)
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
# CONFIGURATION
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
MODEL_NAME = "meta-llama/Llama-3.1-8B"
DATASET_REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"
OUTPUT_DIR = "./sft-output"
MAX_SEQ_LEN = 2048
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
# LOAD MODEL & TOKENIZER
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
tokenizer.pad_token = tokenizer.eos_token
model = AutoModelForCausalLM.from_pretrained(
MODEL_NAME,
torch_dtype=torch.bfloat16,
device_map="auto",
attn_implementation="flash_attention_2",
)
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
# LOAD & FORMAT DATASET
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
def format_instruction(sample):
text = f"### Instruction:\n{sample['instruction']}\n\n### Response:\n{sample['response']}"
return {"text": text}
def tokenize(examples):
return tokenizer(
examples["text"],
truncation=True,
max_length=MAX_SEQ_LEN,
padding="max_length",
)
# Load coding category (use "data/**/*.jsonl" for full 18M+)
train_ds = load_dataset(
DATASET_REPO,
split="train",
data_files="data/coding/*/*.jsonl",
streaming=True,
)
train_ds = train_ds.map(format_instruction).filter(lambda x: len(x["text"]) > 0)
train_ds = train_ds.map(tokenize, batched=True)
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
# TRAIN
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
training_args = TrainingArguments(
output_dir=OUTPUT_DIR,
num_train_epochs=3,
per_device_train_batch_size=4,
gradient_accumulation_steps=4,
warmup_steps=500,
logging_steps=100,
save_steps=2000,
learning_rate=2e-5,
bf16=True,
gradient_checkpointing=True,
optim="adamw_torch",
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_ds,
data_collator=DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False),
)
trainer.train()
trainer.save_model(OUTPUT_DIR)<br>
4๏ธโฃ Curriculum Learning Across Categories
from datasets import load_dataset, interleave_datasets
REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"
# โ Phase 1: Foundation (math + science) โ
phase1_math = load_dataset(REPO, split="train", data_files="data/math/**/*.jsonl", streaming=True)
phase1_sci = load_dataset(REPO, split="train", data_files="data/science/**/*.jsonl", streaming=True)
phase1 = interleave_datasets([phase1_math, phase1_sci])
# โ Phase 2: Add coding traces โ
phase2 = load_dataset(REPO, split="train", data_files="data/coding/**/*.jsonl", streaming=True)
# โ Phase 3: Add distilled reasoning + cybersecurity โ
phase3_distilled = load_dataset(REPO, split="train", data_files="data/distilled/**/*.jsonl", streaming=True)
phase3_cyber = load_dataset(REPO, split="train", data_files="data/cybersecurity/**/*.jsonl", streaming=True)
phase3 = interleave_datasets([phase3_distilled, phase3_cyber])
# Train sequentially
# trainer.train(phase1) # epochs 0-1
# trainer.train(phase2) # epochs 1-2
# trainer.train(phase3) # epochs 2-3<br>
๐ Schema Reference
{
"source": "fable5_2m",
"source_dataset": "Crownelius/Complete-FABLE.5-traces-2M",
"instruction": "<the prompt / question / file path>",
"response": "<the completion / answer / file content>",
"category": "coding"
}๐ Licensing & Limitations
๐ License
The collection as a whole is released under the MIT License.
Each upstream dataset retains its original license. The source_dataset field on every row identifies the upstream โ look it up on Hugging Face to determine its specific license.
โ Intended Use Cases (Our Vision)
- Fine-tuning open-source LLMs for instruction following
- Training coding agents and code-completion models
- Reasoning chain distillation research
- Domain-specific adaptation (math, science, cybersecurity)
- Repository-scale context training (using
archives/)
โ Not Recommended For
- Deploying models without safety evaluation
- Generating harmful, biased, or deceptive content
- High-stakes domains (medical, legal, financial) without expert review
- Claiming models "know" facts โ this is distilled output, not ground truth
โ ๏ธ Limitations
- Field length cap:
instructionandresponsecapped at 4,000 characters. For full content, usearchives/. - Distillation artifacts: Samples are model-generated โ may contain hallucinations or biases.
- Partial recovery: A few upstream datasets (GODCoder variants, Genesisv1.1) had format errors and were partially recovered via raw JSONL parsing.
๐ Citation
@misc{open_distillation_codex_2026,
title = {The Open Distillation Codex: 18M+ samples + 7090 code repositories from 73 sources with Cybersecurity Attack & Defense},
author = {Manusagents},
year = {2026},
url = {https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset},
note = {v8.2 - No skip, full. 516 shards + 7090 archives, 73 sources, 8 categories, 76 GB+}
}๐ Changelog
<div align="center">
<br>
๐ The Open Distillation Codex ๐
73 sources ยท 8 categories ยท 7,090 repositories ยท 516 shards ยท 76 GB+
<br>
No skip. Full. 18M+ samples. Built one archive at a time. Released under MIT.
<br>
"Two layers. Eight categories. Seventy-three sources. One codex. No skip. Full. Armed with cybersecurity attack and defense."
<br>
<img src="https://img.shields.io/badge/Built%20with-Streaming%20Pipeline-blue?style=flat-square" alt="Streaming"> <img src="https://img.shields.io/badge/No-Skipping-brightgreen?style=flat-square" alt="No Skip"> <img src="https://img.shields.io/badge/Full%20Processing-success?style=flat-square" alt="Full"> <img src="https://img.shields.io/badge/Format-JSONL-orange?style=flat-square" alt="JSONL"> <img src="https://img.shields.io/badge/HuggingFace-Dataset-yellow?style=flat-square" alt="HF"> <img src="https://img.shields.io/badge/Cybersecurity-Deep%20Dive-purple?style=flat-square" alt="Cyber">
<br><br>
โ The Open Distillation Codex โ
</div>
