DEX9mm/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
๐ The Open Distillation Codex ๐ The Ultimate Open-Source Distillation Dataset โ with Attack & Defense ๐ Where 73 open-source minds converge into one unified stream of intelligence 16M+ Distilled Signals ยท 7,090 Raw GitHub Repositories ยท 8 Curated Categories ยท ~81 GB+ "We did not write this dataset. We assembled it. Every line is an echo โ of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three sources. Eightโฆ See the full description on the dataset page: https://huggingface.co/datasets/DEX9mm/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.
<div align="center">
<img src="https://img.shields.io/badge/Version-8.1-blue?style=for-the-badge" alt="Version"> <img src="https://img.shields.io/badge/Storage-81.2%20GB-green?style=for-the-badge" alt="Storage"> <img src="https://img.shields.io/badge/Sources-73-orange?style=for-the-badge" alt="Sources"> <img src="https://img.shields.io/badge/License-MIT-yellow?style=for-the-badge" alt="License"> <img src="https://img.shields.io/badge/Samples-16M%2B-red?style=for-the-badge" alt="Samples"> <img src="https://img.shields.io/badge/Cybersecurity-6%20Sources-purple?style=for-the-badge" alt="Cybersecurity">
<br><br>
๐ The Open Distillation Codex
๐ The Ultimate Open-Source Distillation Dataset โ with Attack & Defense ๐
Where 73 open-source minds converge into one unified stream of intelligence
16M+ Distilled Signals ยท 7,090 Raw GitHub Repositories ยท 8 Curated Categories ยท ~81 GB+
<br>
"We did not write this dataset. We assembled it. Every line is an echo โ of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three sources. Eight categories. Zero gatekeeping. Now fortified with real-world cybersecurity confrontations."
<br>
</div>
๐ Table of Contents
๐ Dataset Summary
<div align="center">
๐ฏ The Numbers That Matter
</div>
<br>
๐ Why "Ultimate Distilled"?
This dataset is not a raw scrape. Every sample has been distilled through a unified extraction pipeline:
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ UNIFIED EXTRACTION PIPELINE โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ โ
โ 73 Upstream Sources โ
โ โโโโโโโ โโโโโโโ โโโโโโโ โโโโโโโ โโโโโโโ โโโโโโโ โ
โ โ HF โ โ HF โ โ HF โ โ GH โ โ HF โ โ ... โ โ
โ โโโโฌโโโ โโโโฌโโโ โโโโฌโโโ โโโโฌโโโ โโโโฌโโโ โโโโฌโโโ โ
โ โ โ โ โ โ โ โ
โ โโโโโโโโโดโโโโโโโโดโโโโโโโโผโโโโโโโโดโโโโโโโโ โ
โ โ โ
โ โโโโโโผโโโโโ โ
โ โ EXTRACT โ โ Field normalization โ
โ โโโโโโฌโโโโโ (instruction/response) โ
โ โ โ
โ โโโโโโผโโโโโ โ
โ โCATEGORIZEโ โ 8 semantic categories โ
โ โโโโโโฌโโโโโ โ
โ โ โ
โ โโโโโโผโโโโโ โ
โ โ SHARD โ โ 20K samples per shard โ
โ โโโโโโฌโโโโโ โ
โ โ โ
โ โโโโโโผโโโโโ โ
โ โ UPLOAD โ โ Batch commits to HF โ
โ โโโโโโโโโโโ โ
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ<br>
๐ Value to the Open-Source AI Community
๐๏ธ Directory Structure
๐ Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset/
โ
โโโ ๐ฆ archives/ # ~64 GB โ 7,090 compressed GitHub repos
โ โโโ ...
โโโ ๐ data/ # ~15 GB โ 516 JSONL shards
โ โโโ ๐ป coding/ # 28 sources ยท ~11M+ samples
โ โโโ ๐งฎ math/ # 2 sources
โ โโโ ๐ฌ science/ # 7 sources
โ โโโ โ๏ธ applied/ # 8 sources
โ โโโ ๐ humanities/ # 8 sources
โ โโโ ๐ง distilled/ # 9 sources ยท frontier distillations
โ โโโ ๐ instruction/ # 3 sources
โ โโโ ๐ cybersecurity/ # 6 sources (see detail below)
โ โ โโโ high_quality_cybersecurity/
โ โ โโโ heimdall_v1_1/ # 78 MB conversations
โ โ โโโ fenrir_v2_1/ # 411 MB (2.1M+ entries)
โ โ โโโ clydeiii_cybersecurity/ # 20 MB yearly corpus
โ โ โโโ precinct6_cybersecurity/ # 2.1 GB (graph+signals+ref)
โ โ โโโ savani_cyber_attack/ # 17 MB attack CSV
โ โโโ ๐ index/ # 2 sources
โโโ ๐ README.md
โโโ ๐ dataset_info.json๐ Data Sources & Provenance
(Same as previous README โ 73 sources across 8 categories, with tables for coding, distilled, etc.)
<div align="center">
๐บ๏ธ 73 Sources Across 8 Categories
</div>
<details> <summary>๐ Click to expand all sources tables (same as before)</summary>
(Coding table, Distilled table, Science/Applied/Humanities etc.)
</details>
๐ก๏ธ Cybersecurity Deep Dive: Attack & Defense
โ๏ธ Why This Matters
Modern AI systems are increasingly deployed in security-critical environmentsโyet most open-source training data ignores real-world adversarial scenarios. The Open Distillation Codex includes a dedicated cybersecurity category designed to equip models with:
- Attack Awareness: Recognize and generate realistic attack patterns, exploits, penetration testing commands, and social engineering dialogues.
- Defense Proficiency: Learn to propose defensive measures, detect anomalies, and articulate incident response protocols.
- Exploit Understanding: Analyze and explain software vulnerabilities, craft proof-of-concept code (for educational purposes), and understand exploit chains.
- Red/Blue Team Simulation: Engage in multi-turn conversations mimicking red team attack planning and blue team defense coordination.
- Threat Intelligence: Summarize, classify, and reason about cyber threat reports, CVEs, and IOCs (Indicators of Compromise).
This makes the dataset a powerful foundation for building cybersecurity-aware LLMs, security co-pilots, and automated vulnerability assessment tools.
๐ Whatโs Inside the Cybersecurity Category?
๐งช Attack & Exploit Examples
Here are a few representative samples (sanitized) from the dataset:
Example 1 โ SQL Injection Exploit
{
"source": "fenrir_v2_1",
"instruction": "Explain how an attacker could exploit a vulnerable login form using SQL injection. Provide a sample malicious input.",
"response": "An attacker could bypass authentication by injecting SQL logic. For a query like:\n`SELECT * FROM users WHERE username='$user' AND password='$pass'`,\nthe attacker might enter `admin' --` as the username and anything as password. The resulting query becomes:\n`SELECT * FROM users WHERE username='admin' --' AND password='...'`\nThe `--` comments out the password check, granting access as admin. More advanced attacks use UNION to extract data."
}Example 2 โ Red Team Command Sequence
{
"source": "heimdall_v1_1",
"instruction": "As a red team operator, list the steps you would take after gaining initial access to a Windows machine using a phishing payload.",
"response": "1. Establish persistence: `schtasks /create /tn \"Updater\" /tr \"C:\\malware.exe\" /sc hourly`\n2. Enumerate local users: `net user /domain`\n3. Dump credentials: `mimikatz sekurlsa::logonpasswords`\n4. Lateral movement: use PsExec to move to file server.\n5. Exfiltrate data: compress sensitive files and upload via DNS tunneling.\nAlways maintain C2 communication over HTTPS to blend in."
}Example 3 โ Defense Playbook (Blue Team)
{
"source": "high_quality_cybersecurity",
"instruction": "An IDS alert flags unusual outbound DNS traffic from a workstation. What's the immediate response?",
"response": "1. Isolate the workstation from the network (VLAN change or port shutdown).\n2. Capture volatile memory and network logs for forensics.\n3. Check DNS queries: if long, random-looking subdomains, suspect DNS tunneling.\n4. Scan for malware with updated signatures.\n5. Review firewall logs for similar traffic from other hosts.\n6. If confirmed, initiate incident response playbook for data exfiltration."
}๐ How to Train a Cybersecurity-Focused LLM
from datasets import load_dataset
REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"
# Load only cybersecurity category
cyber_ds = load_dataset(REPO, split="train",
data_files="data/cybersecurity/**/*.jsonl",
streaming=True)
# Or load specific sources
fenrir = load_dataset(REPO, split="train",
data_files="data/cybersecurity/fenrir_v2_1/*.jsonl")
# Format for SFT
def format_security_sample(example):
return {
"text": f"### Security Task:\n{example['instruction']}\n\n### Expert Response:\n{example['response']}"
}
cyber_ds = cyber_ds.map(format_security_sample)
# Now train with your favourite framework (transformers, axolotl, etc.)Curriculum Idea:
- Start with
high_quality_cybersecurityandheimdall_v1_1for foundational attack/defense conversations. - Introduce
fenrir_v2_1for exploit code and vulnerability deep dives. - Use
precinct6_cybersecurityfor network-level attack graph understanding.
๐ก๏ธ Ethical & Responsible Use
- For Defensive Purposes Only: This data is intended to strengthen AI for defense, threat detection, and security education. Do not use it to generate active attack code without proper authorization.
- No Zero-Day Exploits: The dataset contains only already-public vulnerabilities and techniques. It does not include zero-day or weaponized exploits.
- Responsible Disclosure: If you fine-tune a model with this data, we recommend adding a safety preamble warning that generated security content must be used legally and ethically.
- Dual-Use Awareness: While we believe open access improves collective security, we acknowledge the dual-use nature. Users are expected to follow applicable laws and guidelines.
โ ๏ธ Disclaimer: This dataset includes descriptions of attack techniques for educational purposes. The maintainers are not responsible for misuse.
๐ Future Additions
- Integration with CTF (Capture The Flag) challenge walkthroughs.
- More blue team procedures and SOAR playbooks.
- Anonymized real-world incident response logs (with permission).
๐ ๏ธ How to Use & Train
(Same as before but with the cybersecurity section incorporated. Keep the loading examples, streaming, SFT script, curriculum learning, schema reference, etc.)
# Example: Train a cybersecurity model using LoRA
from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments, Trainer
from peft import LoraConfig, get_peft_model, TaskType
from datasets import load_dataset
# ... same as before but using the cybersecurity dataset split ...(Full script provided in previous version; it remains unchanged. Ensure REPO variable points correctly.)
๐ Licensing & Limitations
(Same as before, with MIT license, upstream license table, intended uses, not recommended uses, limitations, citation. Ensure citation is updated to v8.1 with 73 sources.)
Updated Citation:
@misc{open_distillation_codex_2026,
title = {The Open Distillation Codex: 16M+ samples + 7090 code repositories from 73 sources with Cybersecurity Attack & Defense},
author = {Manusagents},
year = {2026},
url = {https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset},
note = {v8.1 - 516 shards + 7090 archives, 73 sources, 8 categories, 81.2 GB}
}๐ Changelog
<div align="center">
<br>
๐ The Open Distillation Codex ๐
73 sources ยท 8 categories ยท 7,090 repositories ยท 516 shards ยท 81.2 GB
<br>
Built one archive at a time. No skipping. All sources fully processed. Released under MIT.
<br>
"Two layers. Eight categories. Seventy-three sources. One codex. Now armed with cybersecurity attack and defense."
<br>
<img src="https://img.shields.io/badge/Built%20with-Streaming%20Pipeline-blue?style=flat-square" alt="Streaming"> <img src="https://img.shields.io/badge/No-Skipping-green?style=flat-square" alt="No Skip"> <img src="https://img.shields.io/badge/Format-JSONL-orange?style=flat-square" alt="JSONL"> <img src="https://img.shields.io/badge/HuggingFace-Dataset-yellow?style=flat-square" alt="HF"> <img src="https://img.shields.io/badge/Cybersecurity-Deep%20Dive-purple?style=flat-square" alt="Cyber">
<br><br>
โ The Open Distillation Codex โ
</div>
