datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.cybersecurity-reasoning-cot-v1
🛡️ Expert Cybersecurity Reasoning Dataset (CoT)
This dataset contains 89 high-fidelity, expert-verified reasoning records focusing on complex cybersecurity attack vectors. It is designed specifically for fine-tuning Large Language Models (LLMs) on sophisticated security analysis and threat logic.
💎 Key Highlights
Niche Rarity 1.0: Covers rare and emerging threats with zero prior representation in open-source datasets.
Advanced Vectors: Includes detailed reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/cybersecurity-reasoning-cot-v1.cybersecurity-controls-instructions
Cybersecurity Controls Instructions
Security control, incident response and risk management guidance from NIST Special Publications, turned into instruction-following examples.
Splits
split
rows
source documents
train
13,106
56
validation
4,840
18
test
5,697
18
Splits are held out by source document. Every chunk yields several
instruction rows, so a random row-level split would place the same passage in
train and test; whole documents are held… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/cybersecurity-controls-instructions.cyber_security
Digital Literacy & Cybersecurity Nepali SFT Dataset
Dataset Overview
This dataset is a Nepali-language Supervised Fine-Tuning (SFT) dataset focused on digital literacy and cybersecurity.
The dataset contains 1,000 valid JSONL records designed for instruction-following tasks. Each record contains a human instruction and a corresponding GPT-generated response.
Dataset Statistics
Property
Value
Total records
1,000
Valid JSONL rows
1,000… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/cyber_security.
