nyxspecter4/kin-cyber-dpo-v2
KIN Cybersecurity DPO v2 Preference Dataset Empirically mined and zero-leak sanitized preference dataset for training cybersecurity and agentic code repair models. Dataset Summary Total DPO Pairs: 1,635 (Updated 2026-09-07) Baseline v4 pairs: 1,495 Hermetic expansion (v5): +140 pairs (AST-invariant vulnerability repair, CWE-79 XSS guards, CWE-89 SQLi, CWE-22 Path Traversal, and supply chain integrity) Format: Direct Preference Optimization (DPO) schema: {prompt… See the full description on the dataset page: https://huggingface.co/datasets/nyxspecter4/kin-cyber-dpo-v2.
fix: sync clean cross-links to canonical model and GGUF
fix: sync clean links to canonical 3B model and GGUF
fix: sync clean links to canonical 3B model and GGUF
#898 deploy upgraded dataset card
docs: use top-level configs schema (viewer multi-subset), 5 configs: dpo (default), dpo-clean, train, sft, grpo-prompts
docs: fix dataset_info features to YAML list form (viewer-compatible), add sft config (1,495 records)
chore: remove internal training log from public dataset
docs: add dataset viewer config (per-file configs + features), correct record counts (dpo 1,637 / clean 432 / train 431 / grpo 15), remove internal training log
docs: clean up companion links and add schema documentation
Updated dataset card with v5 expansion details
docs(dataset): configure default config to dpo.jsonl (1,637 pairs) and eliminate 404 links
Updated dataset card with v5 expansion details
feat(rlhf): append 2 real-world human-in-the-loop DPO pairs from #huntr (Ghost Tools capability mapping) and #github-reviews (empirical branch audit vs sycophancy)
feat(rlhf): append real-world human-in-the-loop pairs from #huntr and #github-reviews
Ship Monday v5 expansion: 1,495 -> 1,635 DPO pairs (+140 AST-invariant pairs)
Phase 3: GRPO prompts for verifiable reward training
Add GRPO training prompts for Phase 3
v5 sft.jsonl update
v5: added 24 new DPO pairs (vuln-finding, exploit-chain, CVE analysis)
Upload zero-leak DPO v2 preference dataset for cybersecurity model fine-tuning
docs(readme): v2 ladder push (kin-cybersec-suite)
diag: v5 training run log
v4: 1471 DPO + 1471 SFT pairs
Expand v2: 1331 DPO pairs + 1331 SFT pairs (48 CVEs + 30 MITRE + 20 concepts)
Upload zero-leak DPO v2 preference dataset for cybersecurity model fine-tuning
Expand v2: 1331 DPO pairs + 1331 SFT pairs (48 CVEs + 30 MITRE + 20 concepts)
Upload zero-leak DPO v2 preference dataset for cybersecurity model fine-tuning
Expand v2: 1331 DPO pairs + 1331 SFT pairs (48 CVEs + 30 MITRE + 20 concepts)
Fix: restore sft.jsonl (939 pairs)
Fix: restore train.jsonl (939 pairs)
Fix: restore dpo.jsonl (939 pairs)
Expand v2: 1331 DPO pairs + 1331 SFT pairs (48 CVEs + 30 MITRE + 20 concepts)
Fix: restore sft.jsonl (939 pairs)
Fix: restore train.jsonl (939 pairs)
Fix: restore dpo.jsonl (939 pairs)
Restore + expand: 949 DPO pairs + 0 SFT pairs
Expand v3: 0 instruction + 0 DPO pairs from NVD JSON feeds
Expand training data: 0 instruction pairs + 0 DPO pairs from NVD
Upload zero-leak DPO v2 preference dataset for cybersecurity model fine-tuning
Upload zero-leak DPO v2 preference dataset for cybersecurity model fine-tuning
initial commit
