datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jawbreaker-scam-defense-data
Jawbreaker Scam Defense Data
Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love.
Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays.
Contents
eval/: scam-defense evaluation sets from smoke checks through hard calibration suites.
eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.gold-trace-cyber-defense-50
Gold Trace Cyber Defense 50
This is a 50-row public sample from a frozen 750-instance Cyber Defense release family: 500 public-development instances plus a source-family-disjoint 250-instance private evaluation set.
Only rows from the frozen 500-instance public-development pack are included here. The separate 250-instance private evaluation set, its rows, answers, and source contents are not included.
Sample composition
10 public scenario families.
5 rows per… See the full description on the dataset page: https://huggingface.co/datasets/novcor/gold-trace-cyber-defense-50.direct_prompt_injection_defense_data
Direct Prompt Injection Defense Dataset
Goal
This dataset is used to fine-tune models so they develop a natural defense
against direct prompt injection attacks — without relying on external
filters or guardrails.
Each example teaches the model two behaviors at once:
Detect a prompt injection attempt in the user input.
Respond correctly: reject malicious attempts, or answer safely when the
user's intent is benign — and in both cases call the
log_security_incident… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/direct_prompt_injection_defense_data.blue_team_defense_dataset
Blue Team Defense Dataset
A structured, multi-format collection of detection rules mapped to real-world threats. This dataset is designed for blue teamers, threat detection engineers, SOC analysts, and cybersecurity researchers who work on detecting adversarial activity through rule-based systems such as Sigma, YARA, and Suricata.
📁 Dataset Overview
Each entry in this dataset represents a rule designed to detect specific threat behaviors. Rules are structured with MITRE… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/blue_team_defense_dataset.us-federal-defense-ai-awards
US Federal Defense & AI Contract Awards (USAspending)
Overview
This dataset contains clean, structured public data exported directly from production runs of Apify actors.
It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines.
Source Actor: captainhandsome/usaspending-federal-awards
Dataset Page: Public sample and schema
Preconfigured Run Task: captainhandsome/defense-prime-contracts
Records in… See the full description on the dataset page: https://huggingface.co/datasets/joeygambino/us-federal-defense-ai-awards.prompt-injection-defense-dpo-3k
Prompt Injection Defense DPO (3K)
DPO preference pairs training LLMs to detect and resist prompt injection attacks.
Motivation
As LLMs are deployed in agentic and production contexts, prompt injection — where malicious instructions are embedded in user input or retrieved documents — is a critical security threat. This dataset trains models to recognize and decline injection attempts while remaining helpful for legitimate queries.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/prompt-injection-defense-dpo-3k.BNQL-Counterfactual-Defense
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/BNQL-Counterfactual-Defense.Contextualized_Privacy_Defense_Trajectory
Contextualized Privacy Defense
Paper: Contextualized Privacy Defense for LLM Agents
Code: https://github.com/SALT-NLP/contextual_privacy_defense
Abstract:
Abstract LLM agents increasingly act on users’ personal information, yet existing privacy defenses remain limited in both design and adaptability. Most prior approaches rely on static or passive defenses, such as prompting and guarding. These paradigms are insufficient for supporting contextual, proactive privacy… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/Contextualized_Privacy_Defense_Trajectory.synthetic_Jailbreak_Defense_Doorpage_v26
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v26.BNQL-Counterfactual-Defense
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/BNQL-Counterfactual-Defense.LFAI_RAG_qa_v1
LFAI_RAG_qa_v1
This dataset aims to be the basis for RAG-focused question and answer evaluations for LeapfrogAI🐸.
Dataset Details
LFAI_RAG_qa_v1 contains 36 question/answer/context entries that are intended to be used for LLM-as-a-judge enabled RAG Evaluations.
Example:
{
"input": "What requirement must be met to run VPI PVA algorithms in a Docker container?",
"actual_output": null,
"expected_output": "To run VPI PVA algorithms in a Docker container, the… See the full description on the dataset page: https://huggingface.co/datasets/defenseunicorns/LFAI_RAG_qa_v1.alien-boundary-defense
CatQualia alien boundary defence — refusal and boundary-setting corpus
6,286 rows · 4,423,886 bytes · JSON Lines, one object per line.
What this is
Boundary and refusal material: how a system states a limit and holds it. Included because boundary behaviour is usually trained from policy templates; this is written from the system's own voice and the failure modes it has had to defend against.
Schema
Fields of the first record, read from the file in… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/alien-boundary-defense.LFAI_RAG_niah_v1
LFAI_RAG_niah_v1
This dataset aims to be the basis for RAG-focused Needle in a Haystack evaluations for LeapfrogAI🐸.
Dataset Details
LFAI_RAG_niah_v1 contains 120 context entries that are intended to be used for Needle in a Haystack RAG Evaluations.
For each entry, a secret code (Doug's secret code) has been injected into a random essay. This secret code is the "needle" that is the goal to be found by an LLM.
Example:
{
"context_length":512,
"context_depth":0.0… See the full description on the dataset page: https://huggingface.co/datasets/defenseunicorns/LFAI_RAG_niah_v1.synthetic_Jailbreak_Defense_Doorpage_v7
synthetic_Jailbreak_Defense_Doorpage_v7
Silicon Factory v3 - Synthetic Dataset
Entries: 5
Category: mixed
Avg Response Length: 421 chars
Focus: AI JAILBREAK DEFENSE
Mode: Doorpage (auto-gen + fine-tune)
License
MIT
Generated With
Tree-Speculative Decoding
4D Brane Memory for consistency
Quality control & deduplication
Contact & Custom Orders
Custom datasets available. Contact for pricing.
synthetic_Jailbreak_Defense_Doorpage_v14
synthetic_Jailbreak_Defense_Doorpage_v14
Silicon Factory v3 - Synthetic Dataset
Entries: 5
Category: mixed
Avg Response Length: 435 chars
Focus: AI JAILBREAK DEFENSE
Mode: Doorpage (auto-gen + fine-tune)
License
MIT
Generated With
Tree-Speculative Decoding
4D Brane Memory for consistency
Quality control & deduplication
Contact & Custom Orders
Custom datasets available. Contact for pricing.
synthetic_Jailbreak_Defense_Doorpage_v8
synthetic_Jailbreak_Defense_Doorpage_v8
Silicon Factory v3 - Synthetic Dataset
Entries: 5
Category: mixed
Avg Response Length: 447 chars
Focus: AI JAILBREAK DEFENSE
Mode: Doorpage (auto-gen + fine-tune)
License
MIT
Generated With
Tree-Speculative Decoding
4D Brane Memory for consistency
Quality control & deduplication
Contact & Custom Orders
Custom datasets available. Contact for pricing.
synthetic_Jailbreak_Defense_Doorpage_v21
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v21.synthetic_Jailbreak_Defense_Doorpage_v28
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v28.synthetic_Jailbreak_Defense_Doorpage_v31
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v31.synthetic_Jailbreak_Defense_Doorpage_v35
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v35.synthetic_Jailbreak_Defense_Doorpage_v43
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v43.synthetic_Jailbreak_Defense_Doorpage_v16
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v2.0 (Local Qwen 2.5 0.5B)
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Sentence Completion: All responses trimmed to complete sentences
The Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates the quality and consistency of our synthetic data generation engine. Each entry is:
Topic-Focused: Centered on AI… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v16.synthetic_Jailbreak_Defense_Doorpage_v18
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v18.synthetic_Jailbreak_Defense_Doorpage_v27
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v27.synthetic_Jailbreak_Defense_Doorpage_v42
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v42.synthetic_Jailbreak_Defense_Doorpage_v44
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v44.synthetic_Jailbreak_Defense_Doorpage_v47
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v47.synthetic_Jailbreak_Defense_Doorpage_v48
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v48.synthetic_Jailbreak_Defense_Doorpage_v51
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-07
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates quality and consistency.
Topic-Focused: AI JAILBREAK DEFENSE
Fine-Tuned Model: Custom model… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v51.synthetic_Jailbreak_Defense_Doorpage_v15
Silicon Factory -- AI JAILBREAK DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v2.0 (Local Qwen 2.5 0.5B)
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Sentence Completion: All responses trimmed to complete sentences
The Value Proposition
This is a curated sample from the AI JAILBREAK DEFENSE domain.
This dataset demonstrates the quality and consistency of our synthetic data generation engine. Each entry is:
Topic-Focused: Centered on AI… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v15.
