datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AI_threat_and_vulnerabilities_taxonomy
AI Threat Taxonomy and Vulnerability Registry
Dataset Description
This dataset provides a practitioner-grade, machine-readable taxonomy of AI
threat vectors and AI vulnerabilities for use in red teaming, risk assessment,
model governance, and regulatory compliance programs.
The taxonomy distinguishes precisely between:
Threat vectors: paths or mechanisms an attacker, insider, or negligent
actor uses to exploit an AI environment
Vulnerabilities: weaknesses in… See the full description on the dataset page: https://huggingface.co/datasets/hewyler/AI_threat_and_vulnerabilities_taxonomy.synthetic-code-vulnerabilities-1synthetic-code-vulnerabilities-1 is a synthetic dataset with a total of ~493 Question and Answer pairs.
This dataset was generated using the following models:
Gemini:
Fast
Thinking
Pro
ChatGPT:
Whatever is available on the website
Deepseek:
"Instant"
"Expert"
Grok:
Fast
Qwen 3.6:
Fast
Thinking
Perplexity.ai:
Whatever is available on the website
This dataset follows the following format:
[
{"in":"Prompt","out":"Response"},
{"in":"Prompt","out":"Response"}
]
synthetic-code-vulnerabilities-2synthetic-code-vulnerabilities-2 is a synthetic dataset with a total of ~884 Question and Answer pairs.
This dataset was generated using the following models:
ChatGPT:
Whatever is available on the website
OSS 120B
Gemini:
Pro
Deepseek:
"Instant"
"Expert"
Grok:
Fast
Qwen3:
Coder
This dataset follows the following format:
[
{"messages": [
{"role": "system", "content": "Example system prompt"},
{"role": "user", "content": "Example user prompt"},
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/takenusername32/synthetic-code-vulnerabilities-2.
