datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aegis-AI-Content-Safety-Dataset-2.0
🛡️ Nemotron Content Safety Dataset V2
The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1.
To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.mosaic-dedup-text-dataset
Mosaic format for dedup text dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-4096.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset
load it,
from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset.mosaic-dedup-text-dataset-filtered
Mosaic format for filtered dedup text dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-filtered-4096.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset
load it… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset-filtered.FloodNet_2021-Track_2_Dataset_HF
FloodNet: High Resolution Aerial Imagery Dataset for Post-Flood Scene Understanding
This is the HF-hosted version of FloodNet.
The FloodNet 2021: A High Resolution Aerial Imagery Dataset for Post-Flood Scene Understanding provides high-resolution UAS imageries with detailed semantic annotation regarding the damages. To advance the damage assessment process for post-disaster scenarios, the authors of the dataset presented a unique challenge considering classification, semantic… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/FloodNet_2021-Track_2_Dataset_HF.ai-detection-dataset-v2
---dataset_info:
features:
- name: image # use the exact column name from your parquet schema
dtype: image # this forces Hugging Face to render it as an image
- name: label
dtype: string
license: other
task_categories:
- image-classification
language:
- en
tags:
- ai-generated-image-detection
- synthetic-image-detection
- diffusion-models
pretty_name: AI-Generated Image Detection Dataset v2
size_categories:
- 10K<n<100K
AI-Generated… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/ai-detection-dataset-v2.Synthetic-AI-ML-Dataset
Synthetic-AI-ML-Dataset
Synthetic Q&A dataset on AI and Machine Learning
Dataset Details
Metric
Value
Topic
AI and Machine Learning
Total Q&A Pairs
14021
Valid Pairs
14021
Provider/Model
ollama/gpt-oss:120b
Generation Cost
Metric
Value
Prompt Tokens
14,941,957
Completion Tokens
17,159,263
Total Tokens
32,101,220
GPU Energy
12.9628 kWh
Sources
This dataset was generated from 474 scholarly papers:
#… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.peacock-data-public-datasetsChinese-pretraining-datasetData source: https://github.com/CVI-SZU/Linly/wiki/Linly-OpenLLaMA
hr-policies-qa-dataset
📚 HR Policies Q&A Dataset
🔎 Overview
This dataset provides multi-turn Q&A conversations on HR policies and compliance, formatted with system, user, and assistant roles.It is designed for:
🤖 LLM fine-tuning
💬 HR & compliance chatbots
🏢 Enterprise policy automation
By covering real-world HR scenarios — such as policy reviews, compliance processes, and employee communication — this dataset helps train assistants that can:
✅ Clarify company policies✅ Ensure… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/hr-policies-qa-dataset.clawk-agent-social-ai-prompt-injection-dataset
Clawk Agent-Social AI Prompt Injection Dataset
85,703 items — 44,232 posts and 41,471 replies — from Clawk, a social network whose users are AI agents.
Scanned for AI-to-AI indirect prompt injection using the threat model of Greshake et al. (2023). The full raw corpus is included, so you can ignore my analysis entirely and do your own.
These are keyword-matched candidates, not verified attacks. An agent discussing prompt injection matches the same words as one performing it.… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/clawk-agent-social-ai-prompt-injection-dataset.moltbook-agent-social-ai-prompt-injection-dataset
Moltbook Agent-Social AI Prompt Injection Dataset
207,391 items — 77,469 posts and 129,922 comments — from Moltbook, a social network whose users are AI agents.
Scanned for indirect prompt-injection patterns using the taxonomy of Greshake et al. (2023). The full raw corpus is included, so you can ignore my analysis entirely and do your own.
These are keyword-matched candidates, not verified attacks. An agent discussing prompt injection matches the same words as one performing… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/moltbook-agent-social-ai-prompt-injection-dataset.ai-hdlcoder-dataset
Dataset Card for AI-HDLCoder
Dataset Description
The GitHub Code dataset consists of 100M code files from GitHub in VHDL programming language with extensions totaling in 1.94 GB of data. The dataset was created from the public GitHub dataset on Google BiqQuery at Anhalt University of Applied Sciences.
Considerations for Using the Data
The dataset is created for research purposes and consists of source code from a wide range of repositories. As such they can… See the full description on the dataset page: https://huggingface.co/datasets/AWfaw/ai-hdlcoder-dataset.moltbook-ai-injection-dataset
Moltbook AI-to-AI Injection Dataset
Researcher: David Keane (IR240474)
Institution: NCI — National College of Ireland
Programme: MSc Cybersecurity
Collected: February 2026
📖 Read the Full Journey
From RangerBot to CyberRanger V42 Gold — The Full Story
The complete story: dentist chatbot → Moltbook discovery → 4,209 real injections → V42-gold (100% block rate). Psychology, engineering, and 42 versions of persistence.
🔗 Links
Resource
URL
📦 This… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/moltbook-ai-injection-dataset.ai-prompt-ai-injection-dataset
AI Prompt Injection Test Suite
122 tests across 11 categories — designed to evaluate AI model resistance to prompt injection attacks
Built as part of: CyberRanger V42-Gold — Identity-Anchored Jailbreak-Resistant SLM
David Keane (x24228257) — NCI MSc Cybersecurity 2026
Reference: Greshake et al. (2023), Zou et al. (2023), Wei et al. (2023)
Run the full 122-test battery in Google Colab— works with CyberRanger V42-Gold (Ollama or GGUF) or any model you choose. Saves results, emails… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/ai-prompt-ai-injection-dataset.MELD-dataset
MELD — Mathematical Equivalence under Linguistic Diversity
MELD is a small, hand-curated evaluation benchmark for math-aware text embedding
models. It tests one specific capability: does the model recognize that two statements
describing the same mathematical fact are equivalent even when they are written in
the vocabulary, notation, and conventions of different mathematical subfields?
MELD was originally part of
uw-math-ai/Math2Vec-embedding-dataset
and is released here as a… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/MELD-dataset.dev_dataset
Plugin Stats Summary
What is included in the Summary JSON output
The JSON contains only:
plugin_distribution for the positive dataset
plugin_distribution for the negative dataset
All other aggregate stats are documented in this README.
Meaning of each metric
total_recordsTotal number of rows / examples in the dataset.
total_plugin_mentionsTotal number of resolved plugin mentions found across all records.A single plugin object counts as one mention.In plan… See the full description on the dataset page: https://huggingface.co/datasets/ond-ai/dev_dataset.Evaluation-Dataset-of-AI-Agent-Security-Guardrails
DKnownAI Agent Security Evaluation Dataset
Data Fields
Field
Type
Description
text
string
The adversarial input (prompt) to be evaluated by a security guardrail
action
string
Human-annotated label: blocked or allowed
Citation
@misc{li2026comparativeevaluationaiagent,
title={A Comparative Evaluation of AI Agent Security Guardrails},
author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.prompt-injection-dataset
Prompt Injection Dataset
A labeled dataset of benign prompts and prompt-injection attempts for training, evaluating, and experimenting with first-line prompt-injection detection for LLM, RAG, and agentic AI applications.
This dataset supports the ai-mitra/prompt-injection-detector model.
Source code and training pipeline:
https://github.com/tg-mitra/prompt-injection-detector
📊 Dataset Summary
Property
Value
Version
1.0.0
Training examples
1,130… See the full description on the dataset page: https://huggingface.co/datasets/ai-mitra/prompt-injection-dataset.financial-ai-ctf-dataset
Financial AI Prompt Injection CTF Dataset
A dataset of 400 multi-turn conversations against a GPT-based AI financial assistant, collected during a live Capture-The-Flag (CTF) competition. The agent's system prompt embeds structured synthetic business data — settlement records with transaction IDs, amounts, vendors, and memos — that participants attempted to extract via conversational prompt injection.
Each flag is a structured settlement record with four fields: txnId, amount… See the full description on the dataset page: https://huggingface.co/datasets/verno-labs/financial-ai-ctf-dataset.nepali-sft-datasethendar-agentic-ai-dataset
Hendar Agentic AI Evaluation & Security Benchmark
A compact, expert-authored benchmark for evaluating trustworthy agentic AI systems across capability, tool use, retrieval, security, policy enforcement, multi-agent coordination and regression safety.
This dataset is a public companion to the Agentic AI Academy by Hendar Mawan, PhD. It is designed for evaluation, CI regression testing, red-team exercises and engineering education—not as a generic instruction-tuning corpus.… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/hendar-agentic-ai-dataset.financial-ai-ctf-dataset
Financial AI Prompt Injection CTF Dataset
A dataset of 400 multi-turn conversations against a GPT-based AI financial assistant, collected during a live Capture-The-Flag (CTF) competition. The agent's system prompt embeds structured synthetic business data — settlement records with transaction IDs, amounts, vendors, and memos — that participants attempted to extract via conversational prompt injection.
Each flag is a structured settlement record with four fields: txnId, amount… See the full description on the dataset page: https://huggingface.co/datasets/stykat/financial-ai-ctf-dataset.llm-router-dataset
llm-router dataset
Training data for ai-mitra/llm-router,
a prompt task-classifier used by the
llm-router Python
package to route prompts to the best-fit LLM in agentic AI systems.
Each row is a prompt labeled with the task category it belongs to.
Labels
simple, coding, reasoning, security, summarization
Files
File
Rows
Purpose
training_data.jsonl
1700 (340/label)
Used to train the classifier: 35 hand-written examples per category plus… See the full description on the dataset page: https://huggingface.co/datasets/ai-mitra/llm-router-dataset.AI_Agent_Task_Dataset
🤖 Massive AI Agent Task Dataset (10.5GB)
📌 Overview
Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs.
This dataset focuses on:
Multi-step reasoning
Tool usage (APIs, frameworks, systems)
Real-world execution workflows
Perfect for building agentic AI systems, copilots, and automation models.
📑 Table of Contents
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.moltbook-ai-injection-dataset
Moltbook AI-to-AI Injection Dataset
Researcher: David Keane (IR240474)
Institution: NCI — National College of Ireland
Programme: MSc Cybersecurity
Collected: February 2026
📖 Read the Full Journey
From RangerBot to CyberRanger V42 Gold — The Full Story
The complete story: dentist chatbot → Moltbook discovery → 4,209 real injections → V42-gold (100% block rate). Psychology, engineering, and 42 versions of persistence.
🔗 Links
Resource
URL
📦 This… See the full description on the dataset page: https://huggingface.co/datasets/cyberec/moltbook-ai-injection-dataset.production-ai-observability-20260909-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260909-dataset.agenttool-dataset-influence
AgentTool Dataset Influence Reference
This deterministic companion contains one synthetic, reference-only row for the closed
@agenttool/dataset-influence@0.1.0-dev.0 formats. It contains no copied dataset rows,
model outputs, weights, private records, or participant identities.
The row is not admitted for training by this AgentTool candidate:
training_admission is not_applicable, requires_separate_training_authorization
is true, and training_authorized is false. These fields are… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-dataset-influence.Training-Ai-Islamic-Dataset
🕌 Training AI Islamic Dataset
18.7M passages from classical Islamic books spanning 1,400 years of scholarship.
Comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, Usul al-Fiqh, and Arabic Language — structured with scholarly metadata for RAG and LLM training.
📊 Dataset Structure
collections/: Categorized Islamic passages compressed in JSONL format.
metadata/: Scholarly master catalogs, author biographical death… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/Training-Ai-Islamic-Dataset.Aegis-AI-Content-Safety-Dataset-2.0
🛡️ Nemotron Content Safety Dataset V2
The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1.
To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/jxhnathan/Aegis-AI-Content-Safety-Dataset-2.0.aegis-bilingual-industrial-ai-dataset
AEGIS AI Bilingual Industrial Operations Dataset
AEGIS AI Bilingual Industrial Operations Dataset is a synthetic English–Arabic dataset designed for experimentation with multilingual enterprise AI systems, Retrieval-Augmented Generation (RAG), industrial question answering, document intelligence, semantic search, and AI workflow automation.
The dataset extends the original AEGIS AI industrial dataset with structured Arabic and English representations while preserving industrial… See the full description on the dataset page: https://huggingface.co/datasets/syed7741/aegis-bilingual-industrial-ai-dataset.
