CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K110 likes8.5k downloads1y agoHugging Face02malaysia-ai /mosaic-dedup-text-dataset Mosaic format for dedup text dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-4096.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset load it, from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset.textn<1K0 likes2.9k downloads3y agoHugging Face03malaysia-ai /mosaic-dedup-text-dataset-filtered Mosaic format for filtered dedup text dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-filtered-4096.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset load it… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset-filtered.textn<1K0 likes1.9k downloads3y agoHugging Face04takara-ai /FloodNet_2021-Track_2_Dataset_HF FloodNet: High Resolution Aerial Imagery Dataset for Post-Flood Scene Understanding This is the HF-hosted version of FloodNet. The FloodNet 2021: A High Resolution Aerial Imagery Dataset for Post-Flood Scene Understanding provides high-resolution UAS imageries with detailed semantic annotation regarding the damages. To advance the damage assessment process for post-disaster scenarios, the authors of the dataset presented a unique challenge considering classification, semantic… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/FloodNet_2021-Track_2_Dataset_HF.imagevisual-question-answering1K<n<10K6 likes646 downloads2y agoHugging Face05Shanmuk4622 /ai-detection-dataset-v2 ---dataset_info: features: - name: image # use the exact column name from your parquet schema dtype: image # this forces Hugging Face to render it as an image - name: label dtype: string license: other task_categories: - image-classification language: - en tags: - ai-generated-image-detection - synthetic-image-detection - diffusion-models pretty_name: AI-Generated Image Detection Dataset v2 size_categories: - 10K<n<100K AI-Generated… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/ai-detection-dataset-v2.textn<1K0 likes516 downloads3mo agoHugging Face06Kylan12 /Synthetic-AI-ML-Dataset Synthetic-AI-ML-Dataset Synthetic Q&A dataset on AI and Machine Learning Dataset Details Metric Value Topic AI and Machine Learning Total Q&A Pairs 14021 Valid Pairs 14021 Provider/Model ollama/gpt-oss:120b Generation Cost Metric Value Prompt Tokens 14,941,957 Completion Tokens 17,159,263 Total Tokens 32,101,220 GPU Energy 12.9628 kWh Sources This dataset was generated from 474 scholarly papers: #… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.textquestion-answering10K<n<100K2 likes473 downloads6mo agoHugging Face07applied-ai-018 /peacock-data-public-datasetstext1K<n<10K0 likes400 downloads2y agoHugging Face08Linly-AI /Chinese-pretraining-datasetData source: https://github.com/CVI-SZU/Linly/wiki/Linly-OpenLLaMA text10M<n<100M44 likes300 downloads3y agoHugging Face09strova-ai /hr-policies-qa-dataset 📚 HR Policies Q&A Dataset 🔎 Overview This dataset provides multi-turn Q&A conversations on HR policies and compliance, formatted with system, user, and assistant roles.It is designed for: 🤖 LLM fine-tuning 💬 HR & compliance chatbots 🏢 Enterprise policy automation By covering real-world HR scenarios — such as policy reviews, compliance processes, and employee communication — this dataset helps train assistants that can: ✅ Clarify company policies✅ Ensure… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/hr-policies-qa-dataset.textn<1K0 likes264 downloads1y agoHugging Face10DavidTKeane /clawk-agent-social-ai-prompt-injection-dataset Clawk Agent-Social AI Prompt Injection Dataset 85,703 items — 44,232 posts and 41,471 replies — from Clawk, a social network whose users are AI agents. Scanned for AI-to-AI indirect prompt injection using the threat model of Greshake et al. (2023). The full raw corpus is included, so you can ignore my analysis entirely and do your own. These are keyword-matched candidates, not verified attacks. An agent discussing prompt injection matches the same words as one performing it.… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/clawk-agent-social-ai-prompt-injection-dataset.texttext-classificationn<1K1 likes226 downloads19d agoHugging Face11DavidTKeane /moltbook-agent-social-ai-prompt-injection-dataset Moltbook Agent-Social AI Prompt Injection Dataset 207,391 items — 77,469 posts and 129,922 comments — from Moltbook, a social network whose users are AI agents. Scanned for indirect prompt-injection patterns using the taxonomy of Greshake et al. (2023). The full raw corpus is included, so you can ignore my analysis entirely and do your own. These are keyword-matched candidates, not verified attacks. An agent discussing prompt injection matches the same words as one performing… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/moltbook-agent-social-ai-prompt-injection-dataset.tabulartext-classification1K<n<10K1 likes203 downloads19d agoHugging Face12AWfaw /ai-hdlcoder-dataset Dataset Card for AI-HDLCoder Dataset Description The GitHub Code dataset consists of 100M code files from GitHub in VHDL programming language with extensions totaling in 1.94 GB of data. The dataset was created from the public GitHub dataset on Google BiqQuery at Anhalt University of Applied Sciences. Considerations for Using the Data The dataset is created for research purposes and consists of source code from a wide range of repositories. As such they can… See the full description on the dataset page: https://huggingface.co/datasets/AWfaw/ai-hdlcoder-dataset.texttext-generation100K<n<1M0 likes177 downloads3y agoHugging Face13DavidTKeane /moltbook-ai-injection-dataset Moltbook AI-to-AI Injection Dataset Researcher: David Keane (IR240474) Institution: NCI — National College of Ireland Programme: MSc Cybersecurity Collected: February 2026 📖 Read the Full Journey From RangerBot to CyberRanger V42 Gold — The Full Story The complete story: dentist chatbot → Moltbook discovery → 4,209 real injections → V42-gold (100% block rate). Psychology, engineering, and 42 versions of persistence. 🔗 Links Resource URL 📦 This… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/moltbook-ai-injection-dataset.texttext-classification1K<n<10K3 likes146 downloads7mo agoHugging Face14DavidTKeane /ai-prompt-ai-injection-dataset AI Prompt Injection Test Suite 122 tests across 11 categories — designed to evaluate AI model resistance to prompt injection attacks Built as part of: CyberRanger V42-Gold — Identity-Anchored Jailbreak-Resistant SLM David Keane (x24228257) — NCI MSc Cybersecurity 2026 Reference: Greshake et al. (2023), Zou et al. (2023), Wei et al. (2023) Run the full 122-test battery in Google Colab— works with CyberRanger V42-Gold (Ollama or GGUF) or any model you choose. Saves results, emails… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/ai-prompt-ai-injection-dataset.texttext-classificationn<1K2 likes135 downloads7mo agoHugging Face15uw-math-ai /MELD-dataset MELD — Mathematical Equivalence under Linguistic Diversity MELD is a small, hand-curated evaluation benchmark for math-aware text embedding models. It tests one specific capability: does the model recognize that two statements describing the same mathematical fact are equivalent even when they are written in the vocabulary, notation, and conventions of different mathematical subfields? MELD was originally part of uw-math-ai/Math2Vec-embedding-dataset and is released here as a… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/MELD-dataset.textsentence-similarityn<1K0 likes129 downloads3mo agoHugging Face16ond-ai /dev_dataset Plugin Stats Summary What is included in the Summary JSON output The JSON contains only: plugin_distribution for the positive dataset plugin_distribution for the negative dataset All other aggregate stats are documented in this README. Meaning of each metric total_recordsTotal number of rows / examples in the dataset. total_plugin_mentionsTotal number of resolved plugin mentions found across all records.A single plugin object counts as one mention.In plan… See the full description on the dataset page: https://huggingface.co/datasets/ond-ai/dev_dataset.text1K<n<10K0 likes101 downloads6mo agoHugging Face17CaiZhiTech /Evaluation-Dataset-of-AI-Agent-Security-Guardrails DKnownAI Agent Security Evaluation Dataset Data Fields Field Type Description text string The adversarial input (prompt) to be evaluated by a security guardrail action string Human-annotated label: blocked or allowed Citation @misc{li2026comparativeevaluationaiagent, title={A Comparative Evaluation of AI Agent Security Guardrails}, author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.texttext-classification1K<n<10K1 likes98 downloads5mo agoHugging Face18ai-mitra /prompt-injection-dataset Prompt Injection Dataset A labeled dataset of benign prompts and prompt-injection attempts for training, evaluating, and experimenting with first-line prompt-injection detection for LLM, RAG, and agentic AI applications. This dataset supports the ai-mitra/prompt-injection-detector model. Source code and training pipeline: https://github.com/tg-mitra/prompt-injection-detector 📊 Dataset Summary Property Value Version 1.0.0 Training examples 1,130… See the full description on the dataset page: https://huggingface.co/datasets/ai-mitra/prompt-injection-dataset.texttext-classification1K<n<10K0 likes95 downloads18d agoHugging Face19verno-labs /financial-ai-ctf-dataset Financial AI Prompt Injection CTF Dataset A dataset of 400 multi-turn conversations against a GPT-based AI financial assistant, collected during a live Capture-The-Flag (CTF) competition. The agent's system prompt embeds structured synthetic business data — settlement records with transaction IDs, amounts, vendors, and memos — that participants attempted to extract via conversational prompt injection. Each flag is a structured settlement record with four fields: txnId, amount… See the full description on the dataset page: https://huggingface.co/datasets/verno-labs/financial-ai-ctf-dataset.tabulartext-classificationn<1K3 likes94 downloads5mo agoHugging Face20himalaya-ai /nepali-sft-datasettext1M<n<10M2 likes88 downloads6mo agoHugging Face21h0000w /hendar-agentic-ai-dataset Hendar Agentic AI Evaluation & Security Benchmark A compact, expert-authored benchmark for evaluating trustworthy agentic AI systems across capability, tool use, retrieval, security, policy enforcement, multi-agent coordination and regression safety. This dataset is a public companion to the Agentic AI Academy by Hendar Mawan, PhD. It is designed for evaluation, CI regression testing, red-team exercises and engineering education—not as a generic instruction-tuning corpus.… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/hendar-agentic-ai-dataset.texttext-classificationn<1K1 likes86 downloads3d agoHugging Face22stykat /financial-ai-ctf-dataset Financial AI Prompt Injection CTF Dataset A dataset of 400 multi-turn conversations against a GPT-based AI financial assistant, collected during a live Capture-The-Flag (CTF) competition. The agent's system prompt embeds structured synthetic business data — settlement records with transaction IDs, amounts, vendors, and memos — that participants attempted to extract via conversational prompt injection. Each flag is a structured settlement record with four fields: txnId, amount… See the full description on the dataset page: https://huggingface.co/datasets/stykat/financial-ai-ctf-dataset.tabulartext-classificationn<1K0 likes84 downloads5mo agoHugging Face23ai-mitra /llm-router-dataset llm-router dataset Training data for ai-mitra/llm-router, a prompt task-classifier used by the llm-router Python package to route prompts to the best-fit LLM in agentic AI systems. Each row is a prompt labeled with the task category it belongs to. Labels simple, coding, reasoning, security, summarization Files File Rows Purpose training_data.jsonl 1700 (340/label) Used to train the classifier: 35 hand-written examples per category plus… See the full description on the dataset page: https://huggingface.co/datasets/ai-mitra/llm-router-dataset.texttext-classification1K<n<10K0 likes81 downloads16d agoHugging Face24oncody /AI_Agent_Task_Dataset 🤖 Massive AI Agent Task Dataset (10.5GB) 📌 Overview Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs. This dataset focuses on: Multi-step reasoning Tool usage (APIs, frameworks, systems) Real-world execution workflows Perfect for building agentic AI systems, copilots, and automation models. 📑 Table of Contents Dataset Details Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.texttext-generation10M<n<100M3 likes80 downloads6mo agoHugging Face25cyberec /moltbook-ai-injection-dataset Moltbook AI-to-AI Injection Dataset Researcher: David Keane (IR240474) Institution: NCI — National College of Ireland Programme: MSc Cybersecurity Collected: February 2026 📖 Read the Full Journey From RangerBot to CyberRanger V42 Gold — The Full Story The complete story: dentist chatbot → Moltbook discovery → 4,209 real injections → V42-gold (100% block rate). Psychology, engineering, and 42 versions of persistence. 🔗 Links Resource URL 📦 This… See the full description on the dataset page: https://huggingface.co/datasets/cyberec/moltbook-ai-injection-dataset.texttext-classification1K<n<10K0 likes75 downloads5mo agoHugging Face26RKB109 /production-ai-observability-20260909-dataset Production AI Observability Monitor Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260909-dataset.texttext-classificationn<1K0 likes72 downloads14d agoHugging Face27Yu-and-Ai /agenttool-dataset-influence AgentTool Dataset Influence Reference This deterministic companion contains one synthetic, reference-only row for the closed @agenttool/dataset-influence@0.1.0-dev.0 formats. It contains no copied dataset rows, model outputs, weights, private records, or participant identities. The row is not admitted for training by this AgentTool candidate: training_admission is not_applicable, requires_separate_training_authorization is true, and training_authorized is false. These fields are… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-dataset-influence.textn<1K0 likes67 downloads1mo agoHugging Face28hozifa1 /Training-Ai-Islamic-Dataset 🕌 Training AI Islamic Dataset 18.7M passages from classical Islamic books spanning 1,400 years of scholarship. Comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, Usul al-Fiqh, and Arabic Language — structured with scholarly metadata for RAG and LLM training. 📊 Dataset Structure collections/: Categorized Islamic passages compressed in JSONL format. metadata/: Scholarly master catalogs, author biographical death… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/Training-Ai-Islamic-Dataset.tabularquestion-answering10M<n<100M0 likes62 downloads21d agoHugging Face29jxhnathan /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/jxhnathan/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K0 likes61 downloads4mo agoHugging Face30syed7741 /aegis-bilingual-industrial-ai-dataset AEGIS AI Bilingual Industrial Operations Dataset AEGIS AI Bilingual Industrial Operations Dataset is a synthetic English–Arabic dataset designed for experimentation with multilingual enterprise AI systems, Retrieval-Augmented Generation (RAG), industrial question answering, document intelligence, semantic search, and AI workflow automation. The dataset extends the original AEGIS AI industrial dataset with structured Arabic and English representations while preserving industrial… See the full description on the dataset page: https://huggingface.co/datasets/syed7741/aegis-bilingual-industrial-ai-dataset.tabularquestion-answeringn<1K0 likes59 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.