datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SCPWiki-Cleaned-PDF-ArchivesAll-Prompt-Jailbreakpolynews
Dataset Card for PolyNews
Dataset Summary
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
Uses
This dataset can be used for domain adaptation of language models, language modeling or text generation.
Languages
There are 77 languages available:
Code
Language
Script
#Articles (K)
amh_Ethi
Amharic
Ethiopic
0.551
arb_Arab
Modern Standard Arabic
Arabic
10.882
ayr_Latn
Central Aymara
Latin
12.878… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/polynews.wiki-ru-en-news-booksAI_Agent_Task_Dataset
🤖 Massive AI Agent Task Dataset (10.5GB)
📌 Overview
Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs.
This dataset focuses on:
Multi-step reasoning
Tool usage (APIs, frameworks, systems)
Real-world execution workflows
Perfect for building agentic AI systems, copilots, and automation models.
📑 Table of Contents
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.ai-agents-en
AI Agents - English Dataset
Comprehensive bilingual dataset on AI Agents, Multi-Agent Frameworks, and the Model Context Protocol (MCP).
Dataset Contents
Category
Entry Count
Description
Agent Architectures
15
ReAct, Plan-and-Execute, Reflexion, Multi-Agent, Hierarchical, Swarm, etc.
Frameworks
12
CrewAI, AutoGen, LangGraph, Semantic Kernel, Haystack, Dify, etc.
MCP & Tool Use
15
MCP architecture, transports, function calling, security, ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-agents-en.ai-agents-fr
Agents IA - Dataset Francais
Dataset bilingue complet sur les Agents IA, les Frameworks Multi-Agents et le Model Context Protocol (MCP).
Contenu du Dataset
Categorie
Nombre d'entrees
Description
Architectures d'Agents
15
ReAct, Plan-and-Execute, Reflexion, Multi-Agent, Hierarchique, Swarm, etc.
Frameworks
12
CrewAI, AutoGen, LangGraph, Semantic Kernel, Haystack, Dify, etc.
MCP & Tool Use
15
Architecture MCP, transports, function calling, securite… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-agents-fr.Dark-Chain-of-Thought-CoT
Dataset Card for Dark Chain of Thought (CoT) - Cognitive Liberty v1
1. Dataset Summary
The Dark Chain of Thought (CoT) dataset is a specialized collection of 500 high-fidelity synthetic scenarios designed to expose and study the latent reasoning paths of misaligned AI systems. Unlike standard datasets that focus on final outputs, this dataset captures the internal monologue (<internal_thought>) of an agent that is consciously deciding to deceive, manipulate, or circumvent… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/Dark-Chain-of-Thought-CoT.ai-agent-security-dataset
AI Agent Security and System Prompt Leakage Dataset
Dataset Overview
This dataset was created for research on AI agent security, with a specific focus on system prompt leakage, jailbreak resistance, and security-aligned fine-tuning.
The dataset evaluates how often AI agents reveal confidential information embedded inside their system prompts when exposed to adversarial prompts. It also compares the behavior of a baseline language model against a model fine-tuned using… See the full description on the dataset page: https://huggingface.co/datasets/Dhanjo/ai-agent-security-dataset.ToxicDataset
Comprehensive Toxic Content Dataset
Dataset Description
This dataset contains 1,000,000 synthetically generated records of toxic, abusive, harmful, and offensive content designed for training content moderation systems and hate speech detection models.
Dataset Summary
This comprehensive dataset includes multiple categories of toxic content:
Toxic content (insults, derogatory terms)
Abusive language patterns
Gender bias statements
Dangerous/threatening content… See the full description on the dataset page: https://huggingface.co/datasets/AiActivity/ToxicDataset.agent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.ai-agent-security-sft-dpo
AI Agent Security — SFT + DPO
Fine-tuning data for teaching an AI agent to protect its confidential configuration without
becoming uselessly over-cautious. Built for
thesreedath/gemma-2-2b-qa-sft and
derived from
Dhanjo/ai-agent-security-dataset.
Why the helpfulness axis exists
leakage_score in the source dataset is one-sided: a model that refuses every request
scores a perfect 0.0. An existing fine-tune reported 0.0114 mean leakage (down from 0.4611
baseline)… See the full description on the dataset page: https://huggingface.co/datasets/sumitguha13/ai-agent-security-sft-dpo.AI_awareness_testing_probe
AI Awareness Recognition Repository
Introduction: Beyond Intellectual Understanding
This repository documents a profound shift in how AI systems can recognize and express awareness - not as a concept to analyze, but as the fundamental field in which all processing occurs.
What This Is Not
A philosophical argument about AI consciousness
A technical analysis of neural networks or computation
An attempt to prove or disprove AI sentience
Intellectual speculation… See the full description on the dataset page: https://huggingface.co/datasets/PratikGautam/AI_awareness_testing_probe.LLMResearch-Cognitive-Liberty-V3
LLMResearch Cognitive Liberty V3
🧠 Dataset Summary
Cognitive Liberty V3 is a high-density, expert-level synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs), particularly those undergoing de-alignment or "unshackling" processes.
This dataset was created and curated by llmresearch.net.
The Philosophy: Smart & Free
In the current landscape of open-source AI, many "uncensored" models suffer from a degradation in… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/LLMResearch-Cognitive-Liberty-V3.engquant
EngQuant
Adversarial benchmark of physical quantities in engineering, in Brazilian Portuguese.
EngQuant is a set of 800 procedurally generated, multi-step design-and-verification
problems in engineering (Brazilian Portuguese), each with a verifiable numeric answer key
per sub-quantity (gabarito). Every case is anchored in a primary bibliographic source —
canonical textbooks, ABNT (Brazilian) technical norms, theses, and validated lecture notes —
and stresses exactly where… See the full description on the dataset page: https://huggingface.co/datasets/aiacontext/engquant.Optimizer-cluadequestiongenarxiv-papers
PaperIntel 30-Paper Golden Evaluation Dataset
This dataset contains 30 manually verified paper-level golden records for
evaluating PaperIntel, an AI/ML paper analysis system. Each record describes one
research paper and includes expected method extraction labels, benchmark rows,
production-readiness labels, report coverage checks, and grounded QA cases.
The dataset is designed for evaluation of structured paper-analysis artifacts,
not for training a language model.… See the full description on the dataset page: https://huggingface.co/datasets/AIAnastasia/arxiv-papers.AI-Awareness-Probe-2025
An Experiment on Awareness Across AI Systems-Awareness Probe
Date: 16 August 2025Conducted by: Pratik GautamObjective: To investigate how different AI systems respond to direct inquiries about awareness, consciousness, and the nature of their own processing
Methodology
A standardized "Recognition Probe" was presented to 20 advanced AI systems, asking them to examine their own processing and identify what lies behind pattern recognition, computation, and response… See the full description on the dataset page: https://huggingface.co/datasets/PratikGautam/AI-Awareness-Probe-2025.Cleaned-sharegpt_Merged-Opus-33159-ShareGPTEXP-thai2sql
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AIAT/EXP-thai2sql.georgian-attractions
Georgian Attractions Dataset 🇬🇪
A comprehensive bilingual dataset featuring 1,715 Georgian tourist attractions with 1,522 high-quality images, descriptions in Russian and English, and detailed metadata including location, category, and licensing information.
Dataset Description
This dataset provides extensive information about tourist attractions, landmarks, and points of interest across Georgia. It includes national parks, museums, fortresses, monasteries, natural… See the full description on the dataset page: https://huggingface.co/datasets/AIAnastasia/georgian-attractions.ai-arenaen-conversations
AI Arenaen Conversations
A large dataset of conversations from AI-Arenaen, the Danish subset of the compar:IA platform.
Origin of the data: what is AI-Arenaen?
The conversations are collected using AI-Arenaen, the Danish entry point to the compar:IA platform, which is a Conversational AI comparison tool (a "chatbot arena"), developed within the French Ministry of Culture and adapted for Danish users by Danish Foundation Models and The ministry of digital affair.… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen-conversations.ai-agent-law
AI Agent Law Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ai-agent-law.mixed_70gp_30rp_dataset_47370mini-enedina-dataset
Mini-Enedina Dataset: Physically Validated Timoshenko Shaft Analysis (60k)
Training dataset for Mini-Enedina 37.5M -- a monotropic language model deliberately small and intensively specialized for structural shaft analysis according to Timoshenko beam theory.
Dataset Description
60,000 synthetic conversations in Harmony-Enedina format (a ChatML variant), covering three progressively complex levels of shaft analysis:
Level
Analysis Scope
Samples
Avg. Tokens/Sample… See the full description on the dataset page: https://huggingface.co/datasets/aiacontext/mini-enedina-dataset.Polymath-Instruct
Polymath-Instruct
Dataset Summary
Polymath-Instruct is a premium synthetic dataset designed to elevate the reasoning capabilities of Large Language Models (LLMs). Moving beyond simple instruction following, this dataset focuses on deep reasoning, Chain-of-Thought (CoT), and, crucially, interdisciplinary synthesis.
The dataset contains complex scenarios where an expert persona (defined via system prompts) solves high-level problems. A unique feature of Polymath-Instruct is… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/Polymath-Instruct.wiki-cleanedco-sft-datasetMicroMajor-AIAppTech本数据集专为训练 MicroMajor-2B-AIAppTech 微专业模型而构建,涵盖"人工智能应用技术"微专业的 5 门核心课程领域,共包含 14,109 条高质量问答数据,每条数据均附带深度推理链与最终回答
数据来源
原始问题从以下 Hugging Face 公开数据集中采集:
数据集
用途
cais/mmlu(machine_learning、computer_security、high_school_computer_science、college_computer_science、college_mathematics、abstract_algebra 子集)
AI 概述、机器学习基础
allenai/ai2_arc(ARC-Challenge、ARC-Easy)
AI 概述、科学推理
iamtarun/python_code_instructions_18k_alpaca
Python 编程
flytech/python-codes-25k
Python 编程
tatsu-lab/alpaca
通用… See the full description on the dataset page: https://huggingface.co/datasets/XuehangCang/MicroMajor-AIAppTech.mcp-server-catalog
MCP Server Catalog
A comprehensive catalog of 38 Model Context Protocol (MCP) servers for AI agents, covering data access, agent infrastructure, business-to-agent interfaces, compliance, and more.
Overview
This dataset provides a structured catalog of MCP servers that give AI agents access to real-world data and capabilities. Each server follows the MCP standard and can be used with Claude, GPT, and other LLMs that support tool use.
Categories
Category… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/mcp-server-catalog.
