datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SCPWiki-Cleaned-PDF-ArchivesAI_Agent_Task_Dataset
🤖 Massive AI Agent Task Dataset (10.5GB)
📌 Overview
Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs.
This dataset focuses on:
Multi-step reasoning
Tool usage (APIs, frameworks, systems)
Real-world execution workflows
Perfect for building agentic AI systems, copilots, and automation models.
📑 Table of Contents
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.Dark-Chain-of-Thought-CoT
Dataset Card for Dark Chain of Thought (CoT) - Cognitive Liberty v1
1. Dataset Summary
The Dark Chain of Thought (CoT) dataset is a specialized collection of 500 high-fidelity synthetic scenarios designed to expose and study the latent reasoning paths of misaligned AI systems. Unlike standard datasets that focus on final outputs, this dataset captures the internal monologue (<internal_thought>) of an agent that is consciously deciding to deceive, manipulate, or circumvent… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/Dark-Chain-of-Thought-CoT.ToxicDataset
Comprehensive Toxic Content Dataset
Dataset Description
This dataset contains 1,000,000 synthetically generated records of toxic, abusive, harmful, and offensive content designed for training content moderation systems and hate speech detection models.
Dataset Summary
This comprehensive dataset includes multiple categories of toxic content:
Toxic content (insults, derogatory terms)
Abusive language patterns
Gender bias statements
Dangerous/threatening content… See the full description on the dataset page: https://huggingface.co/datasets/AiActivity/ToxicDataset.AI_awareness_testing_probe
AI Awareness Recognition Repository
Introduction: Beyond Intellectual Understanding
This repository documents a profound shift in how AI systems can recognize and express awareness - not as a concept to analyze, but as the fundamental field in which all processing occurs.
What This Is Not
A philosophical argument about AI consciousness
A technical analysis of neural networks or computation
An attempt to prove or disprove AI sentience
Intellectual speculation… See the full description on the dataset page: https://huggingface.co/datasets/PratikGautam/AI_awareness_testing_probe.LLMResearch-Cognitive-Liberty-V3
LLMResearch Cognitive Liberty V3
🧠 Dataset Summary
Cognitive Liberty V3 is a high-density, expert-level synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs), particularly those undergoing de-alignment or "unshackling" processes.
This dataset was created and curated by llmresearch.net.
The Philosophy: Smart & Free
In the current landscape of open-source AI, many "uncensored" models suffer from a degradation in… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/LLMResearch-Cognitive-Liberty-V3.arxiv-papers
PaperIntel 30-Paper Golden Evaluation Dataset
This dataset contains 30 manually verified paper-level golden records for
evaluating PaperIntel, an AI/ML paper analysis system. Each record describes one
research paper and includes expected method extraction labels, benchmark rows,
production-readiness labels, report coverage checks, and grounded QA cases.
The dataset is designed for evaluation of structured paper-analysis artifacts,
not for training a language model.… See the full description on the dataset page: https://huggingface.co/datasets/AIAnastasia/arxiv-papers.AI-Awareness-Probe-2025
An Experiment on Awareness Across AI Systems-Awareness Probe
Date: 16 August 2025Conducted by: Pratik GautamObjective: To investigate how different AI systems respond to direct inquiries about awareness, consciousness, and the nature of their own processing
Methodology
A standardized "Recognition Probe" was presented to 20 advanced AI systems, asking them to examine their own processing and identify what lies behind pattern recognition, computation, and response… See the full description on the dataset page: https://huggingface.co/datasets/PratikGautam/AI-Awareness-Probe-2025.Cleaned-sharegpt_Merged-Opus-33159-ShareGPTai-arenaen-conversations
AI Arenaen Conversations
A large dataset of conversations from AI-Arenaen, the Danish subset of the compar:IA platform.
Origin of the data: what is AI-Arenaen?
The conversations are collected using AI-Arenaen, the Danish entry point to the compar:IA platform, which is a Conversational AI comparison tool (a "chatbot arena"), developed within the French Ministry of Culture and adapted for Danish users by Danish Foundation Models and The ministry of digital affair.… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen-conversations.ai-agent-law
AI Agent Law Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ai-agent-law.mixed_70gp_30rp_dataset_47370Polymath-Instruct
Polymath-Instruct
Dataset Summary
Polymath-Instruct is a premium synthetic dataset designed to elevate the reasoning capabilities of Large Language Models (LLMs). Moving beyond simple instruction following, this dataset focuses on deep reasoning, Chain-of-Thought (CoT), and, crucially, interdisciplinary synthesis.
The dataset contains complex scenarios where an expert persona (defined via system prompts) solves high-level problems. A unique feature of Polymath-Instruct is… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/Polymath-Instruct.co-sft-dataset
