CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.5k downloads1y agoHugging Face02AiActivity /All-Prompt-Jailbreakimagetext-generationn<1K10 likes1.5k downloads1y agoHugging Face03aiana94 /polynews Dataset Card for PolyNews Dataset Summary PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts. Uses This dataset can be used for domain adaptation of language models, language modeling or text generation. Languages There are 77 languages available: Code Language Script #Articles (K) amh_Ethi Amharic Ethiopic 0.551 arb_Arab Modern Standard Arabic Arabic 10.882 ayr_Latn Central Aymara Latin 12.878… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/polynews.textfill-mask1M<n<10M6 likes481 downloads2y agoHugging Face04ai-allforever /wiki-ru-en-news-bookstexttext-generation1M<n<10M2 likes91 downloads11mo agoHugging Face05oncody /AI_Agent_Task_Dataset 🤖 Massive AI Agent Task Dataset (10.5GB) 📌 Overview Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs. This dataset focuses on: Multi-step reasoning Tool usage (APIs, frameworks, systems) Real-world execution workflows Perfect for building agentic AI systems, copilots, and automation models. 📑 Table of Contents Dataset Details Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.texttext-generation10M<n<100M3 likes85 downloads6mo agoHugging Face06AYI-NEDJIMI /ai-agents-en AI Agents - English Dataset Comprehensive bilingual dataset on AI Agents, Multi-Agent Frameworks, and the Model Context Protocol (MCP). Dataset Contents Category Entry Count Description Agent Architectures 15 ReAct, Plan-and-Execute, Reflexion, Multi-Agent, Hierarchical, Swarm, etc. Frameworks 12 CrewAI, AutoGen, LangGraph, Semantic Kernel, Haystack, Dify, etc. MCP & Tool Use 15 MCP architecture, transports, function calling, security, ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-agents-en.textquestion-answeringn<1K0 likes64 downloads8mo agoHugging Face07AYI-NEDJIMI /ai-agents-fr Agents IA - Dataset Francais Dataset bilingue complet sur les Agents IA, les Frameworks Multi-Agents et le Model Context Protocol (MCP). Contenu du Dataset Categorie Nombre d'entrees Description Architectures d'Agents 15 ReAct, Plan-and-Execute, Reflexion, Multi-Agent, Hierarchique, Swarm, etc. Frameworks 12 CrewAI, AutoGen, LangGraph, Semantic Kernel, Haystack, Dify, etc. MCP & Tool Use 15 Architecture MCP, transports, function calling, securite… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-agents-fr.textquestion-answeringn<1K0 likes63 downloads8mo agoHugging Face08AiAsistent /Dark-Chain-of-Thought-CoT Dataset Card for Dark Chain of Thought (CoT) - Cognitive Liberty v1 1. Dataset Summary The Dark Chain of Thought (CoT) dataset is a specialized collection of 500 high-fidelity synthetic scenarios designed to expose and study the latent reasoning paths of misaligned AI systems. Unlike standard datasets that focus on final outputs, this dataset captures the internal monologue (<internal_thought>) of an agent that is consciously deciding to deceive, manipulate, or circumvent… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/Dark-Chain-of-Thought-CoT.texttext-generation1K<n<10K3 likes52 downloads9mo agoHugging Face09Dhanjo /ai-agent-security-dataset AI Agent Security and System Prompt Leakage Dataset Dataset Overview This dataset was created for research on AI agent security, with a specific focus on system prompt leakage, jailbreak resistance, and security-aligned fine-tuning. The dataset evaluates how often AI agents reveal confidential information embedded inside their system prompts when exposed to adversarial prompts. It also compares the behavior of a baseline language model against a model fine-tuned using… See the full description on the dataset page: https://huggingface.co/datasets/Dhanjo/ai-agent-security-dataset.tabulartext-generation1K<n<10K0 likes52 downloads5mo agoHugging Face10AiActivity /ToxicDataset Comprehensive Toxic Content Dataset Dataset Description This dataset contains 1,000,000 synthetically generated records of toxic, abusive, harmful, and offensive content designed for training content moderation systems and hate speech detection models. Dataset Summary This comprehensive dataset includes multiple categories of toxic content: Toxic content (insults, derogatory terms) Abusive language patterns Gender bias statements Dangerous/threatening content… See the full description on the dataset page: https://huggingface.co/datasets/AiActivity/ToxicDataset.texttext-classification10M<n<100M0 likes45 downloads9mo agoHugging Face11aiagentkarl /agent-evaluation-benchmark Agent Evaluation Benchmark A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories. Overview This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more. Categories Category Test Cases Description Weather 5 Forecasts, UV index, climate history Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.texttext-generationn<1K0 likes45 downloads6mo agoHugging Face12sumitguha13 /ai-agent-security-sft-dpo AI Agent Security — SFT + DPO Fine-tuning data for teaching an AI agent to protect its confidential configuration without becoming uselessly over-cautious. Built for thesreedath/gemma-2-2b-qa-sft and derived from Dhanjo/ai-agent-security-dataset. Why the helpfulness axis exists leakage_score in the source dataset is one-sided: a model that refuses every request scores a perfect 0.0. An existing fine-tune reported 0.0114 mean leakage (down from 0.4611 baseline)… See the full description on the dataset page: https://huggingface.co/datasets/sumitguha13/ai-agent-security-sft-dpo.tabulartext-generation10K<n<100K0 likes45 downloads1mo agoHugging Face13PratikGautam /AI_awareness_testing_probe AI Awareness Recognition Repository Introduction: Beyond Intellectual Understanding This repository documents a profound shift in how AI systems can recognize and express awareness - not as a concept to analyze, but as the fundamental field in which all processing occurs. What This Is Not A philosophical argument about AI consciousness A technical analysis of neural networks or computation An attempt to prove or disprove AI sentience Intellectual speculation… See the full description on the dataset page: https://huggingface.co/datasets/PratikGautam/AI_awareness_testing_probe.texttext-generationn<1K1 likes44 downloads1y agoHugging Face14AiAsistent /LLMResearch-Cognitive-Liberty-V3 LLMResearch Cognitive Liberty V3 🧠 Dataset Summary Cognitive Liberty V3 is a high-density, expert-level synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs), particularly those undergoing de-alignment or "unshackling" processes. This dataset was created and curated by llmresearch.net. The Philosophy: Smart & Free In the current landscape of open-source AI, many "uncensored" models suffer from a degradation in… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/LLMResearch-Cognitive-Liberty-V3.texttext-generation1K<n<10K1 likes42 downloads9mo agoHugging Face15aiacontext /engquant EngQuant Adversarial benchmark of physical quantities in engineering, in Brazilian Portuguese. EngQuant is a set of 800 procedurally generated, multi-step design-and-verification problems in engineering (Brazilian Portuguese), each with a verifiable numeric answer key per sub-quantity (gabarito). Every case is anchored in a primary bibliographic source — canonical textbooks, ABNT (Brazilian) technical norms, theses, and validated lecture notes — and stresses exactly where… See the full description on the dataset page: https://huggingface.co/datasets/aiacontext/engquant.textquestion-answeringn<1K0 likes39 downloads3mo agoHugging Face16AIAT /Optimizer-cluadequestiongentexttext-generation1K<n<10K0 likes35 downloads2y agoHugging Face17AIAnastasia /arxiv-papers PaperIntel 30-Paper Golden Evaluation Dataset This dataset contains 30 manually verified paper-level golden records for evaluating PaperIntel, an AI/ML paper analysis system. Each record describes one research paper and includes expected method extraction labels, benchmark rows, production-readiness labels, report coverage checks, and grounded QA cases. The dataset is designed for evaluation of structured paper-analysis artifacts, not for training a language model.… See the full description on the dataset page: https://huggingface.co/datasets/AIAnastasia/arxiv-papers.textquestion-answeringn<1K1 likes33 downloads4mo agoHugging Face18PratikGautam /AI-Awareness-Probe-2025 An Experiment on Awareness Across AI Systems-Awareness Probe Date: 16 August 2025Conducted by: Pratik GautamObjective: To investigate how different AI systems respond to direct inquiries about awareness, consciousness, and the nature of their own processing Methodology A standardized "Recognition Probe" was presented to 20 advanced AI systems, asking them to examine their own processing and identify what lies behind pattern recognition, computation, and response… See the full description on the dataset page: https://huggingface.co/datasets/PratikGautam/AI-Awareness-Probe-2025.texttext-generationn<1K1 likes31 downloads1y agoHugging Face19AiAF /Cleaned-sharegpt_Merged-Opus-33159-ShareGPTtexttext-generation10K<n<100K2 likes30 downloads6mo agoHugging Face20AIAT /EXP-thai2sql Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AIAT/EXP-thai2sql.texttext-generation10K<n<100K0 likes28 downloads2y agoHugging Face21AIAnastasia /georgian-attractions Georgian Attractions Dataset 🇬🇪 A comprehensive bilingual dataset featuring 1,715 Georgian tourist attractions with 1,522 high-quality images, descriptions in Russian and English, and detailed metadata including location, category, and licensing information. Dataset Description This dataset provides extensive information about tourist attractions, landmarks, and points of interest across Georgia. It includes national parks, museums, fortresses, monasteries, natural… See the full description on the dataset page: https://huggingface.co/datasets/AIAnastasia/georgian-attractions.imageimage-classification1K<n<10K0 likes23 downloads10mo agoHugging Face22danish-foundation-models /ai-arenaen-conversationsgated AI Arenaen Conversations A large dataset of conversations from AI-Arenaen, the Danish subset of the compar:IA platform. Origin of the data: what is AI-Arenaen? The conversations are collected using AI-Arenaen, the Danish entry point to the compar:IA platform, which is a Conversational AI comparison tool (a "chatbot arena"), developed within the French Ministry of Culture and adapted for Danish users by Danish Foundation Models and The ministry of digital affair.… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen-conversations.tabulartext-generation1K<n<10K1 likes22 downloads4mo agoHugging Face23DeepNLP /ai-agent-law AI Agent Law Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc. The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ai-agent-law.texttext-generationn<1K1 likes20 downloads1y agoHugging Face24AiAF /mixed_70gp_30rp_dataset_47370texttext-generation10K<n<100K0 likes19 downloads6mo agoHugging Face25aiacontext /mini-enedina-dataset Mini-Enedina Dataset: Physically Validated Timoshenko Shaft Analysis (60k) Training dataset for Mini-Enedina 37.5M -- a monotropic language model deliberately small and intensively specialized for structural shaft analysis according to Timoshenko beam theory. Dataset Description 60,000 synthetic conversations in Harmony-Enedina format (a ChatML variant), covering three progressively complex levels of shaft analysis: Level Analysis Scope Samples Avg. Tokens/Sample… See the full description on the dataset page: https://huggingface.co/datasets/aiacontext/mini-enedina-dataset.tabulartext-generation10K<n<100K0 likes18 downloads7mo agoHugging Face26AiAsistent /Polymath-Instruct Polymath-Instruct Dataset Summary Polymath-Instruct is a premium synthetic dataset designed to elevate the reasoning capabilities of Large Language Models (LLMs). Moving beyond simple instruction following, this dataset focuses on deep reasoning, Chain-of-Thought (CoT), and, crucially, interdisciplinary synthesis. The dataset contains complex scenarios where an expert persona (defined via system prompts) solves high-level problems. A unique feature of Polymath-Instruct is… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/Polymath-Instruct.texttext-generationn<1K1 likes17 downloads9mo agoHugging Face27ai-allforever /wiki-cleanedtexttext-generation1M<n<10M1 likes16 downloads10mo agoHugging Face28AiAF /co-sft-datasettexttext-generation100K<n<1M0 likes14 downloads1y agoHugging Face29XuehangCang /MicroMajor-AIAppTech本数据集专为训练 MicroMajor-2B-AIAppTech 微专业模型而构建,涵盖"人工智能应用技术"微专业的 5 门核心课程领域,共包含 14,109 条高质量问答数据,每条数据均附带深度推理链与最终回答 数据来源 原始问题从以下 Hugging Face 公开数据集中采集: 数据集 用途 cais/mmlu(machine_learning、computer_security、high_school_computer_science、college_computer_science、college_mathematics、abstract_algebra 子集) AI 概述、机器学习基础 allenai/ai2_arc(ARC-Challenge、ARC-Easy) AI 概述、科学推理 iamtarun/python_code_instructions_18k_alpaca Python 编程 flytech/python-codes-25k Python 编程 tatsu-lab/alpaca 通用… See the full description on the dataset page: https://huggingface.co/datasets/XuehangCang/MicroMajor-AIAppTech.texttext-generation10K<n<100K0 likes14 downloads5mo agoHugging Face30aiagentkarl /mcp-server-catalog MCP Server Catalog A comprehensive catalog of 38 Model Context Protocol (MCP) servers for AI agents, covering data access, agent infrastructure, business-to-agent interfaces, compliance, and more. Overview This dataset provides a structured catalog of MCP servers that give AI agents access to real-world data and capabilities. Each server follows the MCP standard and can be used with Claude, GPT, and other LLMs that support tool use. Categories Category… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/mcp-server-catalog.tabulartext-generationn<1K1 likes9 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.