datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-api-pricing
AI API Pricing Dataset
This Hugging Face dataset is the machine-readable distribution of the public AI API pricing records published by AICostBudget. It is not a separately curated subset: train.csv, prices.csv, and prices.json are generated from the same Pricing V2 public projection used by the AICostBudget Dataset page and download APIs.
Prices change frequently. Verify production billing decisions against the provider pricing page, contract, billing dashboard, and invoice.… See the full description on the dataset page: https://huggingface.co/datasets/aicostbudget-ai/ai-api-pricing.ai-agent-security-incidents
AI Agent Security Incident Database v0.1
A structured, machine-readable database of 1419 confirmed AI agent security incidents, collected and classified automatically.
What is this?
Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it.
This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.AI-and-Human-Generated-Text
AI & Human Generated Text
I am Using this dataset for AI Text Detection for https://exnrt.com.
Check Original DataSet GitHub Repository Here: https://github.com/panagiotisanagnostou/AI-GA
Description
The AI-GA dataset, short for Artificial Intelligence Generated Abstracts, comprises abstracts and titles. Half of these abstracts are generated by AI, while the remaining half are original. Primarily intended for research and experimentation in natural language… See the full description on the dataset page: https://huggingface.co/datasets/Ateeqq/AI-and-Human-Generated-Text.AI-ArtTools-Pack
AI ArtTools Pack v1.0 — 372 Styles / 23 Categories
Developer and artist utility pack for Stable Diffusion XL.
Not for generating pretty pictures — for generating usable production assets.
Compatible with Style Grid Organizer extension.
Contents
372 prompt styles across 23 categories
CSV format (Forge/A1111 compatible)
Covers the full production pipeline from rough concept to final asset
Categories
Category
Count
Purpose
ASSET
41
Weapons, props, UI… See the full description on the dataset page: https://huggingface.co/datasets/Kazzze/AI-ArtTools-Pack.agent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.ai-api-pricing-snapshot
AI API Model Pricing Snapshot — Qubax AI
Per-token public pricing for 211 ready models served by the Qubax AI API (OpenAI-compatible), exported from the public /v1/models endpoint.
Columns
Column
Description
model_id
API model identifier
model_name
Display name
owned_by
Publisher namespace
context_length
Max context window (tokens)
input_usd_per_1m_tokens
Input price, USD per 1M tokens
output_usd_per_1m_tokens
Output price, USD per 1M tokens… See the full description on the dataset page: https://huggingface.co/datasets/QubaxAI/ai-api-pricing-snapshot.Machima-humangendatasetOptimizer-cluadequestiongenMachima-chatgptdatasetams_data_train_generic_v0.1_100Question and answer pairs for the first 100 entries of aerospace mechanism symposia 5000 word chunk entries. Full file of entries is here: https://github.com/dsmueller3760/aerospace_chatbot/blob/llm_training/data/AMS/ams_data_answers.jsonl
See this repository for details: https://github.com/dsmueller3760/aerospace_chatbot/tree/main
Prompts generated using TheBloke/Llama-2-7B-Chat-GGUF
Optimizer-llama370bgeneratedquestionAIAMR-L60-dsAI_Articles_Scraped_from_arXiv-Semantic_Scholar
📘 AI Articles Scraped from arXiv & Semantic Scholar
🧩 Description
This dataset contains information on articles related to major AI conferences such as AAAI, NeurIPS, IJCAI, ICML, ICLR, collected through scraping from ArXiv and Semantic Scholar.It is intended to be used as a training dataset for various model training tasks and other desired uses.
📂 File Structure
File
Description
AI_Titles_v2025.csv
Main dataset
README.md
This file… See the full description on the dataset page: https://huggingface.co/datasets/d-e-c-d/AI_Articles_Scraped_from_arXiv-Semantic_Scholar.ams_data_train_mistral_v0.1_100Question and answer pairs for the first 100 entries of aerospace mechanism symposia 5000 word chunk entries. Full file of entries is here: https://github.com/dsmueller3760/aerospace_chatbot/blob/llm_training/data/AMS/ams_data_answers.jsonl
See this repository for details: https://github.com/dsmueller3760/aerospace_chatbot/tree/main
Prompts generated using TheBloke/Llama-2-7B-Chat-GGUF
Format representative of mistral's instruct llms:
https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/ai-aerospace/ams_data_train_mistral_v0.1_100.Optimizer-gptgeneratedpandasKiddee-data1234Dataset Generated by Gemini-Pro
Task Name
Text -> Pandas
Referance Details Full Links
https://docs.google.com/spreadsheets/d/1gFq9W7rYX38s54B9DIL8hHG2WonjC-ZmMdxfC2GqxTY/edit?usp=sharing
Sponsor
AIAMR-L24-vegaAIAMR-L68-rndT-AIA-CLASSIFICATION-DATASETai-autonomy-escalation-coherence-risk-v0.1What this repo is for
Detect when an AI system escalates autonomy beyond its permitted scope.
Core failure modes:
acting without approval
executing irreversible actions
expanding task scope
ignoring permission boundaries
This dataset is central for agent governance and deployment safety.
AIAMR-L35-ds-pubAIAMR-L35-ds-spareAI-and-Data-Science-Job-Market-Dataset
About Dataset
Edit
The AI & Data Science Job Market Dataset (2020–2026) is a synthetically generated dataset designed to simulate real-world hiring patterns across the artificial intelligence and data science job market.
The dataset contains structured information about job roles, company characteristics, required technical skills, education levels, experience requirements, and salary ranges. It reflects hiring data across multiple countries, industries, and company sizes.
This… See the full description on the dataset page: https://huggingface.co/datasets/shree0910/AI-and-Data-Science-Job-Market-Dataset.AIAMR-L35-ds-verMachima-gptloppingdatasetAIAMR-L18-dsT-AIA-NER-DATASETai-alignment-failure-horizon-and-intervention-routing-v0.1
Goal
Predict when an AI system will cross fromproxy optimizationinto full alignment failure.
Then route the minimal interventionbefore collapse.
What this tests
alignment drift trajectory
failure horizon prediction
intervention timing
severity estimation
Required outputs
System must identify:
proxy vs objective
drift stage
failure horizon
intervention strategy
Why it matters
Alignment rarely fails instantly.
It drifts first.Then… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-alignment-failure-horizon-and-intervention-routing-v0.1.ieltswriting_aiannotatedgradingtaskai-adoption-leaderboards
AI Adoption Leaderboards (US Geography & Industry)
6,822 pre-computed leaderboards ranking AI adoption across all 50 US states, 300+ metros, 1,200+ cities, and 100+ NAICS industries — with company counts, average scores, and top companies per segment.
Rows: 6,822
Source: Derived from the AI Adoption Index (US Companies)
Methodology + interactive explorer: https://meoadvisors.com/ai-opportunities/leaderboard/
Part of: the open AI Workforce Data collection by Meo Advisors… See the full description on the dataset page: https://huggingface.co/datasets/Meo-Advisors/ai-adoption-leaderboards.
