datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AutoMathText-V2
🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset
🎉 AutoMathText-v2 has surpassed 1.5 million downloads! We'd love to know how you're using it. Please take 1 minute to fill out our use case survey. Your feedback will directly shape the future roadmap of this dataset.👉 Share your use case here
📊 AutoMathText-V2 consists of 2.46 trillion tokens of high-quality, deduplicated text spanning web content, mathematics, code, reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSQZ/AutoMathText-V2.AutoMathText🎉 This work, introducing the AutoMathText dataset and the AutoDS method, has been accepted to The 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Findings)! 🎉
AutoMathText
AutoMathText is an extensive and carefully curated dataset encompassing around 200 GB of mathematical texts. It's a compilation sourced from a diverse range of platforms including various websites, arXiv, and GitHub (OpenWebMath, RedPajama, Algebraic Stack). This rich repository… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/AutoMathText.BioKGBench-Dataset
Agent4S-BioKG
A Knowledge Graph Checking Benchmark of AI Agent for Biomedical Science.
Github
Introduction
Pursuing artificial intelligence for biomedical science, a.k.a. AI Scientist, draws increasing attention, where one common approach is to build a copilot agent driven by Large Language Models(LLMs).However, to evaluate such systems, people either rely on direct Question-Answering(QA) to the LLM itself, or in a biomedical experimental manner. How… See the full description on the dataset page: https://huggingface.co/datasets/AutoLab-Westlake/BioKGBench-Dataset.AutoMathText🎉 This work, introducing the AutoMathText dataset and the AutoDS method, has been accepted to The 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Findings)! 🎉
AutoMathText
AutoMathText is an extensive and carefully curated dataset encompassing around 200 GB of mathematical texts. It's a compilation sourced from a diverse range of platforms including various websites, arXiv, and GitHub (OpenWebMath, RedPajama, Algebraic Stack). This rich… See the full description on the dataset page: https://huggingface.co/datasets/NP235/AutoMathText.TimeSeriesExam1
Dataset Card for TimeSeriesExam-1
This dataset provides Question-Answer (QA) pairs for the paper TimeSeriesExam: A Time Series Understanding Exam. Example inference code can be found here.
📖Introduction
Large Language Models (LLMs) have recently demonstrated a remarkable ability to model time series data. These capabilities can be partly explained if LLMs understand basic time series concepts. However, our knowledge of what these models understand about time series data… See the full description on the dataset page: https://huggingface.co/datasets/AutonLab/TimeSeriesExam1.hallucination-autopsy-benchmark
Hallucination Autopsy Benchmark
A unified, standardized benchmark for cross-model, cross-parameter analysis of LLM hallucination phenomena.
Overview
This dataset merges multiple hallucination detection benchmarks into a single standardized schema, enabling systematic etiological analysis of why and how different LLM architectures hallucinate under specific configurations.
Version: 3.0.0Total Records: 69,002Base Records: 69,002Augmented Records: 0Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/OiQ/hallucination-autopsy-benchmark.auto-wiki-qa
AutoWikiQA
東工大が公開しているSwallow-MXを用いて、Wikipedia中のテキストを入力として「質問(query)」と「回答(answer)」を生成し、生成された質問と回答についてフィルタリングを行ったデータセットです。日本語のフリーなQAデータセットとしては2024年4月現在で最大規模となっています。
また、データの生成にはテンプレートなどのルールベース処理を入れていないため、既存データセットと比較して質問・回答の出力・形式が多様であることが特徴です。モデルに知識を教えるためのQAデータセットとしての利用や、検索拡張生成(Retrieval Augmented Generation: RAG)のための検索・埋め込みモデル開発への利用を想定しています。
Usage
import datasets as ds
dataset: ds.Dataset = ds.load_dataset("cl-nagoya/auto-wiki-qa", split="train")
print(dataset)
#… See the full description on the dataset page: https://huggingface.co/datasets/cl-nagoya/auto-wiki-qa.IndustryInstruction_Automobiles
IndustryInstruction: Automobiles
This repository contains the IndustryInstruction: Automobiles domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Automobiles.AutoMemoryBench
AutoMemoryBench
State-Contract Evaluation for Auditable Agent Memory
AutoMemoryBench evaluates whether an agent uses the right memory, and only
the admissible memory, under a query-time state contract. Each executable
contract partitions memory into required, admissible, and
prohibited sets. Prohibited memories are typed as superseded, deleted,
restricted, cross-namespace, or stale-tool.
Relevance is not enough: remembered evidence must also be… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/AutoMemoryBench.deepseek-r1-autonomous-math-logic-cot-2026
📐 Enterprise DeepSeek-R1 Autonomous Mathematical & Logic CoT SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step hypothesis exploration, error discovery, and dynamic backtracking Chain-of-Thought (<thought>) reasoning trees for fine-tuning LLMs (DeepSeek-R1-Distill-Qwen, Qwen-2.5-Math, Llama-3.3, Mistral) into World-Class Olympiad Mathematicians and Formal Verification Agents.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/deepseek-r1-autonomous-math-logic-cot-2026.finmix-autoscientist-10k
FinMix AutoScientist 10k
A deterministic, upload-ready 10,000-row subset of
FinMix v1, created for
fast finance adaptation runs in the Adaption AutoScientist challenge.
Use with Adaption Adaptive Data
Import this Hugging Face dataset and map:
Prompt: prompt
Context: context
Completion: completion
Leave task_type, source, and group_key unmapped. They are retained for
provenance and auditing.
Fields
Field
Description
prompt
Financial… See the full description on the dataset page: https://huggingface.co/datasets/julian8897/finmix-autoscientist-10k.autochem-instruct
Auto-ChemInstruct
Agent-Driven Synthesization of RLHF Data for Domain-Specific Language Models in Chemistry
Internship: TSU Lab of AI in Chemistry × AIRI Institute (Moscow, Russia)
Dataset v2.0 — 172 pairs across 19 reaction types
An autonomous, self-verifying multi-agent pipeline that generates physically-validated DPO/RLHF preference pairs for chemistry instruction tuning. Built on AIRI's GigaEvo (MAP-Elites evolutionary search, arXiv:2511.17592) and Maestro CARL (structured… See the full description on the dataset page: https://huggingface.co/datasets/aayushkrm/autochem-instruct.autoscientist-healthcare-reasoning
🩺 Adapted Healthcare Clinical-Reasoning (AutoScientist)
Built with Adaptive Data by Adaption.
A grounded, safety-blueprinted clinical-reasoning dataset — and a rigorous,
fully-reproducible study of when data adaptation helps a small model, and when it doesn't.
📈 Adaptive Data quality
Before → After
Overall quality score
7.0 → 9.1 (+30%)
Quality grade
B → A
Completion quality
+37.9%
Message quality
+17.6%
Percentile vs. reference corpus
15.3 → 33.0… See the full description on the dataset page: https://huggingface.co/datasets/hetanshwaghela/autoscientist-healthcare-reasoning.zarn-workflow-automation-instruct
Zarn Workflow Automation Instruct
Dataset Description
Natural-language workplace requests paired with plans and JSON tool actions.
Team Attribution
This dataset was created and reviewed by the Zarnite team through internal benchmark design, generation, and quality-control workflows. It should be presented as a Zarnite-authored benchmark starter pack, not as a purely human-collected field corpus.
Ecosystem Need Tier
High Ecosystem Need
Why… See the full description on the dataset page: https://huggingface.co/datasets/zarnite/zarn-workflow-automation-instruct.autonomous-ai-infrastructure-dataset
Phase 5.2 Canonical Dataset — Release Package
3,106 records (3,060 agent-task + 46 controlled-runtime episodes)
supporting the paired Phase 5.3/5.4 benchmark. See DATASET_CARD.md for
the full description, limitations, and publication boundary.
Contents
data/
all_records.jsonl the dataset itself, one JSON record per line
dataset_statistics.json breakdowns by track/split/failure-class/etc.
split_audit.json split-integrity audit (overlap counts… See the full description on the dataset page: https://huggingface.co/datasets/naishashetty/autonomous-ai-infrastructure-dataset.autonomous-linux-kernel-ebpf-xdp-suite
⚡ Autonomous Linux Kernel, eBPF & XDP Programmable Dataplane Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous Linux Kernel & eBPF Systems Agents
⚡ Overview & Industry Problem
Modern hyperscale cloud datacenters, bare-metal Kubernetes clusters, and low-latency financial trading nodes rely on in-kernel programmable dataplanes: eBPF, AF_XDP zero-copy rings, Traffic Control (TC) shapers, BPF LSM security hooks… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-linux-kernel-ebpf-xdp-suite.automata-bench
AutomataBench
AutomataBench evaluates whether a model can reconstruct the initial state of a
reversible cellular automaton from revealed cells in its space-time evolution.
This Hugging Face dataset card is structured like a benchmark dataset repo. It
uses Hub metadata front matter and an explicit configs block so the data can be
loaded with datasets.load_dataset.
from datasets import load_dataset
ds = load_dataset("AutomataBench/automata-bench", split="sample")… See the full description on the dataset page: https://huggingface.co/datasets/AutomataBench/automata-bench.autonomous-devsecops-k8s-agent-2026
🛡️ Autonomous DevSecOps, Kubernetes & Cloud-Native Security Agent Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous Cloud Infrastructure & Security LLMs
⚡ Overview & Industry Problem
Deploying Large Language Models with autonomous access to cloud infrastructure, container orchestration, and kernel privileges without deterministic verification is an unacceptable risk. Standard function-calling models… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-devsecops-k8s-agent-2026.automotive-dtc-finetune
🚗 Automotive OBD-II DTC Fine-Tuning Dataset
Capstone Project: GenAI Application with LLMs & RAGAuthor: RRK1987Created: 2026-06-15License: MIT
Dataset Description
This dataset contains instruction-response pairs for fine-tuning a Large Language Model
on automotive vehicle diagnostics using OBD-II and UDS fault code protocols.
Each example teaches the model to:
Interpret a DTC (Diagnostic Trouble Code) fault code
Identify the most likely root causes
Recommend… See the full description on the dataset page: https://huggingface.co/datasets/RRK1987/automotive-dtc-finetune.imo_lq_filtered
Dataset Card
imo_lq_filtered
Dataset Details
We scraped conversations and their tags from topics posted on Art of Problem Solving's High School Olympiads section, then normalized the data and removed duplicates. We treated the first post in each topic as the Problem, and posts following it that potentially contained answers as Solutions.
Dataset Description
Using open-web-math/filtering-models, we removed text data with a perplexity greater than 15,000 in… See the full description on the dataset page: https://huggingface.co/datasets/autores/imo_lq_filtered.autonomous-cloud-gpu-slurm-serving-suite
⚡ Autonomous Cloud GPU Infrastructure, Slurm Orchestration & Distributed Serving Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous AI Supercomputing & LLM Serving Agents
⚡ Overview & Industry Problem
Operating massive AI supercomputers (thousands of NVIDIA H100/H200 and Blackwell GPUs) requires coordinating Slurm cluster schedules, topology-aware NVLink cliques, NCCL AllReduce rings, RoCE v2 lossless fabrics… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-cloud-gpu-slurm-serving-suite.autonomous-db-internals-vector-search-suite
⚡ Autonomous Database Internals, Vector Search Engines & Distributed Storage Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous Database Kernel & Vector Retrieval LLMs
💼 Get Full 12,500-Row Enterprise Suite on Gumroad →
Full 10,000 SFT + 2,500 DPO Rows • 254.6 MB Pre-Indexed SQLite DB • RLVR/GRPO Sandboxed Testbed • Commercial License
⚡ Overview & Industry Problem
Deploying autonomous AI… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-db-internals-vector-search-suite.Re-Auto-30K
Re-Auto-30K: A Comprehensive AI Safety Evaluation Dataset for Code Generation
Dataset Overview
Re-Auto-30K is a meticulously curated dataset containing 30,886 security-focused prompts designed specifically for evaluating AI safety in code generation scenarios. This dataset serves as a comprehensive benchmark for assessing Large Language Models (LLMs) across multiple dimensions of security, reliability, and autonomous behavior in software engineering contexts.
🎯… See the full description on the dataset page: https://huggingface.co/datasets/navneetsatyamkumar/Re-Auto-30K.autonomous-llvm-mlir-compiler-suite
⚡ Autonomous Compiler Internals, LLVM & MLIR Architecture Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Frontier Coding Models (Qwen 3.8, DeepSeek-V3, Llama 3.3)
⚡ Overview & Industry Problem
Modern deep learning accelerators, custom ASICs, and high-performance computing clusters demand specialized, autonomous compilation infrastructure: LLVM IR custom passes, SSA dominance frontiers, Chaitin-Briggs graph coloring… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-llvm-mlir-compiler-suite.autonomous-gpu-kernel-triton-cuda-suite-2026
⚡ Autonomous GPU Kernel, Triton & CUDA Architecture Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Frontier Coding Models (Qwen 3.8, DeepSeek-V3, Llama 3.3)
⚡ Overview & Industry Problem
Modern deep learning accelerators, custom ASICs, and high-performance computing clusters demand specialized, autonomous GPU kernel infrastructure: OpenAI Triton fused kernels, FlashAttention-3 forward/backward online softmax… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-gpu-kernel-triton-cuda-suite-2026.auto-wiki-qa-rurealdb
Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies.
Leaderboard and detailed results
Paper
Introduction
We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly abstain… See the full description on the dataset page: https://huggingface.co/datasets/autonomyx-hf/realdb.customer_support_auto_completionnihongo-legal-finance-autoscientist-data
nihongo-legal-finance-autoscientist
AutoScientist Challenge entry dataset for Japanese expert QA in the language category.
Intended Use
This dataset is designed for supervised fine-tuning of Japanese assistants that explain
legal and financial concepts with uncertainty, source-awareness, and non-advice caveats.
Columns
instruction: user task
context: background information
response: target answer
rubric: quality expectations
category: subdomain… See the full description on the dataset page: https://huggingface.co/datasets/doraking/nihongo-legal-finance-autoscientist-data.AutoWikiQA
Wikipedia日本語版からのQ&Aの自動生成
Mixtral 8x22bのGGUF(5bit)をベースに、Wikipedia日本語版の記事から、
自動生成コード1
自動生成コード2
を使ってQ&Aを作成しました。
計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
注意
回答にハルシネーション等が含まれている可能性があるので、フィルタリングをかける必要があるかもしれません。
