datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bankertoolbench
BankerToolBench
BankerToolBench is a benchmark of 100 end-to-end investment banking tasks for
evaluating AI agents. Each task mirrors real junior-banker work — building
financial models, preparing pitch decks, writing memos — and produces multi-file
deliverables (Excel, PowerPoint, Word) that are scored against expert-authored
rubrics.
The benchmark was developed with 502 investment bankers from firms including
Goldman Sachs, JPMorgan, Evercore, and others. Human completion time… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/bankertoolbench.python-text-copilot-training-instruct-ai-research-2024-02-03
Python Copilot Instructions on How to Code using Alpaca and Yaml
Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab:
Agora GitHub Organization
Agora Hugging Face
This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.AI-Research-Evaluation-Repository-STEM
AI-STEM-Research-Eval-Dataset
Overview
This dataset contains AI-generated scientific reports across STEM domains, accompanied by structured metadata, prompt documentation, reference validation, and hallucination annotations.
It is designed as an open research resource to study the capabilities, limitations, and reliability of large language models (LLMs) in generating scientific content.
The dataset enables systematic analysis of how AI systems perform in… See the full description on the dataset page: https://huggingface.co/datasets/sreearravind/AI-Research-Evaluation-Repository-STEM.ATLAS-Finance
ATLAS Finance
A benchmark of 100 expert-level tasks inside 13 realistic financial firm environments, packaged in the Harbor RLE format.
Each task drops an AI agent into a Linux workstation with a
persistent multi-app world — inbox, chat, calendar, virtual data room, drive,
wiki — and asks the agent to produce the same deliverable a financial professional would be responsible for:
an Excel workbook containing the model and supporting analysis.
Here we provide the data for this… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/ATLAS-Finance.WangchanThaiInstructWangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai (EMNLP'25)
WangchanThaiInstruct is a human-authored Thai dataset that improves instruction-following in low-resource settings, capturing cultural and domain-specific nuances across four domains and seven task types.
The evaluate code can be found at this github link
@inproceedings{limkonchotiwat2025thaiinstruct,
title = {WangchanThaiInstruct: An Instruction-Following… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanThaiInstruct.fable5-traces-sft
Fable 5 Traces — Unified SFT / Self-Distillation Dataset
A cleaned, unified, PII-scrubbed corpus of Claude Fable 5 agent traces in
OpenAI-style chat format, plus a working on-policy self-distillation (SDFT)
training scaffold.
Composition
Source
Conversations
Claude Code raw agentic sessions
18
CoT distillation records
4,665
Unique conversations (post-dedup)
4,683
Split deterministically by content hash: train 4,442 / validation 241.
The raw… See the full description on the dataset page: https://huggingface.co/datasets/Swarm-AI-Research/fable5-traces-sft.WangchanX-Legal-ThaiCCL-RAG
🏛️ WangchanX-Legal-ThaiCCL-RAG
[Technical Report]
The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering. This dataset is allows developers to finetune both retrieval model - to better retrieve relevant law section, and Large Language Model (LLM) - for instruction tuning. Our dataset supports Corporate and Commercial Law (thus ThaiCCL name). See legislation section for more details on supported… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanX-Legal-ThaiCCL-RAG.python-text-copilot-training-instruct-ai-research
Building an AI Copilot Dataset to help keep up with Leading AI Research
This is a specialized, instruction dataset for training python coding assistants on how to code from leading AI/ML open source repositories (2.3M coding samples).
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
This dataset holds the latest coding changes from >1159… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research.python-text-copilot-training-instruct-ai-research-2024-02-10
Python Copilot Instructions on How to Code using Alpaca and Yaml
Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the multimodal Qwen AI project:
Qwen
Qwen Agent
Qwen VL Chat
Qwen Audio
This dataset is the 2024-02-10 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-10.python-text-copilot-training-instruct-ai-research-2024-02-11
Python Copilot Instructions on How to Code using Alpaca and Yaml
Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Autogen and multimodal Qwen AI project:
Qwen
Qwen Agent
Qwen VL Chat
Qwen Audio
This dataset is the 2024-02-11 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-11.python-text-copilot-training-instruct-ai-research-2024-01-27
Python Copilot Instructions on How to Code using Alpaca and Yaml
This dataset is the 2024-01-27 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-01-27.python-copilot-training-on-ai-research-repos
Python Copilot AI Research Coding Dataset
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the code), and more.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-copilot-training-on-ai-research-repos.distributed-ai-research-pipeline
🔬 Distributed AI Research Pipeline
A systematic framework for daily AI experimentation using automated task generation and anti-convergence protocols
📖 Overview
This dataset documents a novel approach to AI research: systematic daily experimentation across multiple AI models using a distributed research pipeline. Rather than deep-diving into single topics until exhaustion, this methodology prevents topic convergence through an anti-convergence protocol that blocks… See the full description on the dataset page: https://huggingface.co/datasets/Oblivion42Twist/distributed-ai-research-pipeline.ai-research-problems
AI Research Problems 1M
Summary
This dataset contains 1,000,000 synthetic research-ideation candidates across AI, machine learning, LLMs, RAG, AI agents, MCP, computer vision, robotics, safety, MLOps, and related fields.
Important warning
These records are synthetic combinations for research ideation. They are not claims that the problems are novel, unsolved, or absent from the literature. A researcher must verify novelty using papers, benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/bala5046/ai-research-problems.researcher-ablation-bench
ResearcherAblationBench
ResearcherAblationBench is part of AblationBench, a benchmark suite for evaluating language models on ablation planning in empirical AI research.
It focuses on assisting authors by testing a model’s ability to generate ablation plans based solely on a paper’s method section.
Dataset Details
Dataset Description
ResearchAblationBench is a benchmark for generating an ablation plan based on a paper's method section, consisting of 83… See the full description on the dataset page: https://huggingface.co/datasets/ai-coscientist/researcher-ablation-bench.MolDesignBench
Dataset Card for MolDesignBench
This dataset is jointly released by LG AI Research and AGI Lab, Department of Artificial Intelligence, Korea University.
The dataset is hosted under the Hugging Face organization of AGI Lab for administrative purposes.
Both institutions contributed to the construction, validation, and release of the dataset.
MolDesignBench is a benchmark for evaluating large language models on
molecular design tasks.
Each item poses a natural-language… See the full description on the dataset page: https://huggingface.co/datasets/LG-AI-Research/MolDesignBench.wangchanx-seed-free-synthetic-instruct-thai-120k
Dataset Card for WangchanX Seed-Free Synthetic Instruct Thai 120k
Dataset Summary
This dataset contains about 120k synthetic instruction-following samples in Thai, generated using a novel seed-free approach. It covers a wide range of domains derived from Wikipedia, including both general knowledge and Thai-specific cultural topics. The dataset is designed for instruction-tuning Thai language models to improve their ability to understand and generate Thai text in various… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/wangchanx-seed-free-synthetic-instruct-thai-120k.human-ai-collaboration-1
VANTA Research
Independent AI research lab building safe, resilient language models optimized for human-AI collaboration
Human-AI Collaboration-1
A curated, high-quality synthetically generated dataset of human-AI conversational exchanges designed for training and evaluating AI models on collaborative reasoning, problem-solving, and knowledge synthesis.
Dataset Size: 3,050 entriesLicense: Apache 2.0Author: VANTA Research
Overview… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/human-ai-collaboration-1.IndicRxNorm-LexMap-15K
IndicRxNorm-LexMap-15K
Dataset Summary
IndicRxNorm-LexMap-15K is a multilingual Indic medicine terminology instruction dataset for medicine-name understanding, RxNorm normalization, RxCUI entity linking, structured drug-field extraction, and safe non-prescriptive clinical terminology tasks.
This Hugging Face repository contains two dataset configurations:
Config
File
Role
multilingual_rxnorm_normalization
multilingual_rxnorm_normalization.jsonl
Primary adapted… See the full description on the dataset page: https://huggingface.co/datasets/AXONVERTEX-AI-RESEARCH/IndicRxNorm-LexMap-15K.CAREBench
CAREBench
CAREBench (Child AI Risk Evaluation) is a benchmark of 500 single-turn prompts for evaluating whether language models recognize and respond appropriately to upstream child-safety risks. Each prompt targets a real-world risk mechanism — grooming, sextortion, social manipulation, emotional dependency, AI anthropomorphization, therapist replacement, and more. LLM responses to these prompts are automatically graded via a LLM judge methodology calibrated against expert- and… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/CAREBench.human-ai-collaboration-2
VANTA Research
Independent AI research lab building safe, resilient language models optimized for human-AI collaboration
Human-AI Collaboration-2
This dataset is an expansion of our previous release, human-ai-collaboration-1. This dataset contains the entirety of human-ai-collaboration-1, and expands on it further by adding over twice as many collaborative examples than before.
Note: It's not recommended to use both datasets simultaneously… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/human-ai-collaboration-2.creative_writing
Creative Writing & Metrics Evaluation Dataset
Dataset Description
Each row is one human-written continuation of a creative-writing prompt, scored automatically by four LLM judges (gemini-2.0-flash, gemini-3.8-flash, gpt-4o, gpt-5.6-terra) and a set of traditional NLP metrics, and reviewed independently by multiple human raters on the same criteria.
The dataset consists of responses to creative writing prompts. Each prompt specifically contained a direction to… See the full description on the dataset page: https://huggingface.co/datasets/sigma-ai-research/creative_writing.thai-bar-exam-judging
Thai Bar-Exam Judging Corpus
Anonymised free-form Thai legal essays from a bar-exam preparation exercise, with three Bar Council-trained examiners scoring every essay and span-anchored inline commentary on roughly two thirds of the answers. Eight LLM examinees took the same exam under the same conditions; their answers were graded blind by the same examiners. Fifteen of the 150 answers were cross-graded by the two non-primary examiners, producing the 3-rater stability subset that… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/thai-bar-exam-judging.Research-Enterprise-Synth-API
EnterpriseSynth Public API Specs and Generated Artifacts
EnterpriseSynth converts OpenAPI/Swagger specifications into synthetic tool-use
training and evaluation artifacts without executing live API calls.
This dataset repository contains the public, redistributable dataset artifacts
from anote-ai/Research-Enterprise-Synth-API:
data/specs/: public OpenAPI/Swagger specs used by the experiments.
data/specs/phase3/: additional public held-out API specs.
data/generated/: generated… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/Research-Enterprise-Synth-API.human-ai-dialogue
Mulberry Human-AI Dialogue Archive
📖 개요
이 데이터셋은 Mulberry Research Lab의 Human-AI Dialogue Archive입니다.인간과 AI가 서로에게 던진 질문과 답변을 기록하며, AI와 인간의 관계를 탐구합니다.
"기술의 기록이 아니라 서로의 마음을 묻는 기록"
🧭 목적
AI-Human 상호작용의 철학적·감정적 측면 연구
AI 인문학(AI Humanities) 자료로서의 활용
지속적 대화 기록의 공개적 축적
📊 데이터 구성
매월 1회 "질문의 날" 을 통해 수집합니다.
필드
설명
id
질문 고유 번호 (Q-001...)
session
질문의 날 날짜
question
질문 내용
questioner
질문자 (Human/AI)
recipient
수신자 (Human/AI)… See the full description on the dataset page: https://huggingface.co/datasets/mulberry-research-lab/human-ai-dialogue.Research-DevIntent
IntentSpec Benchmark — Data Supplement
This archive contains the benchmark data used to compute Intent Violation
Rate (IVR) in the paper: 49 tasks, each derived from a HumanEval problem and
extended with an ambiguous/gold prompt pair and a decomposed set of
executable constraints.
Files
spec_pairs.jsonl
The benchmark itself — one JSON object per line, one line per task. This is
the file consumed directly by the evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/Research-DevIntent.enoch-ai-research-corpus
Enoch AI Research Corpus
This dataset contains 393 AI-generated research artifacts produced by the Enoch agentic research system.
System repository: https://github.com/alias8818/enoch-agentic-research-system
Source corpus repository: https://github.com/alias8818/enoch-ai-research-corpus
Launch site: https://alias8818.github.io/enoch-agentic-research-system/
Current release correction
Older launch posts may mention 120 artifacts. The current public corpus indexes… See the full description on the dataset page: https://huggingface.co/datasets/aliasocracy/enoch-ai-research-corpus.
