datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-research-index-iclr-openreview
Conference-paper corpus (AI-Research-Index project)
Private working dataset. Conference/journal paper corpus across OpenReview
venues (ICLR, NeurIPS incl. D&B/position tracks, ICML incl. position, COLM,
TMLR, AISTATS, UAI, ALT, MathAI, and later additions) plus the ACL Anthology
family. Coverage, per-venue availability, decisions semantics, and known
biases are documented authoritatively in the GitHub repo's data/README.md
— read that first; per-venue counts change as the corpus… See the full description on the dataset page: https://huggingface.co/datasets/latkes/ai-research-index-iclr-openreview.bankertoolbench
BankerToolBench
BankerToolBench is a benchmark of 100 end-to-end investment banking tasks for
evaluating AI agents. Each task mirrors real junior-banker work — building
financial models, preparing pitch decks, writing memos — and produces multi-file
deliverables (Excel, PowerPoint, Word) that are scored against expert-authored
rubrics.
The benchmark was developed with 502 investment bankers from firms including
Goldman Sachs, JPMorgan, Evercore, and others. Human completion time… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/bankertoolbench.FinDER
FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation
FinDER is a benchmark dataset designed for evaluating Retrieval-Augmented Generation (RAG) in financial question answering. It consists of 5,703 expert-annotated query–evidence–answer triplets derived from real-world 10-K filings and ambiguous financial queries submitted by industry professionals.
This dataset captures the domain-specific challenges of financial QA, including short… See the full description on the dataset page: https://huggingface.co/datasets/Linq-AI-Research/FinDER.StreamBench
StreamBench paper link: https://arxiv.org/abs/2406.08747 (The links for the original raw datasets on StreamBench can be found in Appendix F)
If you find our work helpful, please cite as:
@article{wu2024streambench,
title={StreamBench: Towards Benchmarking Continuous Improvement of Language Agents},
author={Wu, Cheng-Kuang and Tam, Zhi Rui and Lin, Chieh-Yen and Chen, Yun-Nung and Lee, Hung-yi},
journal={arXiv preprint arXiv:2406.08747},
year={2024}
}
python-text-copilot-training-instruct-ai-research-2024-02-03
Python Copilot Instructions on How to Code using Alpaca and Yaml
Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab:
Agora GitHub Organization
Agora Hugging Face
This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.movie_gen_video_bench
Dataset Card for the Movie Gen Benchmark
Movie Gen is a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio.
Here, we introduce our evaluation benchmark "Movie Gen Bench Video Bench", as detailed in the Movie Gen technical report (Section 3.5.2).
To enable fair and easy comparison to Movie Gen for future works on these evaluation benchmarks, we additionally release the non cherry-picked generated videos from… See the full description on the dataset page: https://huggingface.co/datasets/meta-ai-for-media-research/movie_gen_video_bench.PANORAMArobust-finetuningPlease refer to the following source for the original datasets:
GSM8K: https://huggingface.co/datasets/openai/gsm8k
MATH: https://huggingface.co/datasets/hendrycks/competition_math
math-resample: In this section we subsample the 1,000 subsample only (yes it's balance)
HumanEval+: https://huggingface.co/datasets/evalplus/humanevalplus
MBPP: https://huggingface.co/datasets/google-research-datasets/mbpp
MBPP+: https://huggingface.co/datasets/evalplus/mbppplus
ARC Challenge:… See the full description on the dataset page: https://huggingface.co/datasets/appier-ai-research/robust-finetuning.QIVD
QIVD: Qualcomm Interactive Video Dataset
A collection of 2,900 video clips paired with visual question-answer annotations.
Each clip is associated with exactly one question drawn from one of 13 fine-grained QA categories,
a full-sentence answer, a concise short answer, and a timestamp pinpointing the relevant moment in the video.
Overview
QIVD is a dataset and benchmark for online, situated audio-visual question answering. Unlike existing video QA benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/Qualcomm-AI-Research/QIVD.ResearchArcade-openreview-reviewsATLAS-Finance
ATLAS Finance
A benchmark of 100 expert-level tasks inside 13 realistic financial firm environments, packaged in the Harbor RLE format.
Each task drops an AI agent into a Linux workstation with a
persistent multi-app world — inbox, chat, calendar, virtual data room, drive,
wiki — and asks the agent to produce the same deliverable a financial professional would be responsible for:
an Excel workbook containing the model and supporting analysis.
Here we provide the data for this… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/ATLAS-Finance.ResearchArcade-openreview-papersResearchArcade-openreview-authorsWangchanThaiInstructWangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai (EMNLP'25)
WangchanThaiInstruct is a human-authored Thai dataset that improves instruction-following in low-resource settings, capturing cultural and domain-specific nuances across four domains and seven task types.
The evaluate code can be found at this github link
@inproceedings{limkonchotiwat2025thaiinstruct,
title = {WangchanThaiInstruct: An Instruction-Following… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanThaiInstruct.thai-ser
🇹🇭 THAI-SER Dataset 🎭
[📝 Paper (preprint)]
Published by: AI Research Institute of Thailand (AIResearch)
In collaboration with:
Vidyasirimedhi Institute of Science and Technology (VISTEC)
Digital Economy Promotion Agency (depa)
Department of Computer Engineering, Faculty of Engineering, Chulalongkorn University
Department of Dramatic Arts, Faculty of Arts, Chulalongkorn University
Sponsored by: Advanced Info Services Public Company Limited (AIS), and Siam Commercial… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/thai-ser.fable5-traces-sft
Fable 5 Traces — Unified SFT / Self-Distillation Dataset
A cleaned, unified, PII-scrubbed corpus of Claude Fable 5 agent traces in
OpenAI-style chat format, plus a working on-policy self-distillation (SDFT)
training scaffold.
Composition
Source
Conversations
Claude Code raw agentic sessions
18
CoT distillation records
4,665
Unique conversations (post-dedup)
4,683
Split deterministically by content hash: train 4,442 / validation 241.
The raw… See the full description on the dataset page: https://huggingface.co/datasets/Swarm-AI-Research/fable5-traces-sft.ResearchArcade-openreview-papers-authorsWangchanX-Legal-ThaiCCL-RAG
🏛️ WangchanX-Legal-ThaiCCL-RAG
[Technical Report]
The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering. This dataset is allows developers to finetune both retrieval model - to better retrieve relevant law section, and Large Language Model (LLM) - for instruction tuning. Our dataset supports Corporate and Commercial Law (thus ThaiCCL name). See legislation section for more details on supported… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanX-Legal-ThaiCCL-RAG.WangchanThaiMedicalstudentbench
StudentBench
StudentBench: AI and human tutoring yield equivalent GRE learning gains
Paper · Reproduction code · Project · Files
StudentBench measures how well AI tutors help real students learn. We release the study data to reproduce the paper's results and support open research on learning, lesson planning, practice problems, tutoring conversations, engagement and cost.
Pooled AI tutoring and expert human tutoring produced equivalent GRE learning gains ($p = .015$). Six AI… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/studentbench.python-text-copilot-training-instruct-ai-research
Building an AI Copilot Dataset to help keep up with Leading AI Research
This is a specialized, instruction dataset for training python coding assistants on how to code from leading AI/ML open source repositories (2.3M coding samples).
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
This dataset holds the latest coding changes from >1159… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research.python-text-copilot-training-instruct-ai-research-2024-02-10
Python Copilot Instructions on How to Code using Alpaca and Yaml
Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the multimodal Qwen AI project:
Qwen
Qwen Agent
Qwen VL Chat
Qwen Audio
This dataset is the 2024-02-10 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-10.python-text-copilot-training-instruct-ai-research-2024-02-11
Python Copilot Instructions on How to Code using Alpaca and Yaml
Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Autogen and multimodal Qwen AI project:
Qwen
Qwen Agent
Qwen VL Chat
Qwen Audio
This dataset is the 2024-02-11 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-11.bank-77
Know-No : Bank-77 Dataset
Src: https://github.com/xhz0809/Know-No
ai-research-2026
AI Research Papers 2026
Academic AI research, papers, findings. Updated daily via automated collection pipeline.
Part of the Legion Data Factory — historical AI ecosystem datasets 2026.
Methodology
Automated collection from public sources (HackerNews, RSS feeds, APIs).
Updated daily via cron job. Raw data, minimal processing.
License
CC BY 4.0
🔑 API Access — Updated Daily
Live data via Legion AI API | Documentation
Free: 100 req/day… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-research-2026.python-copilot-training-on-ai-research-repos
Python Copilot AI Research Coding Dataset
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the code), and more.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-copilot-training-on-ai-research-repos.WangchanX-FLAN-v6.1
[!NOTE]
The script for creates a dataset can be found at FLAN-like Dataset Creator
🔎 Dataset Details
An overview of a curated collection of datasets designed for natural language processing tasks, with a focus on Thai language applications. These datasets span a range of tasks including Summarization, Translation, Text Generation, Text Classification, and Question Answering. Each dataset is accompanied by its source, size, intended task, and licensing information, making it easy… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanX-FLAN-v6.1.python-text-copilot-training-instruct-ai-research-2024-01-27
Python Copilot Instructions on How to Code using Alpaca and Yaml
This dataset is the 2024-01-27 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-01-27.ResearchArcade-openreview-arxivResearchArcade-openreview-paragraphs
