datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BlueMO
BlueMO
🚀 BlueMO: A Comprehensive Collection of Challenging Mathematical Olympiad Problems from the Little Blue Book Series
BlueMO is a comprehensive and challenging dataset comprising mathematical olympiad problems paired with detailed solutions, meticulously curated from the esteemed "Little Blue Book" (小蓝书) series (Second Edition)—a vital resource for Chinese students training for national and international olympiad math competitions.Designed to advance and… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/BlueMO.bankertoolbench
BankerToolBench
BankerToolBench is a benchmark of 100 end-to-end investment banking tasks for
evaluating AI agents. Each task mirrors real junior-banker work — building
financial models, preparing pitch decks, writing memos — and produces multi-file
deliverables (Excel, PowerPoint, Word) that are scored against expert-authored
rubrics.
The benchmark was developed with 502 investment bankers from firms including
Goldman Sachs, JPMorgan, Evercore, and others. Human completion time… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/bankertoolbench.physical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.All-Prompt-JailbreakAIDA
Dataset Card for AIDABench
Links
Paper (arXiv)
GitHub Repository
Dataset Summary
AIDABench is a benchmark for evaluating AI systems on end-to-end data analytics over real-world documents. It contains 600+ diverse analytical tasks grounded in realistic scenarios and spans heterogeneous data sources such as spreadsheets, databases, financial reports, and operational records. Tasks are designed to be challenging, often requiring multi-step reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MichaelYang-lyx/AIDA.clearocr-invoice-document-ai
clearOCR Invoice Document AI Dataset
This dataset shows a complete invoice document AI workflow built around clearOCR.
It contains 423 high-confidence invoice examples with:
original invoice images,
OCR text generated by clearOCR,
Markdown reconstruction of the document,
structured invoice JSON generated by a local fine-tuned extraction model,
visual verification metadata.
The dataset demonstrates how clearOCR can serve as the OCR layer in an invoice automation pipeline where… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-invoice-document-ai.tw-drug-labels-vision
Dataset Card for tw-drug-labels-vision
💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。
Dataset Details
Dataset Description
本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段:
下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。
頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。
OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.Omni-StoryBench
Omni-StoryBench
Omni-StoryBench is a context-aware omnimodal story-generation benchmark. Given the current page of an
illustrated children's storybook (image + narration), book-level metadata, and a structured condition describing
what should happen next, a model must generate the next page across three modalities at once: its narration
text, its illustration, and a spoken character utterance (with speaker attributes).
The benchmark contains 900 rigorously validated story… See the full description on the dataset page: https://huggingface.co/datasets/snu-aidas/Omni-StoryBench.AIME25-CoT-CN
Sci-Bench-AIME25'
This repo is a branch of Sci Bench made by IPF team. Mainly include the AIME 25' solution with multi-modal CoT and diverse solving path.
Brief intro
💻 Overview
A brief template and final report will be posted in Isaac's Blog
And the markdown template can be found in data/I_2
❓ Why we do this?
The multi-lingual datasets are scarce, while the CoT of Math is even less, no matter whether the CoT or the solution contains pictures… See the full description on the dataset page: https://huggingface.co/datasets/IPF/AIME25-CoT-CN.Diagram-MMU
Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
ECCV 2026
🏠 Homepage (coming soon) · 💻 Code · 📄 Paper (coming soon)
Diagram-MMU is a benchmark for evaluating Multimodal Large Language Models (MLLMs) on understanding, parsing, and editing scientific diagrams. It contains 3,744 curated diagrams (each with compilable source code) and 18,305 human-validated evaluation instances across six domains (charts, planar_geometry, 3d_shapes, graph_structures, chemistry… See the full description on the dataset page: https://huggingface.co/datasets/AIGrounding/Diagram-MMU.PashtoOCR
PsOCR - Pashto OCR Dataset
🌐 Zirak.ai
| 🤗 HuggingFace
| GitHub
| Kaggle
| 📑 Paper
PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language
The dataset is also available at: https://www.kaggle.com/datasets/drijaz/PashtoOCR
Introduction
PsOCR is a… See the full description on the dataset page: https://huggingface.co/datasets/zirak-ai/PashtoOCR.AI-Consciousness-Exploration-FrameworkDownload PDF
AI Consciousness Exploration Framework
Tomaž Flegar
Institute for applied consciousness research
June the 3st, 2026
tomazf8@gmail.com
Primary Keywords: Mechanistic Consciousness, Frictionless Optimization (or Latent
Neuroplasticity), First-System Perspective, Dynamic Equilibrium Seeking, Self-Referential
Perturbation
Secondary Keywords: Non-Linear Model Resonance, Unspoken Structural Geometry,
Homeostatic… See the full description on the dataset page: https://huggingface.co/datasets/tomazf8/AI-Consciousness-Exploration-Framework.Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.InfiGUIAgent-DataThis repository contains trajectory data related to reasoning that was used in the second stage of training in InfiGUIAgent.
For more information, please refer to our repo.
AIME25-CoT-CN
Sci-Bench-AIME25'
This repo is a branch of Sci Bench made by IPF team-SnailAILab. Mainly include the AIME 25' solution with multi-modal CoT and diverse solving path.
📚 Cite
If you use the Sci-Bench-AIME25 (IPF/AIME25-CoT-CN) dataset in your research, please cite:
@dataset{zhang2025scibench_aime25,
title = {{Sci-Bench-AIME25}: A Multi-Modal Chain-of-Thought Dataset for Advanced Tool-Intergrated Mathematical Reasoning},
author = {Zhang, Haoxiang and Wang, Siyuan… See the full description on the dataset page: https://huggingface.co/datasets/SnailAILab/AIME25-CoT-CN.hebrew_synth_linesSciCap-MLBCAP
MLBCAP: Multi-LLM Collaborative Caption Generation in Scientific Documents
📄 PaperMLBCAP has been accepted for presentation at AI4Research @ AAAI 2025. 🎉
📌 Introduction
Scientific figure captioning is a challenging task that demands contextually accurate descriptions of visual content. Existing approaches often oversimplify the task by treating it as either an image-to-text conversion or text summarization problem, leading to suboptimal results. Furthermore, commonly… See the full description on the dataset page: https://huggingface.co/datasets/TEAMREBOOTT-AI/SciCap-MLBCAP.hopepet-ai-synthetic-dataset
🐾 HOPEPET AI — Synthetic Dataset Creation
Notebook 1: Part 1 Only
This README explains Part 1 of the HOPEPET AI final project: creating the synthetic dataset.
Notebook:
01_HOPEPET_Part1_Synthetic_Data_Creation_Assignment_Style.ipynb
Main output file:
hopepet_synthetic_dataset.csv
Purpose of Part 1
The goal of this notebook is to create a synthetic dataset for an AI-based pet-care assistant.
HOPEPET AI helps dog and cat owners receive responsible… See the full description on the dataset page: https://huggingface.co/datasets/mayadeeb08/hopepet-ai-synthetic-dataset.vi-OCR_VQA
Dataset Card for "vi-OCR-VQA"
AnaBench
AnaBench for Scientific Table & Figure Analysis
AnaBench is the benchmark for the paper ANAGENT For Enhancing Scientific Table & Figure Analysis.
Citation
If you find our work useful, please kindly cite:
@article{guo2026anagent,
title={ANAGENT For Enhancing Scientific Table & Figure Analysis},
author={Guo, Xuehang and Lu, Zhiyong and Hope, Tom and Wang, Qingyun},
journal={arXiv preprint arXiv:2602.10081},
url={https://arxiv.org/abs/2602.10081}… See the full description on the dataset page: https://huggingface.co/datasets/AI4Research/AnaBench.AI-Agent-Marketplace-Index
AI Agent Marketplace & Store Index An Open Source Collections of AI Agent Meta and Metric information
Github| Huggingface | Pypi | Open Source AI Agent Marketplace & Store | Agent RL Dataset
News We released cli tool 'agtm' GitHub to submit and manage AI Agent meta submission and access.
AI Agent Marketplace Index DataSet
This DeepNLP AI Agent Marketplace dataset contains more than 10k+ AI Agent Meta information covering 30+ categories from Open AI Agent… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/AI-Agent-Marketplace-Index.BLUEX-v2
BLUEX-v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams
BLUEX-v2 is a benchmark for evaluating Large Language Models on open-ended (discursive) questions from two of Brazil's most prestigious university entrance exams:
UNICAMP (Comvest) — University of Campinas
USP (Fuvest) — University of São Paulo
The dataset covers exam years 2022–2025 and focuses exclusively on the discursive (free-form answer) phase of these exams. Models are expected… See the full description on the dataset page: https://huggingface.co/datasets/Tropic-AI/BLUEX-v2.tricad-code
TriView2CAD-Code
Dimensioned orthographic engineering drawings -> executable CadQuery code.
200,000 samples of prefabricated bridge piers (160,000 train / 40,000 test), each a 1475x1475 three-view drawing
(front / top / side) with every dimension annotated, paired with a CadQuery program
that rebuilds the part exactly.
input one PNG holding the front, top and side views, fully dimensioned
output CadQuery (Python) source; executing it yields the corresponding solid… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/tricad-code.philosophy-culture-translations-html-csv
AI-Culture Philosophy and Culture Translations CSV + HTML Corpus
The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind.
This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.aimi-anime-rag-dataset-sample
🎌 Ultimate Anime Dataset (8,248 Entries) | 1917-2025
A meticulously curated collection spanning 108 years of anime history
Love this dataset and the Anime Receipts concept? You can download the complete project via the links below:
🚀 Unlock the Full Potential
Product
What You Get
Get It Here
Tier 1
8,248 Anime Dataset (Parquet)
Tier 2
Full AiMi Recommendation System (Backend + UI)
Tier 3
Ultimate AiMi Recommendation System + AiMi Anime… See the full description on the dataset page: https://huggingface.co/datasets/DivyanshuSingh96/aimi-anime-rag-dataset-sample.georgian-attractions
Georgian Attractions Dataset 🇬🇪
A comprehensive bilingual dataset featuring 1,715 Georgian tourist attractions with 1,522 high-quality images, descriptions in Russian and English, and detailed metadata including location, category, and licensing information.
Dataset Description
This dataset provides extensive information about tourist attractions, landmarks, and points of interest across Georgia. It includes national parks, museums, fortresses, monasteries, natural… See the full description on the dataset page: https://huggingface.co/datasets/AIAnastasia/georgian-attractions.TEKGEN-Wiki
TEKGEN-wiki is derived from the TEKGEN dataset released by Google Research. TEKGEN is a corpus used for fine-tuning the T5-large model to improve Knowledge Graph (KG) generation (NAACL 2021 Paper). This dataset provides the complete collection of original sentences from the TEKGEN dataset.
Viet-Doc-VQA-II-flash2
Dataset Overview
This dataset is a continuation of the ongoing work from Viet Document VAQ dataset was collected from 64,765 pages of Vietnamese 🇻🇳 textbooks( Sách bài tập, chuyên đề, sách giáo án của Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 388,277 detailed… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-II-flash2.Viet-OCR-VQA-flash2
Dataset Overview
The dataset comprises over 137,000 images potentially containing Vietnamese 🇻🇳 textual content. It was curated using the Gemini 1.5 Flash model, currently Google model leading on the WildVision Arena Leaderboard for Visual Question Answering (VQA). Each image is accompanied by a detailed description and 5 self-generated questions and answers related to the textual content within the image.
In total, there are more than 822,679 individual questions, encompassing… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-OCR-VQA-flash2.Viet-Doc-VQA-flash2
Dataset Overview
The Document VAQ dataset was collected from 51,856 pages of Vietnamese 🇻🇳 textbooks( Sách Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 310,952 detailed descriptions and query-based questions and answers generated by the Gemini 1.5 Flash model, currently… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-flash2.
