CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01math-ai /BlueMO BlueMO 🚀 BlueMO: A Comprehensive Collection of Challenging Mathematical Olympiad Problems from the Little Blue Book Series   BlueMO is a comprehensive and challenging dataset comprising mathematical olympiad problems paired with detailed solutions, meticulously curated from the esteemed "Little Blue Book" (小蓝书) series (Second Edition)—a vital resource for Chinese students training for national and international olympiad math competitions.Designed to advance and… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/BlueMO.imagequestion-answering1K<n<10K3 likes9.4k downloads8mo agoHugging Face02handshake-ai-research /bankertoolbench BankerToolBench BankerToolBench is a benchmark of 100 end-to-end investment banking tasks for evaluating AI agents. Each task mirrors real junior-banker work — building financial models, preparing pitch decks, writing memos — and produces multi-file deliverables (Excel, PowerPoint, Word) that are scored against expert-authored rubrics. The benchmark was developed with 502 investment bankers from firms including Goldman Sachs, JPMorgan, Evercore, and others. Human completion time… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/bankertoolbench.documenttext-generationn<1K9 likes2.9k downloads4mo agoHugging Face03shi-labs /physical-ai-bench-generation Physical AI Bench - Generation Paper | Code Dataset Description The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.imagevisual-question-answering1K<n<10K5 likes2.8k downloads10mo agoHugging Face04AiActivity /All-Prompt-Jailbreakimagetext-generationn<1K10 likes1.5k downloads1y agoHugging Face05MichaelYang-lyx /AIDA Dataset Card for AIDABench Links Paper (arXiv) GitHub Repository Dataset Summary AIDABench is a benchmark for evaluating AI systems on end-to-end data analytics over real-world documents. It contains 600+ diverse analytical tasks grounded in realistic scenarios and spans heterogeneous data sources such as spreadsheets, databases, financial reports, and operational records. Tasks are designed to be challenging, often requiring multi-step reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MichaelYang-lyx/AIDA.imagequestion-answering1K<n<10K4 likes1k downloads3mo agoHugging Face06Lukaszl /clearocr-invoice-document-ai clearOCR Invoice Document AI Dataset This dataset shows a complete invoice document AI workflow built around clearOCR. It contains 423 high-confidence invoice examples with: original invoice images, OCR text generated by clearOCR, Markdown reconstruction of the document, structured invoice JSON generated by a local fine-tuned extraction model, visual verification metadata. The dataset demonstrates how clearOCR can serve as the OCR layer in an invoice automation pipeline where… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-invoice-document-ai.imageimage-to-textn<1K0 likes443 downloads4mo agoHugging Face07twinkle-ai /tw-drug-labels-vision Dataset Card for tw-drug-labels-vision 💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。 Dataset Details Dataset Description 本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段: 下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。 頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。 OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.imageimage-to-text10K<n<100K4 likes430 downloads5mo agoHugging Face08snu-aidas /Omni-StoryBench Omni-StoryBench Omni-StoryBench is a context-aware omnimodal story-generation benchmark. Given the current page of an illustrated children's storybook (image + narration), book-level metadata, and a structured condition describing what should happen next, a model must generate the next page across three modalities at once: its narration text, its illustration, and a spoken character utterance (with speaker attributes). The benchmark contains 900 rigorously validated story… See the full description on the dataset page: https://huggingface.co/datasets/snu-aidas/Omni-StoryBench.imagetext-generationn<1K0 likes337 downloads4d agoHugging Face09IPF /AIME25-CoT-CN Sci-Bench-AIME25' This repo is a branch of Sci Bench made by IPF team. Mainly include the AIME 25' solution with multi-modal CoT and diverse solving path. Brief intro 💻 Overview A brief template and final report will be posted in Isaac's Blog And the markdown template can be found in data/I_2 ❓ Why we do this? The multi-lingual datasets are scarce, while the CoT of Math is even less, no matter whether the CoT or the solution contains pictures… See the full description on the dataset page: https://huggingface.co/datasets/IPF/AIME25-CoT-CN.imagequestion-answeringn<1K10 likes314 downloads7mo agoHugging Face10AIGrounding /Diagram-MMU Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams ECCV 2026 🏠 Homepage (coming soon) · 💻 Code · 📄 Paper (coming soon) Diagram-MMU is a benchmark for evaluating Multimodal Large Language Models (MLLMs) on understanding, parsing, and editing scientific diagrams. It contains 3,744 curated diagrams (each with compilable source code) and 18,305 human-validated evaluation instances across six domains (charts, planar_geometry, 3d_shapes, graph_structures, chemistry… See the full description on the dataset page: https://huggingface.co/datasets/AIGrounding/Diagram-MMU.imageimage-to-text10K<n<100K0 likes289 downloads3mo agoHugging Face11zirak-ai /PashtoOCR PsOCR - Pashto OCR Dataset 🌐 Zirak.ai &nbsp;&nbsp; | &nbsp;&nbsp;🤗 HuggingFace &nbsp;&nbsp; | &nbsp;&nbsp; GitHub &nbsp;&nbsp; | &nbsp;&nbsp; Kaggle &nbsp;&nbsp; | &nbsp;&nbsp;📑 Paper PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language The dataset is also available at: https://www.kaggle.com/datasets/drijaz/PashtoOCR Introduction PsOCR is a… See the full description on the dataset page: https://huggingface.co/datasets/zirak-ai/PashtoOCR.imagetext-generation10K<n<100K6 likes228 downloads2mo agoHugging Face12tomazf8 /AI-Consciousness-Exploration-FrameworkDownload PDF AI Consciousness Exploration Framework Tomaž Flegar Institute for applied consciousness research June the 3st, 2026 tomazf8@gmail.com Primary Keywords: Mechanistic Consciousness, Frictionless Optimization (or Latent Neuroplasticity), First-System Perspective, Dynamic Equilibrium Seeking, Self-Referential Perturbation Secondary Keywords: Non-Linear Model Resonance, Unspoken Structural Geometry, Homeostatic… See the full description on the dataset page: https://huggingface.co/datasets/tomazf8/AI-Consciousness-Exploration-Framework.documenttext-generationn<1K0 likes148 downloads4mo agoHugging Face13jojo-ai-mst /Myanmar-Tuberculosis-Guidelines-Instructions Myanmar Tuberculosis Guidelines Instructions A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages. Authors: Min Si Thu, Khin Myat Noe Abstract Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.imagequestion-answering1K<n<10K1 likes146 downloads5mo agoHugging Face14InfiX-ai /InfiGUIAgent-DataThis repository contains trajectory data related to reasoning that was used in the second stage of training in InfiGUIAgent. For more information, please refer to our repo. imagetext-generation1K<n<10K6 likes120 downloads2y agoHugging Face15SnailAILab /AIME25-CoT-CN Sci-Bench-AIME25' This repo is a branch of Sci Bench made by IPF team-SnailAILab. Mainly include the AIME 25' solution with multi-modal CoT and diverse solving path. 📚 Cite If you use the Sci-Bench-AIME25 (IPF/AIME25-CoT-CN) dataset in your research, please cite: @dataset{zhang2025scibench_aime25, title = {{Sci-Bench-AIME25}: A Multi-Modal Chain-of-Thought Dataset for Advanced Tool-Intergrated Mathematical Reasoning}, author = {Zhang, Haoxiang and Wang, Siyuan… See the full description on the dataset page: https://huggingface.co/datasets/SnailAILab/AIME25-CoT-CN.imagequestion-answeringn<1K1 likes120 downloads1y agoHugging Face16samaritan-ai /hebrew_synth_linesimagetext-generation100K<n<1M1 likes104 downloads1y agoHugging Face17TEAMREBOOTT-AI /SciCap-MLBCAP MLBCAP: Multi-LLM Collaborative Caption Generation in Scientific Documents 📄 PaperMLBCAP has been accepted for presentation at AI4Research @ AAAI 2025. 🎉 📌 Introduction Scientific figure captioning is a challenging task that demands contextually accurate descriptions of visual content. Existing approaches often oversimplify the task by treating it as either an image-to-text conversion or text summarization problem, leading to suboptimal results. Furthermore, commonly… See the full description on the dataset page: https://huggingface.co/datasets/TEAMREBOOTT-AI/SciCap-MLBCAP.imagetext-generation10K<n<100K19 likes97 downloads2y agoHugging Face18mayadeeb08 /hopepet-ai-synthetic-dataset 🐾 HOPEPET AI — Synthetic Dataset Creation Notebook 1: Part 1 Only This README explains Part 1 of the HOPEPET AI final project: creating the synthetic dataset. Notebook: 01_HOPEPET_Part1_Synthetic_Data_Creation_Assignment_Style.ipynb Main output file: hopepet_synthetic_dataset.csv Purpose of Part 1 The goal of this notebook is to create a synthetic dataset for an AI-based pet-care assistant. HOPEPET AI helps dog and cat owners receive responsible… See the full description on the dataset page: https://huggingface.co/datasets/mayadeeb08/hopepet-ai-synthetic-dataset.imagetext-classificationn<1K0 likes62 downloads2mo agoHugging Face19LR-AI-Labs /vi-OCR_VQA Dataset Card for "vi-OCR-VQA" imagevisual-question-answering10K<n<100K7 likes61 downloads2y agoHugging Face20AI4Research /AnaBench AnaBench for Scientific Table & Figure Analysis AnaBench is the benchmark for the paper ANAGENT For Enhancing Scientific Table & Figure Analysis. Citation If you find our work useful, please kindly cite: @article{guo2026anagent, title={ANAGENT For Enhancing Scientific Table & Figure Analysis}, author={Guo, Xuehang and Lu, Zhiyong and Hope, Tom and Wang, Qingyun}, journal={arXiv preprint arXiv:2602.10081}, url={https://arxiv.org/abs/2602.10081}… See the full description on the dataset page: https://huggingface.co/datasets/AI4Research/AnaBench.imageimage-text-to-text10K<n<100K0 likes61 downloads1mo agoHugging Face21DeepNLP /AI-Agent-Marketplace-Index AI Agent Marketplace & Store Index An Open Source Collections of AI Agent Meta and Metric information Github| Huggingface | Pypi | Open Source AI Agent Marketplace & Store | Agent RL Dataset News We released cli tool 'agtm' GitHub to submit and manage AI Agent meta submission and access. AI Agent Marketplace Index DataSet This DeepNLP AI Agent Marketplace dataset contains more than 10k+ AI Agent Meta information covering 30+ categories from Open AI Agent… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/AI-Agent-Marketplace-Index.imagetext-generation2 likes56 downloads2mo agoHugging Face22Tropic-AI /BLUEX-v2 BLUEX-v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams BLUEX-v2 is a benchmark for evaluating Large Language Models on open-ended (discursive) questions from two of Brazil's most prestigious university entrance exams: UNICAMP (Comvest) — University of Campinas USP (Fuvest) — University of São Paulo The dataset covers exam years 2022–2025 and focuses exclusively on the discursive (free-form answer) phase of these exams. Models are expected… See the full description on the dataset page: https://huggingface.co/datasets/Tropic-AI/BLUEX-v2.imagequestion-answeringn<1K0 likes49 downloads3mo agoHugging Face23AI4Manufacturing /tricad-codegated TriView2CAD-Code Dimensioned orthographic engineering drawings -> executable CadQuery code. 200,000 samples of prefabricated bridge piers (160,000 train / 40,000 test), each a 1475x1475 three-view drawing (front / top / side) with every dimension annotated, paired with a CadQuery program that rebuilds the part exactly. input one PNG holding the front, top and side views, fully dimensioned output CadQuery (Python) source; executing it yields the corresponding solid… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/tricad-code.imageimage-to-text100K<n<1M0 likes35 downloads7d agoHugging Face24AI-Culture-Commons /philosophy-culture-translations-html-csv AI-Culture Philosophy and Culture Translations CSV + HTML Corpus The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind. This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.imagetranslation1K<n<10K2 likes30 downloads1y agoHugging Face25DivyanshuSingh96 /aimi-anime-rag-dataset-sample 🎌 Ultimate Anime Dataset (8,248 Entries) | 1917-2025 A meticulously curated collection spanning 108 years of anime history Love this dataset and the Anime Receipts concept? You can download the complete project via the links below: 🚀 Unlock the Full Potential Product What You Get Get It Here Tier 1 8,248 Anime Dataset (Parquet) Tier 2 Full AiMi Recommendation System (Backend + UI) Tier 3 Ultimate AiMi Recommendation System + AiMi Anime… See the full description on the dataset page: https://huggingface.co/datasets/DivyanshuSingh96/aimi-anime-rag-dataset-sample.imagetext-retrievaln<1K5 likes23 downloads10mo agoHugging Face26AIAnastasia /georgian-attractions Georgian Attractions Dataset 🇬🇪 A comprehensive bilingual dataset featuring 1,715 Georgian tourist attractions with 1,522 high-quality images, descriptions in Russian and English, and detailed metadata including location, category, and licensing information. Dataset Description This dataset provides extensive information about tourist attractions, landmarks, and points of interest across Georgia. It includes national parks, museums, fortresses, monasteries, natural… See the full description on the dataset page: https://huggingface.co/datasets/AIAnastasia/georgian-attractions.imageimage-classification1K<n<10K0 likes22 downloads10mo agoHugging Face27sleeping-ai /TEKGEN-Wiki TEKGEN-wiki is derived from the TEKGEN dataset released by Google Research. TEKGEN is a corpus used for fine-tuning the T5-large model to improve Knowledge Graph (KG) generation (NAACL 2021 Paper). This dataset provides the complete collection of original sentences from the TEKGEN dataset. imagetext-generation1M<n<10M0 likes21 downloads2y agoHugging Face285CD-AI /Viet-Doc-VQA-II-flash2gated Dataset Overview This dataset is a continuation of the ongoing work from Viet Document VAQ dataset was collected from 64,765 pages of Vietnamese 🇻🇳 textbooks( Sách bài tập, chuyên đề, sách giáo án của Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset. There is a set of 388,277 detailed… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-II-flash2.imagevisual-question-answering10K<n<100K6 likes21 downloads8mo agoHugging Face295CD-AI /Viet-OCR-VQA-flash2gated Dataset Overview The dataset comprises over 137,000 images potentially containing Vietnamese 🇻🇳 textual content. It was curated using the Gemini 1.5 Flash model, currently Google model leading on the WildVision Arena Leaderboard for Visual Question Answering (VQA). Each image is accompanied by a detailed description and 5 self-generated questions and answers related to the textual content within the image. In total, there are more than 822,679 individual questions, encompassing… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-OCR-VQA-flash2.imagevisual-question-answering100K<n<1M8 likes17 downloads8mo agoHugging Face305CD-AI /Viet-Doc-VQA-flash2gated Dataset Overview The Document VAQ dataset was collected from 51,856 pages of Vietnamese 🇻🇳 textbooks( Sách Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset. There is a set of 310,952 detailed descriptions and query-based questions and answers generated by the Gemini 1.5 Flash model, currently… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-flash2.imagevisual-question-answering10K<n<100K4 likes13 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.