datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CiQi-VQA
CiQi-Agent
Github | Model | Dataset | Paper
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
Accepted to ECCV 2026
🎯 Overview
CiQi-Agent has been accepted to ECCV 2026.
We present CiQi-Agent, a domain-specific multimodal agent for antique Chinese porcelain connoisseurship. The project is designed to combine fine-grained visual perception, tool-augmented reasoning, and cultural-heritage knowledge… See the full description on the dataset page: https://huggingface.co/datasets/SII-Monument-Valley/CiQi-VQA.CircuitSense
CircuitSense
This dataset is a comprehensive multimodal circuit question-answering benchmark designed to evaluate visual reasoning and problem-solving capabilities across three main domains: Perception, Analysis, and Design. The dataset contains structured question-answer pairs with accompanying visual content, targeting different engineering cognitive levels and reasoning tasks.
Dataset Structure
The dataset is organized into three primary folders, each containing… See the full description on the dataset page: https://huggingface.co/datasets/armanakbari4/CircuitSense.mmlu-cs
Czech MMLU
This is a Czech translation of the original MMLU dataset, created using the WMT 21 En-X model.
The 'auxiliary_train' subset is not included.
The translation was completed for use within the Czech-Bench evaluation framework.
The script used for translation can be reviewed here.
Citation
Original dataset:
@article{hendryckstest2021,
title={Measuring Massive Multitask Language Understanding},
author={Dan Hendrycks and Collin Burns and Steven Basart and… See the full description on the dataset page: https://huggingface.co/datasets/CIIRC-NLP/mmlu-cs.MMSearch-Plus
MMSearch-Plus✨: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
Official repository for the paper "MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents".
🌟 For more details, please refer to the project page with examples: https://mmsearch-plus.github.io/.
[🌐 Webpage] [📖 Paper] [🤗 Huggingface Dataset] [🏆 Leaderboard]
💥 News
[2025.09.26] 🔥 We update the arXiv paperand release all MMSearch-Plus data samples in… See the full description on the dataset page: https://huggingface.co/datasets/Cie1/MMSearch-Plus.Dr-CiK
Dr-CiK: A Testbed for Foresight-Driven Agents
Dr-CiK is a benchmark for evaluating whether agents can retrieve
forecasting-relevant context from a noisy document corpus, filter out
distractors, distill the retrieved context into forecast-useful evidence, and
produce forecasts grounded in that evidence.
Real-world time-series forecasting often depends not only on historical
observations but also on external context that must be actively discovered
from heterogeneous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/Dr-CiK.LimAgents_limitation_data_scientific_papers_with_cited_papers
LimAgents Data
This dataset contains scientific paper metadata and extracted limitation information prepared for use with LLM Agents.The data comes from NeurIPS 2021–2022 papers and related OpenReview reviews, enriched with Cited in and Cited by information.
Dataset Structure
The repository contains two main directories:
1. NeurIPS_21_22_Lim_OPR_with_cited_in_by_papers
This directory includes one JSON file per paper. Each file contains:
title: Original paper… See the full description on the dataset page: https://huggingface.co/datasets/iaadlab/LimAgents_limitation_data_scientific_papers_with_cited_papers.civil-code-phil
Civilex — Philippine Legal RAG & SFT Dataset
Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline.
Contents: 11k+ Supreme Court jurisprudence cases spanning 1949–2025, and 2,270 articles from the Civil Code (Republic Act No. 386).
Dataset structure
.
├── README.md
├──… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.CitationGround-1M
CitationGround-1M (Platinum)
Developer/Publisher: Within US AIVersion: 0.1.0 (sample pack)Created: 2025-12-30T16:53:41Z
What this dataset is
CitationGround-1M is a citation-locked grounded QA/RAG dataset:
Answer using only the provided contexts
Provide span-level citations (doc_id + offsets)
Includes answerable=false hard negatives for abstention behavior
Features / schema (JSONL)
example_id (string)
question (string)
contexts (list of docs)… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/CitationGround-1M.CID-Dataset
CID Dataset
Training and evaluation data for Continuous Interaction Diffusion (CID).
Paper: Continuous Interaction Diffusion: A Diffusion-Native Architecture for Asynchronous Tool-Augmented Reasoning — Yuhang Cao, Yanzhou Mu, Chunrong Fang, and Zhenyu Chen, arXiv:2608.10438 (2026).
CID-Dataset is a structured corpus for asynchronous tool-augmented reasoning. Its supervision is organized as semantic tasks, reviewed TeacherPlans, randomized runtime trajectories, and… See the full description on the dataset page: https://huggingface.co/datasets/fwerkor/CID-Dataset.cibench-experiments
CIBench Experiments
Reproducibility packages for CIBench — the stateless, replayable benchmark engine for the 1M–10M token long-context era.
If a benchmark result cannot be replayed from its manifest alone, it did not happen.
Every sub-directory in this dataset is a self-contained experiment package: per-run manifests, content-addressed canonical JSON, ResultRecord with full scoring + signed provenance, per-item OpenTelemetry gen_ai_* call metrics, retrieved evidence, a… See the full description on the dataset page: https://huggingface.co/datasets/publicus-ai/cibench-experiments.CII-Bench
CII-Bench
🌐 Homepage | 🤗 Dataset | GitHub | 🤗 Paper | 📖 arXiv
Introduction
CII-Bench comprises 698 Chinese images, each accompanied by 1 to 3 multiple-choice questions, totaling 800 questions. CII-Bench encompasses images from six distinct domains: Life, Art, Society, Environment, Politics, and Chinese Traditional Culture. It also features a diverse array of image types, including Illustrations, Memes, Posters, Multi-panel Comics, Single-panel Comics, and… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CII-Bench.wildtrace
WildTrace strict481
WildTrace is a source-internal long-context multi-hop reasoning benchmark
built from natural evidence trails. Unlike reverse-synthetic QA, its tasks are
mined in situ from long source documents before questions are written. The
strict481 release contains 481 locked tasks over 214 public long-form sources,
with full-document, evidence-withheld evaluation. The model under test receives
only the source document and the public question; evidence spans, clue… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/wildtrace.L-CiteEval
L-CITEEVAL: DO LONG-CONTEXT MODELS TRULY LEVERAGE CONTEXT FOR RESPONDING?
Paper Github Zhihu
Benchmark Quickview
L-CiteEval is a multi-task long-context understanding with citation benchmark, covering 5 task categories, including single-document question answering, multi-document question answering, summarization, dialogue understanding, and synthetic tasks, encompassing 11 different long-context tasks. The context lengths for these tasks range from 8K to 48K.… See the full description on the dataset page: https://huggingface.co/datasets/Jonaszky123/L-CiteEval.CiteVQA
CiteVQA
English | 简体中文
CiteVQA is a document visual question answering benchmark for faithful evidence attribution. Unlike conventional DocVQA datasets that only score the final answer, CiteVQA requires a model to answer a question with evidence grounded in the source document at the element level. The benchmark is designed to evaluate whether a system can not only answer correctly, but also cite the right supporting region in long, real-world PDFs.
The dataset contains 1,897… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/CiteVQA.CityCube-Benchgreek_civics_qa
Dataset Card for Greek Civics QA
The Greek Civics QA dataset is a set of 407 question and answer pairs related to Greek highschool courses in civics (Κοινωνική και Πολιτική Αγωγή). The dataset was created by Nikoletta Tsoukala during her 2023 internship at the Institute for Language and Speech Processing/Athena RC, supervised by ILSP Research Associate Vassilis Papavassileiou. The dataset creation process involved mining questions and answers from two civics textbooks used in the… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/greek_civics_qa.SpatiaLab
SpatiaLab: Can Vision–Language Models Perform Spatial Reasoning in the Wild?
Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman, Mahfuz Ahmed Anik, Munem Shahriar, Mohsin Mahmud Topu, Sadia Tasnim Meem, Rahatun Nesa Priti, Sabrina Afroz Mitu, Md. Iqramul Hoque, Shahriyar Zaman Ridoy, Mohammed Eunus Ali, Majd Hawasly, Mohammad Raza, Md Rizwan Parvez
Computational Intelligence and Operations Laboratory (CIOL) • Shahjalal University of Science and Technology (SUST) • Monash… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/SpatiaLab.CIMD
Chinese Instruction Multimodal Data (CIMD)
The dataset contains one million Chinese image-text pairs in total, including detailed image captioning and visual question answering.
Generation Pipeline
Image source
We randomly sample images from two opensource datasets Wanjuan and Wukong
Detailed caption generation
We use Gemini Pro Vision API to generate a detailed description for each image.
Question-answer pairs generation
Based on the generated caption, we use Gemini… See the full description on the dataset page: https://huggingface.co/datasets/jingzi/CIMD.bentoThis dataset is based on MMLU, FLAN, Big Bench Hard and AgiEval English.
The non-"reduced" benchmark is the original benchmark, except for FLAN, which is a sampled version.
The "reduced" benchmark only contains a few representative tasks in the original ones, such that the performance on the "reduced" benchmark can serve as an approximation to the performance on the original ones.
m_lamamLAMA: a multilingual version of the LAMA benchmark (T-REx and GoogleRE) covering 53 languages.MLLM-CITBench
MLLM-CITBench Multimodal Task Benchmarking Dataset
This dataset contains 7 tasks:
OCR: Optical Character Recognition task
art: Art - related task
fomc: Financial and Monetary Policy - related task
math: Mathematical problem - solving task
medical: Medical - related task
numglue: Numerical reasoning task
science: Scientific problem - solving task
Each task has independent training and test splits. The image data is stored in the dataset in the form of full file.
Usage… See the full description on the dataset page: https://huggingface.co/datasets/yueluoshuangtian/MLLM-CITBench.photonic-integrated-circuit-yield
🏭 Photonic Integrated Circuit Yield Dataset
📊 125,000 synthetic (yield query, yield reasoning response) pairs covering process variation, defect density, lithography, and metrology challenges in CMOS-compatible photonic integrated circuit (PIC) manufacturing.
⚠️ Disclaimer: All entries are synthetically generated. Yield figures are computed from textbook models over sampled inputs, and citations are placeholders styled after technical sources; none reference a real… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/photonic-integrated-circuit-yield.k-beauty-ai-citation-dataset
K-Beauty AI Citation Dataset
Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research.
Canonical source: https://kbeautyanswers.com/dataset/
License: CC BY 4.0
Maintainer: K-Beauty Answers (site)
Initial release: 2026-05-23
What's in it
128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.Reverse-circuit-discoveryCipherBank
CipherBank Benchmark
Benchmark description
CipherBank, a comprehensive benchmark designed to evaluate the reasoning capabilities of LLMs in cryptographic decryption tasks.
CipherBank comprises 2,358 meticulously crafted problems, covering 262 unique plaintexts across 5 domains and 14 subdomains, with a focus on privacy-sensitive and real-world scenarios that necessitate encryption. From a cryptographic perspective, CipherBank incorporates 3 major categories of encryption… See the full description on the dataset page: https://huggingface.co/datasets/yu0226/CipherBank.BuySideFinBench
BuySideFinBench
A bilingual benchmark for evaluating large language models on buy-side equity research and valuation tasks.
Overview
BuySideFinBench targets the analytical reasoning that distinguishes a buy-side equity research analyst from a generalist financial reader. Unlike most finance LLM benchmarks that focus on sell-side / news-driven tasks (sentiment, summarization, headline interpretation) or surface-level multiple-choice knowledge, BuySideFinBench… See the full description on the dataset page: https://huggingface.co/datasets/cindy90/BuySideFinBench.citizen_nlu
Dataset Card for citizen_nlu
Dataset Summary
NeuralSpace strives to provide AutoNLP text and speech services, especially for low-resource languages. One of the major services provided by NeuralSpace on its platform is the “Language Understanding” service, where you can build, train and deploy your NLU model to recognize intents and entities with minimal code and just a few clicks.
The initiative of this challenge is created with the purpose of sparkling AI applications to… See the full description on the dataset page: https://huggingface.co/datasets/neuralspace/citizen_nlu.Original-circuit-discoveryTeachArena
TeachArena
TeachArena is a benchmark for evaluating AI tutoring agents across the full teaching
decision chain — from moment-to-moment tutoring dialogue, to pedagogical judgment on
packaged evidence, to multi-step teaching workflows grounded in a learning-management
system. It contains 354 tasks organized into three stages, a mock LMS environment
database, the agent policy documents, and the full scoring logic.
Why three stages
A capable teaching agent must both… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/TeachArena.arc-cs
Czech AI2 Reasoning Challenge
This is a Czech translation of the original ARC dataset, created using the WMT 21 En-X model.
The translation was completed for use within the Czech-Bench evaluation framework.
The script used for translation can be reviewed here.
Citation
Original dataset:
@article{allenai:arc,
author = {Peter Clark and Isaac Cowhey and Oren Etzioni and Tushar Khot and
Ashish Sabharwal and Carissa Schoenick and Oyvind… See the full description on the dataset page: https://huggingface.co/datasets/CIIRC-NLP/arc-cs.
