datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ceval-examC-Eval is a comprehensive Chinese evaluation suite for foundation models. It consists of 13948 multi-choice questions spanning 52 diverse disciplines and four difficulty levels. Please visit our website and GitHub or check our paper for more details.
Each subject consists of three splits: dev, val, and test. The dev set per subject consists of five exemplars with explanations for few-shot evaluation. The val set is intended to be used for hyperparameter tuning. And the test set is for model… See the full description on the dataset page: https://huggingface.co/datasets/ceval/ceval-exam.exams
Dataset Card for [Dataset Name]
Dataset Summary
EXAMS is a benchmark dataset for multilingual and cross-lingual question answering from high school examinations. It consists of more than 24,000 high-quality high school exam questions in 16 languages, covering 8 language families and 24 school subjects from Natural Sciences and Social Sciences, among others.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The languages in the… See the full description on the dataset page: https://huggingface.co/datasets/mhardalov/exams.seggpt-example-datataiwan-examsMachine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), and assets/ (figures).
Usage
Benchmark any model against an exam:
a2b download… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-exams.doc-formats-parquet-1agents-last-exam
Agents Last Exam — Task Card Metadata (v1.1)
A metadata-only release (v1.1) of 152 tasks from the Agents Last Exam (ALE)
benchmark for evaluating computer-use agents on long-horizon professional work.
The Agents Last Exam dataset family
ALE is published as three companion HuggingFace datasets:
Dataset
Contents
Access
Task Card Metadata
One row per task: titles, prompts, taxonomy, input-file descriptors
Open
Task Input Data
The input/ files each task… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam.ceval-exam-zhtw
Dataset Card for "ceval-exam-zhtw"
C-Eval 是一個針對基礎模型的綜合中文評估套件。它由 13,948 道多項選擇題組成,涵蓋 52 個不同的學科和四個難度級別。原始網站和 GitHub 或查看論文以了解更多詳細資訊。
C-Eval 主要的數據都是使用簡體中文來撰寫并且用來評測簡體中文的 LLM 的效能來設計的,本數據集使用 OpenCC 來進行簡繁的中文轉換,主要目的方便繁中 LLM 的開發與驗測。
下載
使用 Hugging Face datasets 直接載入資料集:
from datasets import load_dataset
dataset=load_dataset(r"erhwenkuo/ceval-exam-zhtw",name="computer_network")
print(dataset['val'][0])
# {'id': 0, 'question':… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/ceval-exam-zhtw.Arabic_EXAMShumanitys-second-last-exam
Humanity's Second Last Exam
Benchmark design, curation and release maintenance: Shashank Agnihotri.
Original questions retain their recorded authorship and source attribution.
This owner-reviewed retained release contains 365 target questions, 730
context examples, and 365 ordered target/A/B links: 1,095 question rows.
The owner review concluded on 16 September 2026. This is an owner-reviewed
release after suspected-AI-content exclusions, not a software-certified
guarantee of… See the full description on the dataset page: https://huggingface.co/datasets/shashankskagnihotri/humanitys-second-last-exam.docvqa_1200_examplesoab_examsmultitask_german_examples_32kashraq-esc50-1-dog-example
Dataset Card for "ashraq-esc50-1-dog-example"
More Information needed
math_examplesEXAMS-V
EXAMS-V: ImageCLEF 2025 – Multimodal Reasoning
Dimitar Iliyanov Dimitrov, Hee Ming Shan, Zhuohan Xie, Rocktim Jyoti Das , Momina Ahsan, Sarfraz Ahmad, Nikolay Paev, Ali Mekky, Omar El Herraoui, Rania Hossam, Nurdaulet Mukhituly, Akhmed Sakip, Ivan Koychev, Preslav Nakov
INTRODUCTION
EXAMS-V is a multilingual, multimodal dataset created to evaluate and benchmark the visual reasoning abilities of AI systems, especially Vision-Language Models (VLMs). The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/EXAMS-V.ceval-examC-Eval is a comprehensive Chinese evaluation suite for foundation models. It consists of 13948 multi-choice questions spanning 52 diverse disciplines and four difficulty levels. Please visit our website and GitHub or check our paper for more details.
Each subject consists of three splits: dev, val, and test. The dev set per subject consists of five exemplars with explanations for few-shot evaluation. The val set is intended to be used for hyperparameter tuning. And the test set is for model… See the full description on the dataset page: https://huggingface.co/datasets/duzhaowu/ceval-exam.example_dataset
example_dataset
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
example-vlm-sft-dataset
Sample VLM Reference Dataset
Dataset Description
This is a reference dataset demonstrating the proper VLM SFT (Supervised Fine-Tuning) format for vision-language model training. It contains 10 minimal conversational examples that show the exact structure and formatting required for VLM training pipelines.
⚠️ Important: This dataset is for format reference only - not intended for actual model training. Use this as a template to understand the required data structure for… See the full description on the dataset page: https://huggingface.co/datasets/alay2shah/example-vlm-sft-dataset.external_data_test_exampleceval-examAdd Column 'choices' to the original dataset.
Citation
If you use C-Eval benchmark or the code in your research, please cite their paper:
@article{huang2023ceval,
title={C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models},
author={Huang, Yuzhen and Bai, Yuzhuo and Zhu, Zhihao and Zhang, Junlei and Zhang, Jinghan and Su, Tangjun and Liu, Junteng and Lv, Chuancheng and Zhang, Yikai and Lei, Jiayi and Fu, Yao and Sun, Maosong and He, Junxian}… See the full description on the dataset page: https://huggingface.co/datasets/zacharyxxxxcr/ceval-exam.unsplash-examplesTaiwanese-Minnan-Example-Sentences
Taiwanese Minnan Example Sentences
The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems.
Dataset Features
Source: Ministry of Education, Taiwan (Sutian Resource Center)
Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.Arabic_EXAMS-Redux
Arabic_EXAMS-Redux
A corrected and text-repaired version of OALL/Arabic_EXAMS, the Arabic subset of the EXAMS multilingual high-school examinations benchmark.
What was fixed
Repaired corrupted Arabic text. The upstream benchmark contains widespread PDF-extraction damage to question stems and answer choices: split diacritics, fragmented words, and non-Arabic glyphs replacing standard characters. We restored these to readable Modern Standard Arabic.
Corrected the answer… See the full description on the dataset page: https://huggingface.co/datasets/inceptlabs/Arabic_EXAMS-Redux.entrance-exam-dataset
TO DO Checklist:
Clean Data
Remove duplicates
Handle missing values
Standardize data formats
aya-global-exams-catalanCatalan exams for the Aya Global Exams.
Original data and file available here: link
Github Repo: link
example-documentstaiwan-exams-resultsBenchmark results produced by any-to-bench. One subset here is one taker
configuration — a single model at a single reasoning effort — sat against the
exams in another dataset repo. Every row names the exam repo and subset it was
earned against, so results from several corpora, and from several people, can
live side by side.
results-index.json — the catalog: one headline row per configuration
results-<entry>/entry.json — that configuration's per-paper scores
results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-exams-results.taiwan-professional-exams-115-2Machine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), and assets/ (figures).
Usage
Benchmark any model against an exam:
a2b download… See the full description on the dataset page: https://huggingface.co/datasets/skyhong2002/taiwan-professional-exams-115-2.example_dataset_bboxes
example_dataset
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
ecommerce_last_exam
E-Commerce Last Exam
A benchmark for evaluating LLM agents on 120 real-world travel planning and e-commerce tool-use tasks. Each task runs in an isolated Docker container with domain-specific CLI tools and SQLite databases. Agents must search, analyze, and produce structured recommendations.
Repository: alibaba-flyai/ecommerce_last_exam
Evaluation CLI: flyai-bench (pip install flyai-bench)
Leaderboard: FlyaiLab/ecommerce_last_exam_leaderboard
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/FlyaiLab/ecommerce_last_exam.
