datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ceval-examC-Eval is a comprehensive Chinese evaluation suite for foundation models. It consists of 13948 multi-choice questions spanning 52 diverse disciplines and four difficulty levels. Please visit our website and GitHub or check our paper for more details.
Each subject consists of three splits: dev, val, and test. The dev set per subject consists of five exemplars with explanations for few-shot evaluation. The val set is intended to be used for hyperparameter tuning. And the test set is for model… See the full description on the dataset page: https://huggingface.co/datasets/ceval/ceval-exam.agents-last-exam-data
Agents Last Exam — Task Input Data
Input files (the materials each task hands to the agent at run start) for the
Agents Last Exam (ALE) benchmark. Browsable per-task directory layout.
The Agents Last Exam dataset family
ALE is published as three companion HuggingFace datasets:
Dataset
Contents
Access
Task Card Metadata
One row per task: titles, prompts, taxonomy, input-file descriptors
Open
Task Input Data
The input/ files each task hands the agent at… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data.ale-images-qcow2
ALE QEMU runner image
agentslastexam/ale-qemu is the container-side runtime used by the ALE qemu
provider. It packages QEMU, KVM integration, NAT networking, noVNC, and process
supervision. The Ubuntu or Windows guest is supplied separately as
/storage/data.qcow2.
Docker is the container runtime. Dockur is the upstream QEMU-in-Docker project
whose startup and networking stack this image inherits. ALE adds a stable
runner contract around that upstream image.
The image is based on… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/ale-images-qcow2.exams
Dataset Card for [Dataset Name]
Dataset Summary
EXAMS is a benchmark dataset for multilingual and cross-lingual question answering from high school examinations. It consists of more than 24,000 high-quality high school exam questions in 16 languages, covering 8 language families and 24 school subjects from Natural Sciences and Social Sciences, among others.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The languages in the… See the full description on the dataset page: https://huggingface.co/datasets/mhardalov/exams.ceval-examC-Eval is a comprehensive Chinese evaluation suite for foundation models. It consists of 13948 multi-choice questions spanning 52 diverse disciplines and four difficulty levels.example-imagesRSI-Exam
Benchmarking Recursive Self-Improvement through Executable Research
Can frontier AI agents achieve Recursive Self-Improvement by turning a weak method into one that performs better on hidden data?
Overview
RSI-Exam is a benchmark of 88 executable research tasks for evaluating whether frontier
AI agents can achieve recursive self-improvement. Instead of measuring one-shot problem solving,
each task hands the agent a working but weak method and lets… See the full description on the dataset page: https://huggingface.co/datasets/RSI-Exam/RSI-Exam.agents-last-exam-reference
Agents Last Exam — Reference (Ground-Truth) Data
⚠️ Gated dataset. This repo contains the ground-truth / reference outputs
used to score the Agents Last Exam (ALE) benchmark. Access requires login,
agreement to the terms on the access-request form, and manual approval.
Note (06/16/26): This repository was accidentally deleted and has been recreated. The
previous list of approved requesters could not be restored, so even if you
were granted access before, you will need to… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-reference.Fire3D_examples
Fire3D Website Examples
Paper |
Project page |
Code
Web-optimized interactive examples and qualitative comparison media for the
Fire3D project page.
The website/v1 release contains:
eight interactive iTHOR and Imaginarium scenes;
RGB and instance-colored point clouds with predicted 3D oriented boxes;
separately loadable foreground and background GLBs;
native-panel paper results and supplementary baseline comparisons; and
application videos.
These assets are presentation… See the full description on the dataset page: https://huggingface.co/datasets/hongchi/Fire3D_examples.code-world-model-inference-examples-40
Inference examples
This directory contains 40 numbered, independent inference examples.
Every example uses only its public number; source case names and internal paths
are intentionally omitted.
Each numbered directory contains:
first_frame.png: exact 1536x864 generated RGB first frame used by inference.
prompts/*.txt: the exact rolling long-inference prompts used for the result.
condition/*.npz: ordered lossless condition chunks.
metadata.json: frame count, FPS, prompt windows… See the full description on the dataset page: https://huggingface.co/datasets/NTU-yiwen/code-world-model-inference-examples-40.aya-mm-exams-spanish-medicalMedical Spanish Exams for the Multimodal Aya Exams Projects.
Questions available in file: data.json
Images stored in: /images
Original data and file available here: link
dog-examplelisbet-examplesseggpt-example-dataLyra-Testing-ExampleExamplessounio-code-examples
Sounio Curated Code Examples
Curated compile-clean .sio examples for training and evaluating code models on
Sounio, a self-hosted systems and scientific programming language for epistemic
computing, uncertainty propagation, and algebraic effects.
This directory is the Cx-1 expansion lane for
chiuratto-AIgourakis/sounio-code-examples.
Current batch
Examples: 5,000
Metadata files: 5,000
Compiler gate: bin/souc check pass rate 5,000/5,000
Utility layer: 5,000… See the full description on the dataset page: https://huggingface.co/datasets/chiuratto-AIgourakis/sounio-code-examples.taiwan-examsMachine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), and assets/ (figures).
Usage
Benchmark any model against an exam:
a2b download… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-exams.doc-formats-jsonl-1
[doc] formats - jsonl - 1
This dataset contains one jsonl file at the root.
scenesmith-example-scenes
SceneSmith Example Scenes
Project Page | Paper | Code
Example scenes generated by SceneSmith, a hierarchical agentic framework for constructing simulation-ready indoor environments from natural language prompts.
This dataset contains all scenes from the SceneSmith method (and its ablations) used in the paper evaluations. Each scene is a complete simulation-ready environment with 3D assets (including VLM-estimated physical properties), collision meshes, floor plans, and scene… See the full description on the dataset page: https://huggingface.co/datasets/nepfaff/scenesmith-example-scenes.hypersim-exampleschat_formatted_examplesark_exampleThis dataset contains 11 Field of Views (FOVs), each with 22 channels.thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit Viriyayudhakorn… See the full description on the dataset page: https://huggingface.co/datasets/matichon/thai-onet-m6-exam.thai_exam
Dataset Card for Thai_Exam
ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows:
ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.agents-last-exam-data-archive
Agents Last Exam — Task Data Archive (input + reference)
⚠️ Gated dataset. This repo packages each task's input, software, and
reference (ground-truth) data into a single archive (ale-tasks-data.tar.gz)
for convenient one-shot download — in particular for running ALE locally with
the local Docker provider,
which fetches it and mounts each task's data at run time. Because it includes
the reference outputs used to score runs, access requires login, agreement to
the terms on the… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data-archive.doc-formats-parquet-1doc-formats-csv-1
[doc] formats - csv - 1
This dataset contains one csv file at the root:
data.csv
kind,sound
dog,woof
cat,meow
pokemon,pika
human,hello
The YAML section of the README does not contain anything related to loading the data (only the size category metadata):
---
size_categories:
- n<1K
---
agents-last-exam
Agents Last Exam — Task Card Metadata (v1.1)
A metadata-only release (v1.1) of 152 tasks from the Agents Last Exam (ALE)
benchmark for evaluating computer-use agents on long-horizon professional work.
The Agents Last Exam dataset family
ALE is published as three companion HuggingFace datasets:
Dataset
Contents
Access
Task Card Metadata
One row per task: titles, prompts, taxonomy, input-file descriptors
Open
Task Input Data
The input/ files each task… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam.indian-exams-rawdata
Indian Competitive Exams Raw Dataset (GATE, JEE Main & JEE Advanced)
This dataset contains raw PDF question papers, official answer keys, and extracted/parsed question structured data for major Indian national-level competitive engineering examinations: GATE, JEE Advanced, and JEE Main, spanning multiple years (2007–2025).
Data Sources & Attribution
The data in this repository was scraped and compiled from official conducting authority portals and public… See the full description on the dataset page: https://huggingface.co/datasets/Abhay557/indian-exams-rawdata.
