datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ScreenSpot
Dataset Card for ScreenSpot
GUI Grounding Benchmark: ScreenSpot.
Created researchers at Nanjing University and Shanghai AI Laboratory for evaluating large multimodal models (LMMs) on GUI grounding tasks on screens given a text-based instruction.
Dataset Details
Dataset Description
ScreenSpot is an evaluation benchmark for GUI grounding, comprising over 1200 instructions from iOS, Android, macOS, Windows and Web environments, along with annotated… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/ScreenSpot.AIME_2000_2026_Kimi_K3
AIME 2000–2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2000_2026_Kimi_K3.RICO-WidgetCaptioning
Dataset Card for RICO Widget Captioning
Widget Captioning is a dataset for providing captions for UI elements on mobile screens.
It uses the RICO image database.
Dataset Details
Dataset Sources
Repository:
google-research-datasets/widget-caption
RICO raw downloads
Paper:
Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements
Rico: A Mobile App Dataset for Building Data-Driven Design Applications… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/RICO-WidgetCaptioning.AIME_1983_2026_Kimi_K3
AIME 1983–2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_1983_2026_Kimi_K3.AIME_2026_Kimi_K3
AIME 2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.
AIME… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2026_Kimi_K3.AutoMoT-PDM-Lite-BEV-Encoder-Indexes
AutoMoT PDM-Lite BEV Encoder Indexes
This dataset provides the prepared PDM-Lite JSONL indexes for AutoMoT training.
Files
pdm_lite_2hz_2tp_train_bev_encoder.jsonl
pdm_lite_2hz_2tp_val_bev_encoder.jsonl
Each row contains four historical front-camera paths in image, the current
front-camera path in front, trajectory and route supervision, future-speed
supervision, and a reference to the precomputed current-frame BEV feature:
bev_encoder_feature… See the full description on the dataset page: https://huggingface.co/datasets/HqH1111/AutoMoT-PDM-Lite-BEV-Encoder-Indexes.AIME_2025_Kimi_K3
AIME 2025 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.
AIME… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2025_Kimi_K3.
