datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pubmed-ocr
PubMed-OCR: PMC Open Access OCR Annotations
PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs. Each page is rendered to an image and annotated with Google Cloud Vision OCR, released in a compact JSON schema with word-, line-, and paragraph-level bounding boxes.
Scale (release):
209.5K articles
~1.5M pages
~1.3B words (OCR tokens)
This dataset is intended to support layout-aware modeling, coordinate-grounded QA, and… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/pubmed-ocr.ScreenSpot
Dataset Card for ScreenSpot
GUI Grounding Benchmark: ScreenSpot.
Created researchers at Nanjing University and Shanghai AI Laboratory for evaluating large multimodal models (LMMs) on GUI grounding tasks on screens given a text-based instruction.
Dataset Details
Dataset Description
ScreenSpot is an evaluation benchmark for GUI grounding, comprising over 1200 instructions from iOS, Android, macOS, Windows and Web environments, along with annotated… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/ScreenSpot.AIME_2000_2026_Kimi_K3
AIME 2000–2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2000_2026_Kimi_K3.FinQA
FinQA
A full-fidelity repackaging of the FinQA dataset (Chen et al., EMNLP 2021) for numerical reasoning over financial tables.
FinQA contains questions over earnings reports from S&P 500 companies (1999–2019), sourced from the FinTabNet dataset. Each example pairs a financial table and surrounding text with a question, a human-readable answer, and a structured reasoning program that specifies the arithmetic operations needed to derive the answer.
Why this version… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/FinQA.RICO-ScreenQA
Dataset Card for ScreenQA
Question answering on RICO screens: google-research-datasets/screen_qa.
Citation
BibTeX:
@misc{hsiao2024screenqa,
title={ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots},
author={Yu-Chung Hsiao and Fedir Zubach and Maria Wang and Jindong Chen},
year={2024},
eprint={2209.08199},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
RICO-WidgetCaptioning
Dataset Card for RICO Widget Captioning
Widget Captioning is a dataset for providing captions for UI elements on mobile screens.
It uses the RICO image database.
Dataset Details
Dataset Sources
Repository:
google-research-datasets/widget-caption
RICO raw downloads
Paper:
Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements
Rico: A Mobile App Dataset for Building Data-Driven Design Applications… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/RICO-WidgetCaptioning.RICO-Screen2Words
Dataset Card for Screen2Words
Screen2Words is a dataset providing screen summaries (i.e., image captions for mobile screens).
It uses the RICO image database.
Dataset Details
Dataset Sources
Repository:
google-research-datasets/screen2words
RICO raw downloads
Paper:
Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning
Rico: A Mobile App Dataset for Building Data-Driven Design Applications
Uses
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/RICO-Screen2Words.RICO-ScreenQA-Short
Dataset Card for ScreenQA-Short
Question answering on RICO screens: google-research-datasets/screen_qa.
These are the set of answers that have been machine generated and are designed to be short response.
Citation
BibTeX:
@misc{baechler2024screenai,
title={ScreenAI: A Vision-Language Model for UI and Infographics Understanding},
author={Gilles Baechler and Srinivas Sunkara and Maria Wang and Fedir Zubach and Hassan Mansoor and Vincent Etter and Victor… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/RICO-ScreenQA-Short.RICO-SCA
Dataset Card for RICO SCA (SeeClick cache)
This is the SeeClick cache of a syntehtically generated dataset following RICO SCA's generation procedure.
It consists of approximately 170k captions across 70k widgets and 18k screens.
Dataset Details
Dataset Description
This is a widget captioning (referring expression comprehension/generation) dataset.
Curated by: Google Research, Nanjing University
Language(s) (NLP): en
License: apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/RICO-SCA.websrc
Dataset Card for WebSRC v1.0
WebSRC v1.0 is a dataset for reading comprehension on structural web pages.
The task is to answer questions about web pages, which requires a system to have a comprehensive understanding of the spatial structure and logical structure.
WebSRC consists of 6.4K web pages and 400K question-answer pairs about web pages.
This cached copy of the dataset is focused on Q&A using the web screenshots (HTML and other metadata are omitted).
Questions in WebSRC… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/websrc.ENTRANT
ENTRANT
A HuggingFace mirror of the ENTRANT dataset
(Zenodo record 10667088, CC-BY-4.0): 6.7 million structured financial tables
extracted from ~330,000 SEC EDGAR filings spanning 10 filing types.
Original paper: Gialitsis et al., Scientific Data 2024,
10.1038/s41597-024-03605-5.
License
CC-BY-4.0 (dataset annotations and structural metadata).
The underlying source documents are SEC EDGAR filings -- publicly filed
corporate documents. The factual content of… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/ENTRANT.TABMEpp
Dataset Card for TABME++
The TABME dataset is a synthetic collection of business document folders generated from the Truth Tobacco Industry Documents archive, with preprocessing and OCR results included, designed to simulate real-world digitization tasks.
TABME++ extends TABME by enriching it with commercial-quality OCR (Microsoft OCR).
Dataset Details
Dataset Description
The TABME dataset is a synthetic collection created to simulate the… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/TABMEpp.beverages_catalogue_ru_beirThis is a copy of https://huggingface.co/datasets/jinaai/beverages_catalogue_ru reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/beverages_catalogue_ru_beir.MultiHiertt
MultiHiertt
A repackaging of the MultiHiertt dataset (Zhao et al., ACL 2022) for numerical reasoning over documents containing multiple hierarchical financial tables.
MultiHiertt is built on FinTabNet (CDLA-Permissive-1.0), extracting 4,791 multi-page documents from S&P 500 annual reports (1999-2019). Each document contains 2-6 hierarchical HTML tables and surrounding text. Questions require reasoning across multiple tables and/or text passages, with answers derived via… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/MultiHiertt.RICO-ScreenAnnotation
Dataset Card for RICO Screen Annotations
This is a standardization of Google's Screen Annotation dataset on a subset of RICO screens, as described in their ScreenAI paper.
It retains location tokens as integers.
Dataset Details
Dataset Description
This is an image-to-text annotation format first proscribed in Google's ScreenAI paper.
The idea is to standardize an expected text output that is reasonable for the model to follow,
and fuses together… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/RICO-ScreenAnnotation.websrc-test
Dataset Card for WebSRC Test Split
This is the test split (without ground truth) for WebSRC. See WebSRC for the full train and dev splits, with answers.
robocasa_20260430T030150Z_full_run_beverage_organization ---
pretty_name: RoboCasa Trajectories Single
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
---
# RoboCasa Trajectories Single
This dataset contains one row per RoboCasa trajectory / episode.
## Structure
Each row is one trajectory / episode.
Episode-level JSON is stored inline:
adapted_trajectory
original_trajectory
execution_metadata
Step-level data is stored in aligned sequence columns:… See the full description on the dataset page: https://huggingface.co/datasets/DorianAtSchool/robocasa_20260430T030150Z_full_run_beverage_organization.RICO-ScreenAnnotation-f
Dataset Card for RICO Screen Annotations
This is a standardization of Google's Screen Annotation dataset on a subset of RICO screens, as described in their ScreenAI paper.
Unlike the original, this version transforms integer-based bounding boxes into floating-point-based bounding boxes of 2 decimal precision.
Dataset Details
Dataset Description
This is an image-to-text annotation format first proscribed in Google's ScreenAI paper.
The idea is to… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/RICO-ScreenAnnotation-f.SciTSR-pd
SciTSR-PD
A public-domain subset of SciTSR, a large-scale table structure recognition dataset of scientific tables extracted from arXiv LaTeX source files.
This subset contains only tables whose source papers carry a CC0 or equivalent public domain dedication — no attribution required, no restrictions on commercial or derivative use.
Dataset Details
Split
Tables
Papers
train
89
—
test
19
—
total
108
52
Source paper licenses present: CC0… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/SciTSR-pd.AIME_1983_2026_Kimi_K3
AIME 1983–2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_1983_2026_Kimi_K3.CISOL
CISOL
A HuggingFace mirror of the CISOL dataset
(Zenodo record 10829550, CC-BY-4.0): German construction-industry steel ordering
lists annotated for table detection and table structure recognition.
Original paper: Tschirschwitz et al., WACV 2025 (arXiv:2501.15469).
License
CC-BY-4.0. The peer-reviewed paper (WACV 2025) states: "The data is licensed
under the Creative Commons Attribution 4.0 International license to support full
research use, subject to proper… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/CISOL.RICO-ScreenQA-Complex
Dataset Card for ScreenQA-Complex
Question answering on RICO screens: google-research-datasets/screen_qa.
These are the test-only complex questions.
Citation
BibTeX:
@misc{baechler2024screenai,
title={ScreenAI: A Vision-Language Model for UI and Infographics Understanding},
author={Gilles Baechler and Srinivas Sunkara and Maria Wang and Fedir Zubach and Hassan Mansoor and Vincent Etter and Victor Cărbune and Jason Lin and Jindong Chen and Abhanshu… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/RICO-ScreenQA-Complex.SciTSR-cc-by-nc-sa
SciTSR-CC-BY-NC-SA
A license-filtered subset of SciTSR, a large-scale table structure recognition dataset of scientific tables extracted from arXiv LaTeX source files.
This subset contains tables whose source papers are compatible with a CC-BY-NC-SA 4.0 open-weight model release — covering public domain, CC-BY, CC-BY-NC, and CC-BY-NC-SA licensed papers. The dataset itself is released under CC-BY-NC-SA 4.0.
Dataset Details
Split
Tables
Papers
train
697
—… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/SciTSR-cc-by-nc-sa.HiTab-StatCan-NSF
HiTab-StatCan-NSF
A commercially-permissive subset of Microsoft HiTab
(ACL 2022) containing only the Statistics Canada and NSF table sources.
The Wikipedia/ToTTo subset (CC-BY-SA-4.0, share-alike) is excluded.
License
NSF tables (~14% of rows): sourced from US National Science Foundation
reports. US federal government works are public domain under
17 U.S.C. § 105.
StatCan tables (~86% of rows): sourced from Statistics Canada reports
under the Statistics Canada… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/HiTab-StatCan-NSF.bevy_selectedus-food-beverage-plant-meatpacking-agriculture-layoffs-warn-act-notices-daily
US food plant, meatpacking and agriculture layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-22. 1,280 layoff and closure notices filed by
food and beverage manufacturers, meat and poultry packers, bakeries and snack plants, dairies and creameries, bottlers and brewers, seafood processors, farms, growers and packinghouses, and grain, feed and sugar mills with US state labor departments — 173,217 workers,
670 employers, 43 states, 1989–2026.
476 of… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-food-beverage-plant-meatpacking-agriculture-layoffs-warn-act-notices-daily.bevyengine_bevyAIME_2026_Kimi_K3
AIME 2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.
AIME… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2026_Kimi_K3.Amazon-beverage-reviews-with-ratingsAIME_2025_Kimi_K3
AIME 2025 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.
AIME… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2025_Kimi_K3.
