reading
Datasets
All datasets matching “reading”foia-reading-room-documents
Foia Reading Room Documents
Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act.
Every document here was published by a US federal agency and is a work of the
United States government. Nothing has been altered: files are byte-identical to
what the agency posted, and the checksum in metadata.parquet is of the
original bytes.
Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.meta-active-readingvisual_qa_multipanel
Do you “see" what I “see"? A Multi-panel Visual Question and Answer Dataset for Large Language Model Chart Analysis
Publication: Accepted to TPDL 2026 (linked on publication)
GitHub Repo: https://github.com/ReadingTimeMachine/LLM_VQA_MultiPanel
This is a multi-panel figure dataset for visual question and answer (VQA) to test large language/multimodal models (LMMs).
Data contains synthetically generated multi-panel figures with histogram, scatter… See the full description on the dataset page: https://huggingface.co/datasets/ReadingTimeMachine/visual_qa_multipanel.sat-reading
Dataset Card for "sat-reading"
This dataset contains the passages and questions from the Reading part of ten publicly available SAT Practice Tests.
For more information see the blog post Language Models vs. The SAT Reading Test.
For each question, the reading passage from the section it is contained in is prefixed.
Then, the question is prompted with Question #:, followed by the four possible answers.
Each entry ends with Answer:.
Questions which reference a diagram, chart, table… See the full description on the dataset page: https://huggingface.co/datasets/emozilla/sat-reading.SAT_Writting_Reading_Assessment_Question_Bank
Dataset Card for SAT Reading and Writing Dataset
This dataset card aims to be a base template for the SAT Reading and Writing Dataset, optimized for use with Hugging Face's datasets library.
Dataset Details
Dataset Description
This dataset contains SAT Reading and Writing assessment questions sourced from the College Board's SAT Suite Question Bank, intended for use in training and evaluating Language Models like LLMs.
Curated by: College Board
License:… See the full description on the dataset page: https://huggingface.co/datasets/betterMateusz/SAT_Writting_Reading_Assessment_Question_Bank.rtm-sgt-ocr-v1
Data Introduction
Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic Data from the ar𝜒iv for OCR Post Correction of Historic Scientific Articles".
Synthetic ground truth (SGT) sentences have been mined from the ar𝜒iv Bulk Downloads source documents,
and Optical Character Recognition (OCR)
sentences have been generated with the Tesseract OCR engine on the PDF pages generated from compiled source documents.… See the full description on the dataset page: https://huggingface.co/datasets/ReadingTimeMachine/rtm-sgt-ocr-v1.
