datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
meta-active-readingsat-reading
Dataset Card for "sat-reading"
This dataset contains the passages and questions from the Reading part of ten publicly available SAT Practice Tests.
For more information see the blog post Language Models vs. The SAT Reading Test.
For each question, the reading passage from the section it is contained in is prefixed.
Then, the question is prompted with Question #:, followed by the four possible answers.
Each entry ends with Answer:.
Questions which reference a diagram, chart, table… See the full description on the dataset page: https://huggingface.co/datasets/emozilla/sat-reading.SAT_Writting_Reading_Assessment_Question_Bank
Dataset Card for SAT Reading and Writing Dataset
This dataset card aims to be a base template for the SAT Reading and Writing Dataset, optimized for use with Hugging Face's datasets library.
Dataset Details
Dataset Description
This dataset contains SAT Reading and Writing assessment questions sourced from the College Board's SAT Suite Question Bank, intended for use in training and evaluating Language Models like LLMs.
Curated by: College Board
License:… See the full description on the dataset page: https://huggingface.co/datasets/betterMateusz/SAT_Writting_Reading_Assessment_Question_Bank.rtm-sgt-ocr-v1
Data Introduction
Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic Data from the ar𝜒iv for OCR Post Correction of Historic Scientific Articles".
Synthetic ground truth (SGT) sentences have been mined from the ar𝜒iv Bulk Downloads source documents,
and Optical Character Recognition (OCR)
sentences have been generated with the Tesseract OCR engine on the PDF pages generated from compiled source documents.… See the full description on the dataset page: https://huggingface.co/datasets/ReadingTimeMachine/rtm-sgt-ocr-v1.cia-declassified-reading-room
CIA Declassified Reading Room HF Library
Target account: manus4oHER
This project is a streaming pipeline for building a Hugging Face dataset mirror
of public CIA declassified Reading Room / CREST records without staging the
full corpus on this laptop.
The laptop stores only scripts, small manifests, and logs. Bulk crawling should
run in Hugging Face Jobs, one bounded page range per job. Each job uploads its
own shard and then exits.
Dataset Shape… See the full description on the dataset page: https://huggingface.co/datasets/manus4oHER/cia-declassified-reading-room.chunkr-reading-order-bench-oss
Chunkr Reading Order Bench - Open Source Subset
Open-source subset of the Chunkr Reading Order benchmark dataset, containing 733 professionally annotated documents with detailed reading order annotations across diverse document layouts.
This dataset benchmarks reading order detection models on complex, real-world documents including financial reports, legal contracts, research papers, medical records, and more. Each document includes ground truth reading order sequences essential… See the full description on the dataset page: https://huggingface.co/datasets/ChunkrAI/chunkr-reading-order-bench-oss.ReadingBank
ReadingBank
ReadingBank is a benchmark dataset for reading order detection built with weak supervision from WORD documents, which contains 500K document images with a wide range of document types as well as the corresponding reading order information.
Our paper "LayoutReader: Pre-training of Text and Layout for Reading Order Detection" has been accepted by EMNLP 2021.
Refer to the official repo for more details: https://github.com/doc-analysis/ReadingBank
parsinlu_reading_comprehension
Dataset Card for PersiNLU (Reading Comprehension)
Dataset Summary
A Persian reading comprehenion task (generating an answer, given a question and a context paragraph).
The questions are mined using Google auto-complete, their answers and the corresponding evidence documents are manually annotated by native speakers.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The text dataset is in Persian (fa).
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/parsinlu_reading_comprehension.french-bench-grammar-vocab-reading
Dataset Card for "french-bench-grammar-vocab-reading"
More Information needed
Meter_Readinglunde_nor_nob_reading_optimisedTest only - not for training.
First version - 0.1 of lunde_nor_nob_reading_optimised
This dataset does not contain any audio data.
Export Details
Train samples: 10040932
Validation samples: 0
Test samples: 0
Dataset created using search datasets:lunde_nor_nob_reading_optimised.
reading-comprehension-training-pool
Reading comprehension training pool
Public reading comprehension questions from six datasets, each a question about a passage with an
answer that is a span of it, a number or a date, read at the pinned revisions named below and laid
out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 310728 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/reading-comprehension-training-pool.chinese-reading-comprehensionreading-with-intent
📑 Paper | 📑 Blog
We introduce the Reading with Intent task and prompting method and accompanying datasets.
The goal of this task is to have LLMs read beyond the surface level of text and integrate an understanding of the underlying sentiment of a text when reading it. The focus of this work is on sarcastic text.
We've released:
The code used creating the sarcastic datasets
The sarcasm-poisoned dataset
The reading with intent prompting method
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Symblai/reading-with-intent.hat_asr_sixian_reading_clean
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Sixian
客語_四縣
203.03
106,493
2,137,244
6.86
2.92
Total
-
203.03
106,493
2,137,244
6.86
2.92
reading-comprehension-qa-training-pool
Reading comprehension question answering training pool
Public question answering data from five datasets, every row a question, the passages it is
answered from and every acceptable answer, read at the pinned revisions named below and laid out
twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 311462 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/reading-comprehension-qa-training-pool.law-reading-comprehension-qa-filteredlaw-reading-comprehension-qaTaiwan_Mandarin_Speech_Data_by_Mobile_Phone_Reading
Dataset Card for Nexdata/Taiwan_Mandarin_Speech_Data_by_Mobile_Phone_Reading
Dataset Summary
This dataset is just a sample of Taiwan Mandarin Speech dataset(paid dataset) by mobile phone reading.The data collects 204 Taiwan residents with 450 sentences for each speaker. The recorded is rich in content, including economy, entertainment, news, spoken language, numbers, letters, etc., covering general scenes and human-computer interaction scenes. Manual transcription of text… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/Taiwan_Mandarin_Speech_Data_by_Mobile_Phone_Reading.hat_asr_hailu_reading_clean
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Hailu
客語_海陸
197.58
104,555
2,018,861
6.80
2.84
Total
-
197.58
104,555
2,018,861
6.80
2.84
houseplant-light-needs-and-indoor-lux-readings
Houseplant Light Needs and Indoor Lux Readings
Two files, both CC BY 4.0. Archived with a DOI on Zenodo: 10.5281/zenodo.22023337.
Note for loaders: both CSVs open with commented header lines (#) carrying the snapshot date, the licence and the band definitions. Skip them when reading.
Files
species-light-levels.csv — one row per houseplant species in the GrowSpot care library, with the light tier it belongs to. Columns: common_name, scientific_name… See the full description on the dataset page: https://huggingface.co/datasets/Growspot/houseplant-light-needs-and-indoor-lux-readings.africa-synth-education-early-grade-reading-proficiency-all
Africa Synth Education Early Grade Reading Proficiency All | Africa (World Bank)
Size category: 10K<n<100K - Formats: csv - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Education datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-early-grade-reading-proficiency-all.right-reading-5bd6e9
right-reading-5bd6e9
Synthetic products test data: 39 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/EmberTrail299/right-reading-5bd6e9.law-reading-comprehension-qahat_asr_sixian_reading_clean_r
hat_asr_sixian_reading_clean_r
This dataset is an enhanced -R variant of formospeech/hat_asr_sixian_reading_clean.
Summary
Subset: Hakka_Sixian
Dialect: 客語四縣
Train samples: 106493
Audio: enhanced 24 kHz WAV
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Sixian
客語_四縣
203.03
106,493
5,597,645
6.86
7.66
Total
-
203.03
106,493
5,597,645
6.86
7.66
Processing
Start from the original… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/hat_asr_sixian_reading_clean_r.law-reading-comprehension-qareadingbank
ReadingBank (HF conversion)
Source paper: https://arxiv.org/abs/2108.11591
Original data: https://mail2sysueducn-my.sharepoint.com/:u:/g/personal/huangyp28_mail2_sysu_edu_cn/Efh3ZWjsA-xFrH2FSjyhSVoBMak6ypmbABWmJEmPwtKhhw?e=tbthMD
Created with: https://github.com/albertklor/reading-bank
Fields:
file_name (name of the file): str
page_number (index of the page number): int
bounding_boxes (normalized bounding boxes in [x0, y0, x1, y1] format): list[list[int]]
text… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/readingbank.hat_asr_sixian_reading_cm
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_sx
Hakka_Sixian
204.60
105,513
2,152,732
6.98
2.92
0
0
Total
-
204.60
105,513
2,152,732
6.98
2.92
0
0
hat_asr_hailu_reading_clean_r
hat_asr_hailu_reading_clean_r
This dataset is an enhanced -R variant of formospeech/hat_asr_hailu_reading_clean.
Summary
Subset: Hakka_Hailu
Dialect: 客語海陸
Train samples: 104555
Audio: enhanced 24 kHz WAV
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Hailu
客語_海陸
197.58
104,555
5,287,046
6.80
7.43
Total
-
197.58
104,555
5,287,046
6.80
7.43
Processing
Start from the original… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/hat_asr_hailu_reading_clean_r.nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing
nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 2,418 サンプル
validation: 302 サンプル
test: 303 サンプル
総サンプル数: 3,023
ソース
生成元: ./datasets/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing/
サンプルデータ
{
"instruction": "次の漢字の訓読み(くんよみ)をひらがなで答えてください。",
"input": "「究」の訓読みは?",
"think": "この漢字は「究」です。 小学3年生で習う漢字です。 意味は「research」などです。 訓読み(くんよみ)は日本語の読み方です。 この漢字の訓読みは「きわ」です。"… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing.
