pre-1900
Datasets
All datasets matching “pre-1900”pre1900-training
Pre-1900 Training Corpus
Chunked and resharded pre-1900 English text corpus, ready for language model training.
Format
266 parquet shards (265 train + 1 validation)
12.8M documents (chunks of ≤8,000 characters)
~22B tokens estimated
Text-only — single text column per row
Row groups divisible by 8 for even DDP distribution across GPUs
Last shard (shard_00265) is the validation split
Processing Pipeline
Built from the full pre-1900 filtered corpus through:
OCR… See the full description on the dataset page: https://huggingface.co/datasets/mhla/pre1900-training.pre1900-corpus
Pre-1900 Corpus
The training corpus for GPT-1900 — a cleaned collection of pre-1900 English-language texts with full metadata. Every document in this corpus was published before the year 1900.
Schema
Column
Type
Description
text
string
Full document text
year
int64
Publication year
title
string
Book title or newspaper name
source
string
Source dataset identifier
ocr_score
float64
OCR confidence score (-1.0 if unavailable)
legibility
float64
Legibility… See the full description on the dataset page: https://huggingface.co/datasets/mhla/pre1900-corpus.pre1900-verifiable-physicspre1900-academics-corpuspre1900-catalogue
Pre-1900 English & German Catalogue
Universal catalogue of pre-1900 (or undated) English and German works: 2,344,644 works, 2,047,100 with full text (869.4GB).
Browse & read: https://quasar7-pre1900-catalogue-site.static.hf.space (serverless — searches and reads run in your browser against this repo)
Query here: use the SQL Console on this page (the catalogue config is the works table)
Layout
path
contents
site/catalog_flat.parquet
the catalogue: one… See the full description on the dataset page: https://huggingface.co/datasets/quasar7/pre1900-catalogue.
