CoolFace
Datasetpublic

mhla/pre1900-corpus

Pre-1900 Corpus The training corpus for GPT-1900 — a cleaned collection of pre-1900 English-language texts with full metadata. Every document in this corpus was published before the year 1900. Schema Column Type Description text string Full document text year int64 Publication year title string Book title or newspaper name source string Source dataset identifier ocr_score float64 OCR confidence score (-1.0 if unavailable) legibility float64… See the full description on the dataset page: https://huggingface.co/datasets/mhla/pre1900-corpus.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
4likes152downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
mhla/pre1900-corpus · CoolFace