datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 5.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 1.3000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 4.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 2.2000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 3.1000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 3.4000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 7.0000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 8.0000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think.historical-training-manuals
Historical Training Manuals
1,597 US government and government-adjacent training manuals and technical publications
sourced from the Internet Archive, spanning roughly 1800-2021. Records carry
bibliographic metadata; a subset also carries extracted full text and a machine-generated
summary.
Loading
from datasets import load_dataset
ds = load_dataset("robworks-software/historical-training-manuals")
Splits
Split
Rows
train
1,277… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/historical-training-manuals.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 6.2000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-10m-historical-20t-think.HistoricalPortugueseCorpora
Historical Portuguese Corpora
Resource for the development of PortOldBERT, the first Portuguese Historical Language Models
How to load the dataset:
from datasets import load_dataset
dataset = load_dataset("LIACC/HistoricalPortugueseCorpora")
Citation
When using or citing this model, kindly cite the following publication:
@inproceedings{osorio-lopes-cardoso-2026-portoldbert,
title = "{P}ort{O}ld{BERT}: {P}ortuguese Historical Language Models",
author =… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/HistoricalPortugueseCorpora.swahili-historical-corpus-pd
Swahili Historical Corpus — Public Domain Sources
Historical Swahili linguistic data from 19th century public domain dictionaries and texts.
Structured for NLP training, language model development, and cultural AI applications.
Swahili is spoken by 200+ million people across East and Central Africa.
This corpus addresses the documented data gap in Swahili NLP resources.
Sources (All Public Domain)
All works published before 1928:
Author
Work
Date
Status… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/swahili-historical-corpus-pd.english-historical-corpus-1800-1875
TimeCapsuleLLM World English 1800-1875
This dataset is the sharded pretraining corpus prepared for TimeCapsuleLLM-World-English-1800-1875-v1. It is a large English-language historical text collection centered on the years 1800-1875, with a mix of books, pamphlets, reports, newspapers, prose extracts, and other running text suitable for language-model pretraining.
The corpus was assembled from several cleaned source groups and then normalized into compressed JSONL shards for… See the full description on the dataset page: https://huggingface.co/datasets/haykgrigorian/english-historical-corpus-1800-1875.
