CoolFace
Datasetpublic

jbduran/bartholomew-dataset-v1

BART Dataset v1 The first version of the BART pretraining corpus: pre-1930 English books drawn from Institutional Books 1.0 and filtered hard on OCR quality, language, date, and tokenizability. Documents 160,263 Characters 118,745,375,871 Tokens ~27B (estimated) Shards 473 (472 train + 1 val) Source Institutional Books 1.0 (242B tokens, ~983K documents) Schema single string column text Lineage — three cumulative filtering stages over the same corpus:… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v1.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
1likes1.2kdownloads
16 commits on main
63697151mo ago

Add MIT license to dataset card metadata

jbduran
9e7efae1mo ago

Rewrite dataset card: BART v-series naming, provenance narrative, lineage links

jbduran
8e502164mo ago

Update README.md

jbduran
bbdbf654mo ago

Update README.md

jbduran
465d0c84mo ago

Add README + manifest

jbduran
cd49bc34mo ago

Add 23 shard(s): shard_00450.parquet..shard_00472.parquet

jbduran
fff38484mo ago

Add 50 shard(s): shard_00400.parquet..shard_00449.parquet

jbduran
66c6ab64mo ago

Add 50 shard(s): shard_00350.parquet..shard_00399.parquet

jbduran
babce0f4mo ago

Add 50 shard(s): shard_00300.parquet..shard_00349.parquet

jbduran
91b41394mo ago

Add 50 shard(s): shard_00250.parquet..shard_00299.parquet

jbduran
c59838e4mo ago

Add 50 shard(s): shard_00200.parquet..shard_00249.parquet

jbduran
dd5d07b4mo ago

Add 50 shard(s): shard_00150.parquet..shard_00199.parquet

jbduran
de8c7d14mo ago

Add 50 shard(s): shard_00100.parquet..shard_00149.parquet

jbduran
fe2e8334mo ago

Add 50 shard(s): shard_00050.parquet..shard_00099.parquet

jbduran
63809aa4mo ago

Add 50 shard(s): shard_00000.parquet..shard_00049.parquet

jbduran
96e76ec4mo ago

initial commit

jbduran