bart
Datasets
All datasets matching “bart”bartholomew-dataset-v3
BART Dataset v3
The final version of the BART pretraining corpus, focused on removing anything that betrays a
post-1930 origin. This is our best vintage dataset yet.
Documents
146,031 (97.52% of v2)
Characters
102,798,688,961 (96.73% of v2)
Tokens
~23B (estimated)
Shards
473 (one per v2 shard, same basename)
Source
BART Dataset v2
Cutoff
1930
Schema
single string column text
Lineage — three cumulative filtering stages over the same corpus:… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v3.bartholomew-dataset-v1
BART Dataset v1
The first version of the BART pretraining corpus: pre-1930 English books drawn from
Institutional Books 1.0
and filtered hard on OCR quality, language, date, and tokenizability.
Documents
160,263
Characters
118,745,375,871
Tokens
~27B (estimated)
Shards
473 (472 train + 1 val)
Source
Institutional Books 1.0 (242B tokens, ~983K documents)
Schema
single string column text
Lineage — three cumulative filtering stages over the same corpus:… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v1.variouscryptodata
variouscryptodata
Crypto market datasets collected as a by-product of our own research and
published so they are not lost. One sub-folder per dataset; each appended
nightly where collection is still running.
folder
what
coverage
cadence
polymarket_updown_orderbook/
Polymarket Up/Down (5m/15m) order books, 10 levels, BTC/ETH/SOL/XRP/DOGE/HYPE/BNB, with Binance spot reference
2026-05-24 → present
appended nightly (previous UTC day)
hyperliquid_trades/
Hyperliquid perp… See the full description on the dataset page: https://huggingface.co/datasets/Barthel/variouscryptodata.bartholomew-dataset-v2
BART Dataset v2
The second version of the BART pretraining corpus, focused on stripping low-quality text —
boilerplate and OCR corruption — out of
v1.
Documents
149,745 (93.44% of v1)
Characters
106,274,384,672 (89.50% of v1)
Tokens
~24B (estimated)
Shards
473 (one per v1 shard, same basename)
Source
BART Dataset v1
Schema
single string column text
Lineage — three cumulative filtering stages over the same corpus:
Institutional Books 1.0 →
v1 →
v2 →
v3… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v2.EventHubDatasetspeech_commands
Dataset Card for "speech_commands"
More Information needed
