jbduran/bartholomew-dataset-v1
BART Dataset v1 The first version of the BART pretraining corpus: pre-1930 English books drawn from Institutional Books 1.0 and filtered hard on OCR quality, language, date, and tokenizability. Documents 160,263 Characters 118,745,375,871 Tokens ~27B (estimated) Shards 473 (472 train + 1 val) Source Institutional Books 1.0 (242B tokens, ~983K documents) Schema single string column text Lineage — three cumulative filtering stages over the same corpus:… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v1.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face