zachnorton03/bart-midtrain
BART Midtrain The midtraining corpus for BART — pre-1930 mathematics, science, technology, and medicine — plus the full pipeline that built it and every training mixture it was blended into. Built by Unbounded Labs. Corpus documents 11,409 Corpus characters 2,543,809,124 Corpus tokens ~604M Removed by cleaning 24% of documents (15,075 → 11,409) Subject focus math, science, technology, medicine Cutoff 1930 Schema single string column text… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/bart-midtrain.
This repository belongs to zachnorton03 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
