datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
institutional-newspapers-bpl
📰 Institutional Newspapers: Boston Public Library
A structured dataset derived from the Boston Public Library's public domain
newspapers collection, produced by the Institutional Data
Initiative in collaboration with Boston Public Library.
1,473,635 public domain newspaper scans, published between 1795 and 1930
83,147,041 individual crops segmented from those scans
16.3 billion o200k_base tokens of VLM OCR text, and 14.7 billion from Tesseract
Data for each crop: bbox… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-newspapers-bpl.institutional-books-hl-visual-elements
📚 Institutional Books: Harvard Library — Visual Elements
22 million visual elements extracted from the volumes that comprise the Institutional Books: Harvard Library dataset.
22,622,060 visual elements extracted from 983,004 volumes
766,992,447 o200k_base tokens in AI-generated captions
6 high-level classes of visual elements organized in splits
5 processing steps: Detection, Classification, Deduplication, Captioning, and Rotation
The Institutional Data Initiative at Harvard… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-visual-elements.
