institutional
institutional-newspapers-bpl
📰 Institutional Newspapers: Boston Public Library
A structured dataset derived from the Boston Public Library's public domain
newspapers collection, produced by the Institutional Data
Initiative in collaboration with Boston Public Library.
1,473,635 public domain newspaper scans, published between 1795 and 1930
83,147,041 individual crops segmented from those scans
16.3 billion o200k_base tokens of VLM OCR text, and 14.7 billion from Tesseract
Data for each crop: bbox… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-newspapers-bpl.institutional-books-hl
📚 Institutional Books: Harvard Library
Institutional Books is a growing corpus of public domain books. This release is comprised of 983,004 public domain books digitized as part of Harvard Library's participation in the Google Books project and refined by the Institutional Data Initiative. Use of this data is governed by the IDI Terms of Use for Early-Access.
983K books, published largely in the 19th and 20th centuries
242B o200k_base tokens
386M pages of text, available in… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl.institutional-books-hl-visual-elements
📚 Institutional Books: Harvard Library — Visual Elements
22 million visual elements extracted from the volumes that comprise the Institutional Books: Harvard Library dataset.
22,622,060 visual elements extracted from 983,004 volumes
766,992,447 o200k_base tokens in AI-generated captions
6 high-level classes of visual elements organized in splits
5 processing steps: Detection, Classification, Deduplication, Captioning, and Rotation
The Institutional Data Initiative at Harvard… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-visual-elements.institutional-books-hl-enriched-text
📚 Institutional Books: Harvard Library — Enriched Text
Institutional Books is a growing corpus of public domain books.
This release (IB-HL-ET) is a version of the text present in the Institutional Books: Harvard Library
(IB-HL) dataset
that has been further processed, filtered and optimized for computational access and model training.
This includes:
983K books, published largely in the 19th and 20th centuries
217B o200k_base tokens
7B sentences in 250 languages, grouped into… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-enriched-text.institutional-portfolio-13f
ZipLime Institutional Fund Portfolios (PIT)
[!CAUTION]
A 13F is stale by construction. It reports a quarter that ended up to 45
days earlier, and amendments restate it for months afterwards — the longest in
this dataset arrived 452 days late. Use knowledge_date, never
event_date, to decide what a strategy could act on.
Every Form 13F holdings report filed with the U.S. SEC since structured filing
began, turned into point-in-time portfolios: what each institutional manager
held… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/institutional-portfolio-13f.institutional-holdings-13f-quarterlyUS Institutional 13F Holdings Record-level Embeddings Dataset (Merged Quarterly Snapshots 1980–2024)
This dataset provides a single merged Hugging Face DatasetDict combining all 179 quarterly snapshots of U.S. institutional 13F filings (1980 Q1 → 2024 Q3).
Every quarter’s folder has been concatenated into three splits:
holdings (113,114,724 rows): record-level positions with fields:
• mgrno (string) – Institutional manager ID (SEC)
• permco (string) – Permanent company identifier
•… See the full description on the dataset page: https://huggingface.co/datasets/kurry/institutional-holdings-13f-quarterly.
