biglam/loc_beyond_words
Dataset Card for Beyond Words Dataset Summary The Beyond Words dataset is a crowdsourced collection of bounding box annotations on World War I-era historical newspaper pages from the Library of Congress’s Chronicling America collection. Volunteers marked seven types of visual content — photographs, illustrations, maps, comics, editorial cartoons, headlines, and advertisements — enabling the training of the visual content recognition model behind the Newspaper… See the full description on the dataset page: https://huggingface.co/datasets/biglam/loc_beyond_words.
switch to parquet version of dataset (#3)
Upload dataset (#2)
Update README.md (#1)
Update README.md
Update loc_beyond_words.py
Update README.md
Update README.md
Update README.md
add readme
draft dataset
initial commit
