dell-research-harvard/AmericanStories
American Stories offers high-quality structured data from historical newspapers suitable for pre-training large language models to enhance the understanding of historical English and world knowledge. It can also be integrated into external databases of retrieval-augmented language models, enabling broader access to historical information, including interpretations of political events and intricate details about people's ancestors. Additionally, the structured article texts facilitate the application of transformer-based methods for popular tasks like detecting reproduced content, significantly improving accuracy compared to traditional OCR methods. American Stories serves as a substantial and valuable dataset for advancing multimodal layout analysis models and other multimodal applications.
Update AmericanStories.py
Amend dataset (#14)
Create LICENSE
Readme
Fixing date and article id problems
updating for data streaming
fixing all_years option
Add lam tag (#2)
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Updated Dataset Card
add raw data option
added raw text
error handling
error handling
finalise docs
doc changes
test
test
test
test
test
test
test
finalise args
path changes
path changes
test
modified url dict
final changes
final testing
modified loader
testing_loader
changed_loader
Remaining files added
First version of the AmericanStories dataset.
add data files - 1
initial commit
