meettilavat/InternetArchive_1899_Chunked
Internet Archive Historical Texts - Chunked (0001-1899) TL;DR 163 million text chunks extracted from historical public-domain documents sourced from the Internet Archive Content dated 0001-1899, sorted by download popularity to prioritize high-quality, frequently accessed materials 2,445 Zstandard-compressed Parquet shards totaling ~217 GB on disk, ~594 billion characters uncompressed Optimized chunk size of ~3,600 characters (target: 4,000) for efficient… See the full description on the dataset page: https://huggingface.co/datasets/meettilavat/InternetArchive_1899_Chunked.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face