CoolFace
Datasetpublic

meettilavat/InternetArchive_1899_Chunked

Internet Archive Historical Texts - Chunked (0001-1899) TL;DR 163 million text chunks extracted from historical public-domain documents sourced from the Internet Archive Content dated 0001-1899, sorted by download popularity to prioritize high-quality, frequently accessed materials 2,445 Zstandard-compressed Parquet shards totaling ~217 GB on disk, ~594 billion characters uncompressed Optimized chunk size of ~3,600 characters (target: 4,000) for efficient… See the full description on the dataset page: https://huggingface.co/datasets/meettilavat/InternetArchive_1899_Chunked.

sourceHugging Faceotherupdated 11mo agoView on Hugging Face
0likes178downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
meettilavat/InternetArchive_1899_Chunked · CoolFace