internet-archive
InternetArchive_1899_Large
Internet Archive Historical Texts (0001-1899)
TL;DR
711,680 cleaned public-domain style documents harvested from the Internet Archive via a high-throughput text-to-parquet pipeline.
Coverage targets items that contain textual content dated between 0001 and 1899, ranked by download counts; ~715k IDs were attempted, ~4.1k were filtered during preprocessing.
Stored in 620 Zstandard-compressed Parquet shards (shard_00000.parquet ... shard_00619.parquet) occupying ~240 GB on… See the full description on the dataset page: https://huggingface.co/datasets/meettilavat/InternetArchive_1899_Large.InternetArchive_1899_Chunked
Internet Archive Historical Texts - Chunked (0001-1899)
TL;DR
163 million text chunks extracted from historical public-domain documents sourced from the Internet Archive
Content dated 0001-1899, sorted by download popularity to prioritize high-quality, frequently accessed materials
2,445 Zstandard-compressed Parquet shards totaling ~217 GB on disk, ~594 billion characters uncompressed
Optimized chunk size of ~3,600 characters (target: 4,000) for efficient language model… See the full description on the dataset page: https://huggingface.co/datasets/meettilavat/InternetArchive_1899_Chunked.internet_archive_azerbaijaniInternet-Archive-Unfiltered-Sharegpt
The Dataset is provided ""AS IS"" and ""AS AVAILABLE"" without warranty of any kind, express or implied, including but not limited to warranties of merchantability, fitness for a particular purpose, title, or non-infringement.
The Provider disclaims all liability for any damages or losses resulting from the use or misuse of the Dataset, including but not limited to any damages or losses arising from the use of the Dataset for purposes other than those intended by the Provider.
The Provider… See the full description on the dataset page: https://huggingface.co/datasets/NewEden-Forge/Internet-Archive-Unfiltered-Sharegpt.
