CoolFace
Datasetpublic

amazingvince/survivor-library-text

Survivor Library Text Corpus This dataset contains OCR-extracted text from the Survivor Library, a collection of public domain books focused on practical knowledge and skills from the pre-industrial and early industrial era. Dataset Description The Survivor Library is a collection of books that would be useful in rebuilding civilization after a catastrophic event. It focuses on practical, hands-on knowledge from the 1800s and early 1900s. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/amazingvince/survivor-library-text.

sourceHugging Facecc0-1.0updated 1y agoView on Hugging Face
2likes214downloads
Dataset Card

Survivor Library Text Corpus

This dataset contains OCR-extracted text from the Survivor Library, a collection of public domain books focused on practical knowledge and skills from the pre-industrial and early industrial era.

Dataset Description

The Survivor Library is a collection of books that would be useful in rebuilding civilization after a catastrophic event. It focuses on practical, hands-on knowledge from the 1800s and early 1900s.

Dataset Summary

  • —Total Files: 0
  • —Total Size: 0.00 GB
  • —Categories: 187
  • —Source: survivorlibrary.com
  • —Processing Date: 2025-07-02

Categories and File Counts

CategoryFiles
Accounting0
Aeroplanes0
Airships0
Archery0
Architecture0
Astronomy0
Baking0
Banking0
Basketry0
BeeJournalAmerican0
BeeJournalBritish0
Beekeeping0
Berries0
Blacksmithing0
Boat_Building0
Boilermaker0
Bookbinding0
BooksforBoysandGirls0
BooksforYoung_Children0
Botany0
BoyScoutManuals0
BrewingandDistilling0
BrewingandMalting0
Brickmaking0
BridgesandDams0
Bronze_Working0
Building0
Butchering0
Canning0
Carburetor0
Carpentry0
Carriage_Building0
Cement0
Charcoal0
Cheese0
CheeseandButter0
Chemistry0
Chlorine0
Christmas0
Clockmaking0
Clocks0
Clothing0
Coal_Gas0
CoalandMining0
CoffeeandTea0
Concrete0
ConductofLife0
Construction0
CookingandCookbooks0
Coopering0
Cotton0
Courthhouses0
Cutlery0
CyclesBiTri_Motor0
Dairying0
Dentistry0
Dictionaries0
Diesel0
Distilling0
Dogs0
Drilling0
Dyeing0
Economics0
Embalming0
Encyclopedias0
Engineering_Drainage0
Engineering_Electrical0
Engineering_General0
Engineering_Hydraulics0
EngravingandWoodcuts0
Ethics0
Farming0
Farming_Corn0
Farming_Fish0
FarmingPotatoandSweetPotato0
Firearms_Books0
Firearms_Manuals0
Fishing0
Food0
Forestry0
ForgingandCasting0
Formulas0
Fuels0
Geodesy0
Geography0
Glassmaking0
GrapesWineRaisins0
Great_Books0
GunpowderandExplosives0
Hatmaking0
Heating0
HeavyIndustrialMachinery0
HempandFlax0
Herbalism0
History_American0
Home_Economics0
Horses0
Journalism0
KnittingLaceNeedlepoint0
Laundry0
Law0
Leather0
LeisureGamesand_Sports0
LeisureRecreationMagazine0
Leisure_Whist0
Lithography0
Livestock_Cattle0
LivestockRabbitsand_Cavies0
Livestock_Sheep0
Livestock_Swine0
Machine_Tools0
Machinerys_Reference0
MasterpiecesofEloquence0
Mathematics0
Mechanical_Drawing0
Medical_Anesthesia0
MedicalCoursesUS_Army0
Medical_Diagnostics0
Medical_Emergency0
Medical_Hypnotism0
MedicalMedicine1900-19220
Medical_Microscopy0
Medical_Nursing0
MedicalObstetrics1900-19220
MedicalSurgery1900-19220
MedicalSurgery20
MedicalXRays0
Meteorology0
Mimeograph0
Miscellaneous0
Monasticism0
Morality0
Mushrooms0
Musical_Instruments0
NBC0
Navigation0
Opium0
Optometry0
Painting0
Papermaking0
Photography0
Pottery0
Poultry0
Primers0
Printing0
Radio0
Radio73Magazine0
Railroads0
Rat_Control0
Refrigeration0
Sanitation0
ScientificAmericanSeries_10
ScientificAmericanSeries_20
Sewage0
Sewing0
Shelter0
Shipbuilding0
Shoemaking0
Shorthand0
Silk_Culture0
Sliderules0
Smithing0
Steam_Engines0
StoneandMasonry0
Surveying0
Survival_Individual0
Teaching0
Teaching_Arithmetic0
Teaching_Civics0
Teaching_Phonics0
Teaching_Readers0
TeachingReadersMcGuffey0
TelegraphandTelephone0
Thanksgiving0
Tobacco0
Toys0
TrappingandHunting0
TurpentineGlueSolvents0
Veterinary0
WagonsandCoaches0
Weaving0
Welding0
WindandWater0
Wood_Carpentry0
Wood_Carving0
Wood_Furniture0
World_Depression0

Dataset Structure

The dataset is organized by category, with each category containing text files extracted from PDFs:

text_outputs/
├── Accounting/
│   ├── book1.txt
│   ├── book2.txt
│   └── ...
├── Agriculture/
│   └── ...
└── ...

Data Processing

  1. 1.Original PDFs downloaded from survivorlibrary.com
  2. 2.Text extracted using OCR/PDF text extraction
  3. 3.Organized by category
  4. 4.Uploaded to Hugging Face Hub

Usage

python
from datasets import load_dataset

# Load the entire dataset
dataset = load_dataset("amazingvince/survivor-library-text")

# Load specific categories
accounting_texts = dataset.filter(lambda x: x['category'] == 'Accounting')

Considerations

  • —These texts are historical and may contain outdated or potentially dangerous information
  • —Always verify information with modern sources before practical application
  • —Some OCR errors may be present in the extracted text
  • —Original formatting may not be perfectly preserved

License

The original books are in the public domain. This dataset compilation is released under CC0 1.0 Universal.

Citation

If you use this dataset, please cite:

bibtex
@misc{survivor_library_text,
  title={Survivor Library Text Corpus},
  author={Survivor Library},
  year={2024},
  publisher={Hugging Face},
  url={https://huggingface.co/datasets/amazingvince/survivor-library-text}
}
amazingvince/survivor-library-text · CoolFace