dotoking/wwmad
Source Viastoica.com Project GutenbergOCR’d and cleaned using Tesseract in Google Colab. Intended use: Grounding philosophical Q&A responses in a Retrieval-Augmented Generation (RAG) system supporting the WWMAD bot. License: This dataset includes material believed to be in the public domain or publicly redistributable. Users are advised to verify content license terms before republishing or using commercially. Cleaning: Removed links, author signatures, and misrecognized footers. OCR… See the full description on the dataset page: https://huggingface.co/datasets/dotoking/wwmad.
Source
- Viastoica.com
- Project Gutenberg OCR’d and cleaned using Tesseract in Google Colab.
Intended use: Grounding philosophical Q&A responses in a Retrieval-Augmented Generation (RAG) system supporting the WWMAD bot.
License: This dataset includes material believed to be in the public domain or publicly redistributable. Users are advised to verify content license terms before republishing or using commercially.
Cleaning: Removed links, author signatures, and misrecognized footers. OCR artifacts may remain and will be refined in future versions.
