datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wwmadSource
Viastoica.com
Project GutenbergOCR’d and cleaned using Tesseract in Google Colab.
Intended use: Grounding philosophical Q&A responses in a Retrieval-Augmented Generation (RAG) system supporting the WWMAD bot.
License: This dataset includes material believed to be in the public domain or publicly redistributable. Users are advised to verify content license terms before republishing or using commercially.
Cleaning: Removed links, author signatures, and misrecognized footers. OCR… See the full description on the dataset page: https://huggingface.co/datasets/dotoking/wwmad.wikimia
