CoolFace
Datasetpublic

dotoking/wwmad

Source Viastoica.com Project GutenbergOCR’d and cleaned using Tesseract in Google Colab. Intended use: Grounding philosophical Q&A responses in a Retrieval-Augmented Generation (RAG) system supporting the WWMAD bot. License: This dataset includes material believed to be in the public domain or publicly redistributable. Users are advised to verify content license terms before republishing or using commercially. Cleaning: Removed links, author signatures, and misrecognized footers. OCR… See the full description on the dataset page: https://huggingface.co/datasets/dotoking/wwmad.

sourceHugging Faceotherupdated 1y agoView on Hugging Face
0likes9downloads
filechroma_db_export.zip19.6 MBdownload
fileWWMAD_raw_data.zip467 KBdownload

dotoking/wwmad · main · files are served by the source, never re-hosted here