CoolFace
Datasetpublic

dotoking/wwmad

Source Viastoica.com Project GutenbergOCR’d and cleaned using Tesseract in Google Colab. Intended use: Grounding philosophical Q&A responses in a Retrieval-Augmented Generation (RAG) system supporting the WWMAD bot. License: This dataset includes material believed to be in the public domain or publicly redistributable. Users are advised to verify content license terms before republishing or using commercially. Cleaning: Removed links, author signatures, and misrecognized footers. OCR… See the full description on the dataset page: https://huggingface.co/datasets/dotoking/wwmad.

sourceHugging Faceotherupdated 1y agoView on Hugging Face
0likes9downloads
7 commits on main
84114a41y ago

Upload WWMAD_chunk_embedpersistant.ipynb

dotoking
c1ab32b1y ago

Upload chroma_db_export.zip

dotoking
f99b0061y ago

Upload wwmad_ocr_and_clean.ipynb

dotoking
c0244a11y ago

Update README.md

dotoking
42af1911y ago

Upload WWMAD_raw_data.zip

dotoking
e9dd7831y ago

Create README.md

dotoking
76097521y ago

initial commit

dotoking