CoolFace
Datasetpublic

contextlab/baum-corpus

ContextLab L. Frank Baum Corpus Dataset Description This dataset contains works of L. Frank Baum (1856-1919), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025). The corpus includes 14 books by L. Frank Baum, including The Wonderful Wizard of Oz series (14 books). All text has been converted to lowercase and… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/baum-corpus.

sourceHugging Facemitupdated 11mo agoView on Hugging Face
1likes149downloads

contextlab/baum-corpus · main · files are served by the source, never re-hosted here