contextlab/baum-corpus
ContextLab L. Frank Baum Corpus Dataset Description This dataset contains works of L. Frank Baum (1856-1919), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025). The corpus includes 14 books by L. Frank Baum, including The Wonderful Wizard of Oz series (14 books). All text has been converted to lowercase and… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/baum-corpus.
1149
