LumberChunker/GutenQA_Semantic
š GutenQA-Semantic GutenQA-Semantic consists on the same 100 Public Domain Narrative Books used in GutenQA (the proposed benchmark to the paper LumberChunker: Long-Form Narrative Document Segmentation, and serves as one of the baseline chunking approaches utilized on the LumberChunker paper. In this version, passages are segmented with Semantic Chunking, which utilizes embeddings to cluster semantically similar text segments. The dataset is organized into the following columns:⦠See the full description on the dataset page: https://huggingface.co/datasets/LumberChunker/GutenQA_Semantic.
012
