Svngoku/mande-ancient-treasures-de-grunne-van-dyke-2016
mande-ancient-treasures-de-grunne-van-dyke-2016 Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline. Dataset Summary Metric Value Total chunks 268 Avg chars/chunk 722 Avg images/chunk 0.14 Source files 1 Duplicates removed 0 Quality filtered 6 Schema Column Type Description chunk_id string Unique identifier: filename_chunk_N text string Raw markdown chunk with image refs text_clean… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/mande-ancient-treasures-de-grunne-van-dyke-2016.
mande-ancient-treasures-de-grunne-van-dyke-2016
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Schema
Source Files
Document Structure
Chapters found in the source documents:
- 04
- 06
- 08
- 10
- 13
- 15
- 18
- 19
- 23
- 24
- 28
- Preface
- Préface
Pipeline
This dataset was processed through the PDF2Dataset pipeline:
- OCR -- Mistral OCR extracts text and images from PDF/image files
- Chunking -- Structure-aware recursive splitting preserves document hierarchy
- Validation -- Schema validation ensures every chunk has required fields
- Deduplication -- Content-hash based dedup removes identical chunks
- Quality Filtering -- Removes empty, too-short, or malformed chunks
