Svngoku/Kongo-bernard-de-grunne
Kongo-bernard-de-grunne Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline. Dataset Summary Metric Value Total chunks 222 Avg chars/chunk 709 Avg images/chunk 0.72 Source files 1 Duplicates removed 0 Quality filtered 4 Schema Column Type Description chunk_id string Unique identifier: filename_chunk_N text string Raw markdown chunk with image refs text_clean string Cleaned text without… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/Kongo-bernard-de-grunne.
Kongo-bernard-de-grunne
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Schema
Source Files
Document Structure
Chapters found in the source documents:
- 02
- 03
- 04
- 05
- 06
- 08
- 11
- 13
- 14
- 15
- 16
- 18
- 22
- 23
- 25
- 26
- 28
- 30
- 32
- 34
- 36
- 40
- 44
- 47
- 48
- 52
- 53
- 55
- 56
- 58
- ... and 14 more
Pipeline
This dataset was processed through the PDF2Dataset pipeline:
- OCR -- Mistral OCR extracts text and images from PDF/image files
- Chunking -- Structure-aware recursive splitting preserves document hierarchy
- Validation -- Schema validation ensures every chunk has required fields
- Deduplication -- Content-hash based dedup removes identical chunks
- Quality Filtering -- Removes empty, too-short, or malformed chunks
