Mdean77/quickb-kb
quickb-kb Generated using QuicKB, a tool developed by Adam Lucek. QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization. Chunking Configuration Chunker: RecursiveTokenChunker Parameters: chunk_size: 400 chunk_overlap: 0 length_type: 'character' separators: ['\n\n', '\n', '.', '?', '!', ' ', ''] keep_separator: True… See the full description on the dataset page: https://huggingface.co/datasets/Mdean77/quickb-kb.
quickb-kb
Generated using QuicKB, a tool developed by Adam Lucek.
QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization.
Chunking Configuration
- Chunker: RecursiveTokenChunker
- Parameters:
- chunk_size:
400 - chunk_overlap:
0 - length_type:
'character' - separators:
['\n\n', '\n', '.', '?', '!', ' ', ''] - keep_separator:
True - is_separator_regex:
False
Dataset Statistics
- Total chunks: 526
- Average chunk size: 57.8 words
- Source files: 1
Dataset Structure
This dataset contains the following fields:
text: The content of each text chunksource: The source file path for the chunkid: Unique identifier for each chunk
