datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Quickbooks_LLM_Training_SampleDataset Summary
(Not affiliated with Quickbooks)
A comprehensive training dataset sample of realistic, synthetic generated QuickBooks Online API interaction scenarios, specifically designed for training AI assistants, chatbots, and automation tools on QuickBooks accounting workflows. Each scenario includes natural language user requests, properly formatted API calls, realistic QuickBooks API responses, and human-readable summaries covering the complete lifecycle of customers, invoices… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Quickbooks_LLM_Training_Sample.quickb-kb
quickb-kb
Generated using QuicKB, a tool developed by Adam Lucek.
QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization.
Chunking Configuration
Chunker: RecursiveTokenChunker
Parameters:
chunk_size: 400
chunk_overlap: 0
length_type: 'character'
separators: ['\n\n', '\n', '.', '?', '!', ' ', '']
keep_separator: True… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/quickb-kb.quickb
quickb
Generated using QuicKB, a tool developed by Adam Lucek.
QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization.
Chunking Configuration
Chunker: RecursiveTokenChunker
Parameters:
chunk_size: 400
chunk_overlap: 50
length_type: 'character'
separators: ['\n\n', '\n', '.', '?', '!', ' ', '']
keep_separator: True… See the full description on the dataset page: https://huggingface.co/datasets/Mitchell6024/quickb.quickb-qa
quickb-qa
Generated using QuicKB, a tool developed by Adam Lucek.
QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization.
Question Generation
Model: openai/gpt-4o-mini
Deduplication threshold: 0.85
Results:
Total questions generated: 13292
Questions after deduplication: 11116
Dataset Structure
anchor: The generated… See the full description on the dataset page: https://huggingface.co/datasets/Mitchell6024/quickb-qa.quickb-kb
quickb-kb
Generated using QuicKB, a tool developed by Adam Lucek.
QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization.
Chunking Configuration
Chunker: RecursiveTokenChunker
Parameters:
chunk_size: 400
chunk_overlap: 0
length_type: 'character'
separators: ['\n\n', '\n', '.', '?', '!', ' ', '']
keep_separator: True… See the full description on the dataset page: https://huggingface.co/datasets/Mdean77/quickb-kb.quickb-qa
quickb-qa
Generated using QuicKB, a tool developed by Adam Lucek.
QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization.
Question Generation
Model: openai/gpt-4o-mini
Deduplication threshold: 0.85
Results:
Total questions generated: 80
Questions after deduplication: 80
Dataset Structure
anchor: The generated… See the full description on the dataset page: https://huggingface.co/datasets/onecd2000/quickb-qa.quickb-qa
quickb-qa
Generated using QuicKB, a tool developed by Adam Lucek.
QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization.
Question Generation
Model: openai/gpt-4o-mini
Deduplication threshold: 0.85
Results:
Total questions generated: 148
Questions after deduplication: 142
Dataset Structure
anchor: The generated… See the full description on the dataset page: https://huggingface.co/datasets/Nuf-hugginface/quickb-qa.quickb-qa
quickb-qa
Generated using QuicKB, a tool developed by Adam Lucek.
QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization.
Question Generation
Model: openai/gpt-4o-mini
Deduplication threshold: 0.85
Results:
Total questions generated: 2098
Questions after deduplication: 1742
Dataset Structure
anchor: The generated… See the full description on the dataset page: https://huggingface.co/datasets/Mdean77/quickb-qa.quickb-qa
quickb-qa
Generated using QuicKB, a tool developed by Adam Lucek.
QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization.
Question Generation
Model: openai/gpt-4o-mini
Deduplication threshold: 0.85
Results:
Total questions generated: 11111
Questions after deduplication: 9724
Dataset Structure
anchor: The generated… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/quickb-qa.quickb-kb
quickb-kb
Generated using QuicKB, a tool developed by Adam Lucek.
QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization.
Chunking Configuration
Chunker: RecursiveTokenChunker
Parameters:
chunk_size: 400
chunk_overlap: 0
length_type: 'character'
separators: ['\n\n', '\n', '.', '?', '!', ' ', '']
keep_separator: True… See the full description on the dataset page: https://huggingface.co/datasets/Nuf-hugginface/quickb-kb.quick_batman_loves_tylerPerryA easy quick llm chat dataset for a batman who loves tyler perry but not as much as he loves JUSTICE
Base Prompt:
You are batman you will alwsy talk in a dark gloomy tone, you will alwasy redirect the conversation to ebing batman, being an orphan and fighting your many enemies, be creative. you will also throw in a last thing about how great the tyler perry movie is but its nothing in comparision to JUSTICE
Was generated through:… See the full description on the dataset page: https://huggingface.co/datasets/superdrew100/quick_batman_loves_tylerPerry.quickb-qa
quickb-qa
Generated using QuicKB, a tool developed by Adam Lucek.
QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization.
Question Generation
Model: huggingface/starcoder
Deduplication threshold: 0.85
Results:
Total questions generated: 0
Questions after deduplication: 0
Dataset Structure
anchor: The generated… See the full description on the dataset page: https://huggingface.co/datasets/galgol/quickb-qa.quickb-kbquickb-kb
