datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mulesoft-documentation-embeddings
mulesoft-documentation-embeddings
MuleSoft Documentation Embeddings for RAG Applications
Dataset Information
Version: 1.0.0
Created: 2025-09-16T02:41:16.352809
Source: Vector Database
License: MIT
Language: en
Task Categories
question-answering, retrieval, knowledge-base
Dataset Statistics
SkillPilotDataSet_v11
Total Objects: 6430
Unique Properties: 13
Knowledge Sources: mulesoft, user_defined_docs
Average Content Length: 5079… See the full description on the dataset page: https://huggingface.co/datasets/BassemE/mulesoft-documentation-embeddings.Developer-Documentation-QAsugarcrm_130_documentation
Source: Sugarcrm 13.0 Dev Documentation
The chunks in the files are diffrent splittet based on the tokenizer conained in the name of the file
cl100k_base: 400 Tokens per chunk
p50k_base: 200 Tokens per chunk
Groovy_documentation_QA
Dataset Card for Groovy_documentation_QA
A dataset consisting of 2900+ question/answer pairs generated from the Apche groovy documentation
Dataset Details
Each row in the dataset consists of the following features:
topic: 2-3 word description of the topic
question: A question about the Groovy Programming language
answer: The answer to the question
ebf-onboarder-documentation
Dataset Card for ebf-onboarder-documentation
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/MichaelPrimez/ebf-onboarder-documentation/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/MichaelPrimez/ebf-onboarder-documentation.materials-project-documentationUser-Documentation-QA
Keboola QA Pairs
Dataset SummaryA concise set of Q&A pairs derived from Keboola user documentation, each question and answer is self-contained and domain-specific. Suitable for training or evaluation of QA and retrieval-augmented generation models.
Metadata
Dataset Name: keboola-qa-pairs
Language: English (en)
Task Categories: Question-Answering, Retrieval-Augmented Generation
Size: ~3000
Source: Keboola documentation excerpts
Sample Data
{
"question": "How… See the full description on the dataset page: https://huggingface.co/datasets/Keboola/User-Documentation-QA.
