datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
starcoder2-documentation
Dataset Card
This dataset is the code documenation dataset used in StarCoder2 pre-training, and it is also part of the-stack-v2-train-extras descried in the paper.
Dataset Details
Overview
This dataset comprises a comprehensive collection of crawled documentation and code-related resources sourced from various package manager platforms and programming language documentation sites. It focuses on popular libraries, free programming books, and other relevant… See the full description on the dataset page: https://huggingface.co/datasets/SivilTaram/starcoder2-documentation.aws_bedrock_documentation_demo
Aws Bedrock Documentation Demo
This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections.
Pipeline Steps
ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps
summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction
chunking: Split texts into token-based single-hop and multi-hop… See the full description on the dataset page: https://huggingface.co/datasets/yourbench/aws_bedrock_documentation_demo.software-documentation-zsm-bitextmining
software-documentation-zsm-bitextmining
Deduplicated copy of kornwtp/software-documentation-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/software-documentation-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-zsm-bitextmining.aws_bedrock_documentation_demo
Aws Bedrock Documentation Demo
This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections.
Pipeline Steps
ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps
summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction
chunking: Split texts into token-based single-hop and multi-hop… See the full description on the dataset page: https://huggingface.co/datasets/sg-c/aws_bedrock_documentation_demo.software-documentation-tha-bitextmining
software-documentation-tha-bitextmining
Deduplicated copy of kornwtp/software-documentation-tha-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/software-documentation-tha-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-tha-bitextmining.software-documentation-vie-bitextmining
software-documentation-vie-bitextmining
Deduplicated copy of kornwtp/software-documentation-vie-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/software-documentation-vie-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-vie-bitextmining.software-documentation-ind-bitextmining
software-documentation-ind-bitextmining
Deduplicated copy of kornwtp/software-documentation-ind-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/software-documentation-ind-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-ind-bitextmining.kubernetes-documentation-dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This dataset consists of the Kubernetes data that has been scraped from the web(https://kubernetes.io/docs/concepts/services-networking/)
License: [MIT]
Dataset Sources [optional]
Repository: [https://github.com/keethu12345/Kubernetes_ML-Model]
Uses
This… See the full description on the dataset page: https://huggingface.co/datasets/keethu/kubernetes-documentation-dataset.documentation-kubernetes
Documentation-Kubernetes
Made with ❤️ using 🦥 Unsloth Studio
kubernetes documentation was generated with Unsloth Recipe Studio. It contains 99 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("makayel/documentation-kubernetes", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 99
📋 Columns: 3
✅ Completion: 99.0% (100 requested)
📋 Schema & Statistics… See the full description on the dataset page: https://huggingface.co/datasets/makayel/documentation-kubernetes.pandas-documentation
Dataset Card for "pandas-documentation"
More Information needed
pandas_documentation3 datasets:
Web scraped pandas documentation, where each instance is a code example generated by gpt-3.5-turbo based on the examples from documentation. Each instance is one pandas method, type, class, etc.
DS1000 is DS-1000 samples that contain pandas code
OSS-Instruct is Magicoder's dataset where pandas occur.
Filtering and scraping is available here.
software-documentation-tha-bitextminingsoftware-documentation-vie-bitextminingmulesoft-documentation-embeddings
mulesoft-documentation-embeddings
MuleSoft Documentation Embeddings for RAG Applications
Dataset Information
Version: 1.0.0
Created: 2025-09-16T02:41:16.352809
Source: Vector Database
License: MIT
Language: en
Task Categories
question-answering, retrieval, knowledge-base
Dataset Statistics
SkillPilotDataSet_v11
Total Objects: 6430
Unique Properties: 13
Knowledge Sources: mulesoft, user_defined_docs
Average Content Length: 5079… See the full description on the dataset page: https://huggingface.co/datasets/BassemE/mulesoft-documentation-embeddings.software-documentation-zsm-bitextminingsoftware-documentation-ind-bitextminingGroovy_documentation_QA
Dataset Card for Groovy_documentation_QA
A dataset consisting of 2900+ question/answer pairs generated from the Apche groovy documentation
Dataset Details
Each row in the dataset consists of the following features:
topic: 2-3 word description of the topic
question: A question about the Groovy Programming language
answer: The answer to the question
ebf-onboarder-documentation
Dataset Card for ebf-onboarder-documentation
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/MichaelPrimez/ebf-onboarder-documentation/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/MichaelPrimez/ebf-onboarder-documentation.documentation_datasetsynthetic-documentationsmodel-documentation-scoreboardprocess-documentationservices-de-documentation-et-sieges-des-bibliotheques-de-lenseignement-superieur
Services de documentation et sièges des bibliothèques de l'Enseignement supérieur
Source
Source officielle : https://www.data.gouv.fr/datasets/services-de-documentation-et-sieges-des-bibliotheques-de-lenseignement-superieur
Identifiant du jeu de données data.gouv.fr : 5b120868b5950870b30303f0
Slug data.gouv.fr : services-de-documentation-et-sieges-des-bibliotheques-de-lenseignement-superieur
Licence indiquée dans les métadonnées data.gouv.fr : lov2… See the full description on the dataset page: https://huggingface.co/datasets/Data-Gouv-ML/services-de-documentation-et-sieges-des-bibliotheques-de-lenseignement-superieur.
