datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tech-docs
Technical Documentation Dataset
A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices.
Dataset Overview
This dataset includes documentation across multiple domains:
Cloud Platforms: GCP (83 docs), EKS (33 docs)
Kubernetes Ecosystem:… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/tech-docs.sunbird_salt_docs
COCIS WEB INFO
Dataset Summary
This dataset contains text chucks scraped from its official website and corresponding websites.
The dataset consists of JSON chunks, designed for high-performance streaming and parallel processing. Each chunk represents a discrete unit of data structured for machine learning tasks.
By sharding the data into chuck files, this repository supports the datasets library's streaming mode, allowing users to train models without… See the full description on the dataset page: https://huggingface.co/datasets/jimjunior/sunbird_salt_docs.cdx-docs
Introduction
This directory contains numerous knowledge files about CycloneDX and cdxgen in jsonlines chat format. The data is useful for training and fine-tuning (LoRA and QLoRA) LLM models.
Data Generation
We used Google Gemini 2.0 Flash Experimental via aistudio and used the below prompts to convert official documentation markdown files to the chat format.
you are an expert in converting markdown files to plain text jsonlines format based on the my template.… See the full description on the dataset page: https://huggingface.co/datasets/CycloneDX/cdx-docs.Gradio-Docs
Gradio Docs
These markdown docs were taken from https://github.com/gradio-app/gradio/tree/main/guides
I just wanted to have a copy in the Hub 🤗
k8s-docs-rag-bench
k8s-docs-rag-bench
Paper: Analyzing Quality--Latency--Resource Trade-offs in a Technical Documentation RAG Assistant Using LoRA Adaptation (arXiv:2605.28222)
Code: github.com/EugPal/rag-lora-tradeoffs
A small, fully-grounded benchmark for retrieval-augmented question answering
(RAG) over the official Kubernetes documentation, together with the full
set of LLM-judge labels used in the accompanying preprint
"Analyzing Quality-Latency-Resource Trade-offs in a Technical… See the full description on the dataset page: https://huggingface.co/datasets/evgenypal/k8s-docs-rag-bench.redhat-docs_dataset
🖥️ Red Hat Technical Documentation Dataset
📌 Overview
This dataset contains 55,741 structured technical documentation entries sourced from Red Hat, covering:✅ System Administration Guides – User management, permissions, kernel tuning✅ Networking & Security – Firewall rules, SELinux, VPN setup✅ Virtualization & Containers – KVM, Podman, OpenShift, Kubernetes✅ Enterprise Software Documentation – RHEL, Ansible, Satellite, OpenStack
📊 Dataset Details
This… See the full description on the dataset page: https://huggingface.co/datasets/mtpti5iD/redhat-docs_dataset.lemone-docs-embedded
Lemone-embedded, pre-built embeddings dataset for French taxation.
This database presents the embeddings generated by the Lemone-embed-pro model and aims at a large-scale distribution of the model even for the GPU-poor.
This sentence transformers model, specifically designed for French taxation, has been fine-tuned on a dataset comprising 43 million tokens, integrating a blend of semi-synthetic and fully synthetic data generated by GPT-4 Turbo and Llama 3.1 70B, which have… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/lemone-docs-embedded.mongodb-docs
Overview
This dataset consists of a small subset of MongoDB's technical documentation.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the document.
url: Link to the article.
action: Action taken on the article.
body: Content of the article in Markdown format.
format: Format of the content.
metadata: Metadata such as tags, content type etc. associated with the document.
title: Title of the document.
updated: The last updated… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/mongodb-docs.mongodb-docs-embedded
Overview
This dataset consists of chunked and embedded versions of a small subset of MongoDB's technical documentation.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the document.
url: Link to the article.
action: Action taken on the article.
body: Content of the article in Markdown format.
format: Format of the content.
metadata: Metadata such as tags, content type etc. associated with the document.
title: Title of the… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/mongodb-docs-embedded.go-effective-docs-qa
Effective Go Instruction Dataset
Overview
This dataset was created from the official Effective Go documentation. The content was extracted from:
https://go.dev/doc/effective_go
and transformed into instruction-following samples consisting of:
instruction
input
output
Dataset Structure
Split
Examples
Train
2,643
Validation
293
Features
instruction (string)
input (string)
output (string)
How this… See the full description on the dataset page: https://huggingface.co/datasets/farid678/go-effective-docs-qa.vietnamese-legal-docs
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
2,393 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/nhn309261/vietnamese-legal-docs.KZ-RAG-single-docs-final-gold
🇰🇿 Kazakh Analytical RAG and Document-Based QA
📖 Overview
This dataset is a high-density collection of 4,522 analytical samples designed for Retrieval-Augmented Generation (RAG) tasks in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
4,522
Total Words (approx.)
5,978,950
Avg. Words per Sample
1,322
Word Count Distribution (Per Field)
The dataset features… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/KZ-RAG-single-docs-final-gold.mirror-tech-docs
Technical Documentation Dataset
A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices.
Dataset Overview
This dataset includes documentation across multiple domains:
Cloud Platforms: GCP (83 docs), EKS (33 docs)
Kubernetes… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-tech-docs.hugr-docs-qa
Hugr Docs QA Dataset
This dataset contains a collection of question-and-answer pairs derived from the Hugr documentation. It is designed to evaluate the performance of documentation assistants and RAG (Retrieval-Augmented Generation) systems.
Summary
The dataset consists of 10 QA pairs categorized by difficulty (basic, intermediate, advanced), covering core concepts such as subagent composition, crate scaffolding, and the runtime resume mechanism.… See the full description on the dataset page: https://huggingface.co/datasets/Wauplin/hugr-docs-qa.docs-instruct-unity-20260601-2007
docs-instruct-unity-20260601-2007
Synthetic instruction-tuning dataset generated by the DownFTuner pipeline.
Source: documentation site (unity), crawled and chunked.
Generator: DeepSeek-V4-Pro via Ollama, with EmbeddingGemma dedup + hallucination filter.
Format: chat-format JSONL (messages field), split into train.jsonl and valid.jsonl.
Source URLs are preserved in each row's source metadata where available.
docs-instruct-nextjs-20260601-0306
docs-instruct-20260601-0306
Synthetic instruction-tuning dataset generated by the DownFTuner pipeline.
Source: random Wikipedia articles (en), one run.
Generator: LLM-synthesized instruction/answer pairs grounded in each article.
Format: chat-format JSONL (messages field), split into train.jsonl and valid.jsonl.
License: CC-BY-SA-4.0 (inherits from Wikipedia source).
Source URLs are preserved in each row's source metadata.
qa_program_modules_docs
