CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saidsef /tech-docs Technical Documentation Dataset A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices. Dataset Overview This dataset includes documentation across multiple domains: Cloud Platforms: GCP (83 docs), EKS (33 docs) Kubernetes Ecosystem:… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/tech-docs.textquestion-answering1K<n<10K2 likes308 downloads2y agoHugging Face02jimjunior /sunbird_salt_docs COCIS WEB INFO Dataset Summary This dataset contains text chucks scraped from its official website and corresponding websites. The dataset consists of JSON chunks, designed for high-performance streaming and parallel processing. Each chunk represents a discrete unit of data structured for machine learning tasks. By sharding the data into chuck files, this repository supports the datasets library's streaming mode, allowing users to train models without… See the full description on the dataset page: https://huggingface.co/datasets/jimjunior/sunbird_salt_docs.textquestion-answering1K<n<10K0 likes300 downloads3mo agoHugging Face03CycloneDX /cdx-docs Introduction This directory contains numerous knowledge files about CycloneDX and cdxgen in jsonlines chat format. The data is useful for training and fine-tuning (LoRA and QLoRA) LLM models. Data Generation We used Google Gemini 2.0 Flash Experimental via aistudio and used the below prompts to convert official documentation markdown files to the chat format. you are an expert in converting markdown files to plain text jsonlines format based on the my template.… See the full description on the dataset page: https://huggingface.co/datasets/CycloneDX/cdx-docs.textquestion-answeringn<1K0 likes147 downloads1y agoHugging Face04Nymbo /Gradio-Docs Gradio Docs These markdown docs were taken from https://github.com/gradio-app/gradio/tree/main/guides I just wanted to have a copy in the Hub 🤗 textquestion-answeringn<1K1 likes109 downloads2y agoHugging Face05evgenypal /k8s-docs-rag-bench k8s-docs-rag-bench Paper: Analyzing Quality--Latency--Resource Trade-offs in a Technical Documentation RAG Assistant Using LoRA Adaptation (arXiv:2605.28222) Code: github.com/EugPal/rag-lora-tradeoffs A small, fully-grounded benchmark for retrieval-augmented question answering (RAG) over the official Kubernetes documentation, together with the full set of LLM-judge labels used in the accompanying preprint "Analyzing Quality-Latency-Resource Trade-offs in a Technical… See the full description on the dataset page: https://huggingface.co/datasets/evgenypal/k8s-docs-rag-bench.tabularquestion-answering100K<n<1M0 likes72 downloads4mo agoHugging Face06mtpti5iD /redhat-docs_dataset 🖥️ Red Hat Technical Documentation Dataset 📌 Overview This dataset contains 55,741 structured technical documentation entries sourced from Red Hat, covering:✅ System Administration Guides – User management, permissions, kernel tuning✅ Networking & Security – Firewall rules, SELinux, VPN setup✅ Virtualization & Containers – KVM, Podman, OpenShift, Kubernetes✅ Enterprise Software Documentation – RHEL, Ansible, Satellite, OpenStack 📊 Dataset Details This… See the full description on the dataset page: https://huggingface.co/datasets/mtpti5iD/redhat-docs_dataset.texttext-retrieval10K<n<100K1 likes65 downloads2y agoHugging Face07louisbrulenaudet /lemone-docs-embedded Lemone-embedded, pre-built embeddings dataset for French taxation. This database presents the embeddings generated by the Lemone-embed-pro model and aims at a large-scale distribution of the model even for the GPU-poor. This sentence transformers model, specifically designed for French taxation, has been fine-tuned on a dataset comprising 43 million tokens, integrating a blend of semi-synthetic and fully synthetic data generated by GPT-4 Turbo and Llama 3.1 70B, which have… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/lemone-docs-embedded.textquestion-answering10K<n<100K3 likes48 downloads2y agoHugging Face08MongoDB /mongodb-docs Overview This dataset consists of a small subset of MongoDB's technical documentation. Dataset Structure The dataset consists of the following fields: sourceName: The source of the document. url: Link to the article. action: Action taken on the article. body: Content of the article in Markdown format. format: Format of the content. metadata: Metadata such as tags, content type etc. associated with the document. title: Title of the document. updated: The last updated… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/mongodb-docs.textquestion-answeringn<1K1 likes37 downloads2y agoHugging Face09MongoDB /mongodb-docs-embedded Overview This dataset consists of chunked and embedded versions of a small subset of MongoDB's technical documentation. Dataset Structure The dataset consists of the following fields: sourceName: The source of the document. url: Link to the article. action: Action taken on the article. body: Content of the article in Markdown format. format: Format of the content. metadata: Metadata such as tags, content type etc. associated with the document. title: Title of the… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/mongodb-docs-embedded.textquestion-answeringn<1K0 likes30 downloads2y agoHugging Face10farid678 /go-effective-docs-qa Effective Go Instruction Dataset Overview This dataset was created from the official Effective Go documentation. The content was extracted from: https://go.dev/doc/effective_go and transformed into instruction-following samples consisting of: instruction input output Dataset Structure Split Examples Train 2,643 Validation 293 Features instruction (string) input (string) output (string) How this… See the full description on the dataset page: https://huggingface.co/datasets/farid678/go-effective-docs-qa.texttext-generation1K<n<10K0 likes30 downloads2mo agoHugging Face11nhn309261 /vietnamese-legal-docs Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 2,393 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/nhn309261/vietnamese-legal-docs.texttext-classification1M<n<10M0 likes26 downloads6mo agoHugging Face12farabi-lab /KZ-RAG-single-docs-final-goldgated 🇰🇿 Kazakh Analytical RAG and Document-Based QA 📖 Overview This dataset is a high-density collection of 4,522 analytical samples designed for Retrieval-Augmented Generation (RAG) tasks in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 4,522 Total Words (approx.) 5,978,950 Avg. Words per Sample 1,322 Word Count Distribution (Per Field) The dataset features… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/KZ-RAG-single-docs-final-gold.textquestion-answering1K<n<10K0 likes21 downloads2mo agoHugging Face13alucent /mirror-tech-docsgated Technical Documentation Dataset A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices. Dataset Overview This dataset includes documentation across multiple domains: Cloud Platforms: GCP (83 docs), EKS (33 docs) Kubernetes… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-tech-docs.textquestion-answering1K<n<10K0 likes21 downloads2mo agoHugging Face14Wauplin /hugr-docs-qa Hugr Docs QA Dataset This dataset contains a collection of question-and-answer pairs derived from the Hugr documentation. It is designed to evaluate the performance of documentation assistants and RAG (Retrieval-Augmented Generation) systems. Summary The dataset consists of 10 QA pairs categorized by difficulty (basic, intermediate, advanced), covering core concepts such as subagent composition, crate scaffolding, and the runtime resume mechanism.… See the full description on the dataset page: https://huggingface.co/datasets/Wauplin/hugr-docs-qa.textquestion-answeringn<1K0 likes12 downloads3mo agoHugging Face15jokernifty /docs-instruct-unity-20260601-2007 docs-instruct-unity-20260601-2007 Synthetic instruction-tuning dataset generated by the DownFTuner pipeline. Source: documentation site (unity), crawled and chunked. Generator: DeepSeek-V4-Pro via Ollama, with EmbeddingGemma dedup + hallucination filter. Format: chat-format JSONL (messages field), split into train.jsonl and valid.jsonl. Source URLs are preserved in each row's source metadata where available. texttext-generation1K<n<10K0 likes10 downloads4mo agoHugging Face16jokernifty /docs-instruct-nextjs-20260601-0306 docs-instruct-20260601-0306 Synthetic instruction-tuning dataset generated by the DownFTuner pipeline. Source: random Wikipedia articles (en), one run. Generator: LLM-synthesized instruction/answer pairs grounded in each article. Format: chat-format JSONL (messages field), split into train.jsonl and valid.jsonl. License: CC-BY-SA-4.0 (inherits from Wikipedia source). Source URLs are preserved in each row's source metadata. texttext-generation1K<n<10K0 likes7 downloads4mo agoHugging Face17iloncka /qa_program_modules_docsgatedtextquestion-answering10K<n<100K0 likes3 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.