CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SivilTaram /starcoder2-documentation Dataset Card This dataset is the code documenation dataset used in StarCoder2 pre-training, and it is also part of the-stack-v2-train-extras descried in the paper. Dataset Details Overview This dataset comprises a comprehensive collection of crawled documentation and code-related resources sourced from various package manager platforms and programming language documentation sites. It focuses on popular libraries, free programming books, and other relevant… See the full description on the dataset page: https://huggingface.co/datasets/SivilTaram/starcoder2-documentation.text10K<n<100K10 likes219 downloads2y agoHugging Face02yourbench /aws_bedrock_documentation_demo Aws Bedrock Documentation Demo This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections. Pipeline Steps ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction chunking: Split texts into token-based single-hop and multi-hop… See the full description on the dataset page: https://huggingface.co/datasets/yourbench/aws_bedrock_documentation_demo.tabular1K<n<10K0 likes103 downloads1y agoHugging Face03puttatidam /software-documentation-zsm-bitextmining software-documentation-zsm-bitextmining Deduplicated copy of kornwtp/software-documentation-zsm-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/software-documentation-zsm-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: train What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-zsm-bitextmining.text1K<n<10K0 likes95 downloads8d agoHugging Face04sg-c /aws_bedrock_documentation_demo Aws Bedrock Documentation Demo This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections. Pipeline Steps ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction chunking: Split texts into token-based single-hop and multi-hop… See the full description on the dataset page: https://huggingface.co/datasets/sg-c/aws_bedrock_documentation_demo.textn<1K0 likes90 downloads10mo agoHugging Face05puttatidam /software-documentation-tha-bitextmining software-documentation-tha-bitextmining Deduplicated copy of kornwtp/software-documentation-tha-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/software-documentation-tha-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: train What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-tha-bitextmining.text1K<n<10K0 likes88 downloads8d agoHugging Face06puttatidam /software-documentation-vie-bitextmining software-documentation-vie-bitextmining Deduplicated copy of kornwtp/software-documentation-vie-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/software-documentation-vie-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: train What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-vie-bitextmining.text1K<n<10K0 likes87 downloads8d agoHugging Face07puttatidam /software-documentation-ind-bitextmining software-documentation-ind-bitextmining Deduplicated copy of kornwtp/software-documentation-ind-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/software-documentation-ind-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: train What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-ind-bitextmining.text1K<n<10K0 likes82 downloads8d agoHugging Face08keethu /kubernetes-documentation-dataset Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description This dataset consists of the Kubernetes data that has been scraped from the web(https://kubernetes.io/docs/concepts/services-networking/) License: [MIT] Dataset Sources [optional] Repository: [https://github.com/keethu12345/Kubernetes_ML-Model] Uses This… See the full description on the dataset page: https://huggingface.co/datasets/keethu/kubernetes-documentation-dataset.texttext-generationn<1K0 likes58 downloads2y agoHugging Face09makayel /documentation-kubernetes Documentation-Kubernetes Made with ❤️ using 🦥 Unsloth Studio kubernetes documentation was generated with Unsloth Recipe Studio. It contains 99 generated records. 🚀 Quick Start from datasets import load_dataset # Load the main dataset dataset = load_dataset("makayel/documentation-kubernetes", "data", split="train") df = dataset.to_pandas() 📊 Dataset Summary 📈 Records: 99 📋 Columns: 3 ✅ Completion: 99.0% (100 requested) 📋 Schema & Statistics… See the full description on the dataset page: https://huggingface.co/datasets/makayel/documentation-kubernetes.textn<1K0 likes52 downloads6mo agoHugging Face10pacovaldez /pandas-documentation Dataset Card for "pandas-documentation" More Information needed text1K<n<10K0 likes43 downloads3y agoHugging Face11poludmik /pandas_documentation3 datasets: Web scraped pandas documentation, where each instance is a code example generated by gpt-3.5-turbo based on the examples from documentation. Each instance is one pandas method, type, class, etc. DS1000 is DS-1000 samples that contain pandas code OSS-Instruct is Magicoder's dataset where pandas occur. Filtering and scraping is available here. text1K<n<10K1 likes41 downloads2y agoHugging Face12kornwtp /software-documentation-tha-bitextminingtext1K<n<10K0 likes40 downloads2y agoHugging Face13kornwtp /software-documentation-vie-bitextminingtext1K<n<10K0 likes32 downloads2y agoHugging Face14BassemE /mulesoft-documentation-embeddings mulesoft-documentation-embeddings MuleSoft Documentation Embeddings for RAG Applications Dataset Information Version: 1.0.0 Created: 2025-09-16T02:41:16.352809 Source: Vector Database License: MIT Language: en Task Categories question-answering, retrieval, knowledge-base Dataset Statistics SkillPilotDataSet_v11 Total Objects: 6430 Unique Properties: 13 Knowledge Sources: mulesoft, user_defined_docs Average Content Length: 5079… See the full description on the dataset page: https://huggingface.co/datasets/BassemE/mulesoft-documentation-embeddings.tabularquestion-answering1K<n<10K0 likes29 downloads1y agoHugging Face15kornwtp /software-documentation-zsm-bitextminingtext1K<n<10K0 likes27 downloads2y agoHugging Face16kornwtp /software-documentation-ind-bitextminingtext1K<n<10K0 likes24 downloads2y agoHugging Face17DarioArena87 /Groovy_documentation_QA Dataset Card for Groovy_documentation_QA A dataset consisting of 2900+ question/answer pairs generated from the Apche groovy documentation Dataset Details Each row in the dataset consists of the following features: topic: 2-3 word description of the topic question: A question about the Groovy Programming language answer: The answer to the question textquestion-answering1K<n<10K0 likes15 downloads8mo agoHugging Face18MichaelPrimez /ebf-onboarder-documentation Dataset Card for ebf-onboarder-documentation This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/MichaelPrimez/ebf-onboarder-documentation/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/MichaelPrimez/ebf-onboarder-documentation.texttext-generationn<1K0 likes10 downloads1y agoHugging Face19trackio /documentation_datasettabularn<1K0 likes8 downloads9mo agoHugging Face206StringNinja /synthetic-documentationstextn<1K0 likes7 downloads1y agoHugging Face21remyxai /model-documentation-scoreboardtabular1K<n<10K1 likes6 downloads1y agoHugging Face22rainiernate /process-documentationtextn<1K0 likes4 downloads2y agoHugging Face23Data-Gouv-ML /services-de-documentation-et-sieges-des-bibliotheques-de-lenseignement-superieur Services de documentation et sièges des bibliothèques de l'Enseignement supérieur Source Source officielle : https://www.data.gouv.fr/datasets/services-de-documentation-et-sieges-des-bibliotheques-de-lenseignement-superieur Identifiant du jeu de données data.gouv.fr : 5b120868b5950870b30303f0 Slug data.gouv.fr : services-de-documentation-et-sieges-des-bibliotheques-de-lenseignement-superieur Licence indiquée dans les métadonnées data.gouv.fr : lov2… See the full description on the dataset page: https://huggingface.co/datasets/Data-Gouv-ML/services-de-documentation-et-sieges-des-bibliotheques-de-lenseignement-superieur.textn<1K0 likes4 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.