datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
markdown-documentation-transformers
Hugging Face Transformers documentation as markdown dataset
This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown.
This dataset can be used to create RAG applications, which want to use the transformers documentation.
Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.python_documentation_codeDataset created using https://github.com/jeffmeloy/py2dataset using the Python code from the following:
https://github.com/ansible/ansible
https://github.com/apache/airflow
https://github.com/arogozhnikov/einops
https://github.com/arviz-devs/arviz
https://github.com/astropy/astropy
https://github.com/biopython/biopython
https://github.com/bjodah/chempy
https://github.com/bokeh/bokehhttps://github.com/CalebBell/thermo
https://github.com/camDavidsonPilon/lifelines
https://github.com/coin-or/pulp… See the full description on the dataset page: https://huggingface.co/datasets/jeffmeloy/python_documentation_code.library-documentationThe library documentation retrieval source for code-rag-bench, contains all documentation for Python libraries available on devdocs.io.
aws-documentation-chunkedsugarcrm_130_documentation
Source: Sugarcrm 13.0 Dev Documentation
The chunks in the files are diffrent splittet based on the tokenizer conained in the name of the file
cl100k_base: 400 Tokens per chunk
p50k_base: 200 Tokens per chunk
mz94-documentation
MZ94 - LLM Training Dataset
Overview
This dataset contains crawled documentation from https://infozone.atlassian.net/wiki/spaces/MD94/, formatted for LLM training and RAG systems.
Dataset Statistics
Total Pages: 4109
Total Words: 1005892
Total Chunks: 2420
Crawled: 2025-06-24 04:55:10
Directory Structure
/llm_ready/
Plain text files optimized for LLM training:
Clean, formatted text content
Consistent structure with headers
Document… See the full description on the dataset page: https://huggingface.co/datasets/ratanon/mz94-documentation.langchain-documentationmz93-documentation
MZ93 - LLM Training Dataset
Overview
This dataset contains crawled documentation from https://infozone.atlassian.net/wiki/spaces/MD93/, formatted for LLM training and RAG systems.
Dataset Statistics
Total Pages: 3722
Total Words: 943002
Total Chunks: 2276
Crawled: 2025-06-24 05:02:33
Directory Structure
/llm_ready/
Plain text files optimized for LLM training:
Clean, formatted text content
Consistent structure with headers
Document… See the full description on the dataset page: https://huggingface.co/datasets/ratanon/mz93-documentation.seo_documentation_ymanim_community_and_documentation_code
Dataset Description: Manim Code Examples
Overview:This dataset consists of 210 code examples demonstrating the usage of Manim, a powerful mathematical animation engine for creating precise and visually appealing animations. Each example is structured as a JSON object containing three key components: a prompt, a response, and metadata. The dataset serves as a comprehensive resource for learning, teaching, or experimenting with Manim's capabilities in animating mathematical concepts… See the full description on the dataset page: https://huggingface.co/datasets/shekhar98/manim_community_and_documentation_code.powershell-documentation-datasetprocedures_documentationdocumentationcomplex_code_documentation_datasetdocumentationa_datasetDocumentationSO-documentationseo_documentation_gaveva_public_documentation_01EDIT:
i'm letting this dataset public because it is technically valid so you can use it as a test if you want, but know that its quality is very poor, the models tuned with it have very poor performances.
i created another dataset from PML coding only (public) documents, so if you want a better quality dataset, you can try this one: YagCed/Aveva_PML_test
A 13000 lines dataset created from a few public PDF documents for PML and Aveva products.
This is a test i made to learn how to create a… See the full description on the dataset page: https://huggingface.co/datasets/YagCed/aveva_public_documentation_01.ndn-cxx-documentation
