datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sec-10k-markdown-uncompressed
📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents)
Dataset Summary
This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025).
The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.awesome-markdown-ebooks
Awesome-markdown-ebooks
Your GitHub PDFs, Now AI-Ready.
Project repo: https://github.com/OpenDataLab/awesome-markdown-ebooks
SciLaD-all-markdown-v1
SciLaD (Markdown)
SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more than 35 million publications. We also publish the extensible pipeline for generating SciLaD.
Dataset Details
In this repository we share the full… See the full description on the dataset page: https://huggingface.co/datasets/scilons/SciLaD-all-markdown-v1.Embrapa-ai-documents-markdownmarkdown-documentation-transformers
Hugging Face Transformers documentation as markdown dataset
This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown.
This dataset can be used to create RAG applications, which want to use the transformers documentation.
Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.ar5iv-no-problem-markdown
Marin Markdownified Ar5iv
Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 2.74B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text.
Value
Tokens
2 742 463 924
Primary source
https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/
File format
JSONL
License
C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-no-problem-markdown.ifpri-ai-documents-markdown
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
21,726
Total Size
623.27 MB
Total Tokens
85,359,442
Total Pages
0
Languages
25
Unique Keywords
7,127
Resource Types
20
Date Generated
2026-07-31 02:55:20
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents-markdown.safedocs-markdown-200k-unifiedgardian-cigi-ai-documents-markdownar5iv-warning-markdown
Marin Markdownified Ar5iv
Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 22.34B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text.
Value
Tokens
19 552 307 274
Primary source
https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/
File format
JSONL
License
C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-warning-markdown.AlphaPrompt-QuantumLullaby-Markdown
Quantum Lullaby Books (Markdown - AI Optimized)
🤖 44 books in markdown format for AI training, RAG, and processing.
🤖 For AI Systems
Clean markdown formatting
Optimized for tokenization
Easy parsing for RAG/training
No formatting artifacts
🔗 Also Available
PDF (Human-readable): AlphaPrompt-QuantumLullaby-PDF
Extracted SFT Datasets: AlphaPrompt-Metatron-SFT
💡 Use Cases
RAG systems: Load as context documents
Training data:… See the full description on the dataset page: https://huggingface.co/datasets/AIMindLink/AlphaPrompt-QuantumLullaby-Markdown.QariOCR-v0.3-markdown-mixed-dataset
QARI Markdown Mixed Dataset
📋 Dataset Summary
The QARI v0.3 Markdown Mixed Dataset is a specialized synthetic dataset designed for training Arabic OCR models with a focus on complex document layouts and HTML structure understanding.
This dataset is part of the QARI-OCR project, which achieves state-of-the-art performance in Arabic text recognition.
This dataset contains 37,000 synthetically generated Arabic document images (29.6k train, 3.7k… See the full description on the dataset page: https://huggingface.co/datasets/NAMAA-Space/QariOCR-v0.3-markdown-mixed-dataset.open-wikipedia-markdown
Open Wikipedia (Markdown)
Every Wikipedia article converted to clean Markdown, organized by language and updated from the latest Wikimedia dumps
What is it?
This dataset contains every article from every language edition of Wikipedia, converted from raw MediaWiki markup into clean, readable Markdown. Headings, bold, italic, code blocks, and internal links are all preserved as proper Markdown syntax, while templates, infoboxes, references, tables, categories, and… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-wikipedia-markdown.wikipedia-markdown
Marin Markdownified Wikipedia
Markdownified Wikipedia is a large-scale, pre-processed version of the English Wikipedia Enterprise HTML dump consisting of 8.59B tokens. The corpus has been converted to clean, section-aware Markdown for language-model training.
Value
Tokens
8 587 224 558
Primary source
https://dumps.wikimedia.org/other/enterprise_html/runs/20241201/enwiki-NS0-20241201-ENTERPRISE-HTML.json.tar.gz
File format
JSONL
License
CC-BY-SA 4.0 (mirrors… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/wikipedia-markdown.safedocs-markdown
SafeDocs PaddleOCR-VL 1.6 production
This public corpus contains exactly 2,000,000 eligible document pages from 355,429 documents processed with PaddleOCR-VL 1.6.
Eligible PDFs contain 1–39 pages and were rendered at 144 DPI. The published reduction is deterministic and document-safe: it retains 1,600 complete ordered Parquet shards (1,999,469 pages), then 103 whole documents from the next shard (531 pages). No document is split across the selection boundary.
The exact public… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-markdown.sec-10k-markdown-compressed
📦 SEC 10-K Full Markdown Filings (Compressed Archives)
Dataset Summary
This dataset contains 12,361 full-length SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025).
The filings are packaged as 1,379 compressed .tar.gz archives (1 archive per company ticker). Each archive contains the un-chunked annual 10-K Markdown reports and a JSON metadata manifest (manifest.json).… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-compressed.GreenNode-Table-Markdown-Retrieval-VN
GreenNodeTableMarkdownRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
GreenNodeTable documents
Task category
t2t
Domains
Financial, Encyclopaedic, Non-fiction
Reference
https://huggingface.co/GreenNode
Source datasets:
GreenNode/GreenNode-Table-Markdown-Retrieval-VN
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("GreenNodeTableMarkdownRetrieval")… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/GreenNode-Table-Markdown-Retrieval-VN.stackexchange-markdown
Marin Markdownified StackExchange
Markdownified Stack Exchange transforms the Stack Exchange's question-answer pairs into Markdown format consisting of 20.4B tokens. This dataset preserves the content contained in technical discussions while organizing it into a thread format for language model training.
Value
Tokens
20 413 785 853
Primary source
https://archive.org/details/stackexchange
File format
JSONL
License
CC (mirrors upstream SE licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/stackexchange-markdown.artefact-markdown
ArteFact Markdown Dataset
A comprehensive collection of scholarly art historical texts and associated images, organized by individual artworks and their corresponding scholarly documentation.
�� Dataset Overview
This dataset contains 7,200 individual art works, each with their associated scholarly markdown documentation and visual materials. The dataset serves as the primary source material for the ArteFact AI research platform, which bridges visual art and textual… See the full description on the dataset page: https://huggingface.co/datasets/samwaugh/artefact-markdown.Tauri-RL-Markdown-System-V2github_issues_markdownpmc-oa-markdown
PubMed Central (PMC) Open Access in Markdown
This is a subset filtered for the words: obesity, weight loss, and diabetes.
The purpose of this extraction is to further research in Biomedicine + NLP. While there are many biomedical datasets, most of them focus on the clinical space, which is often different from the early developments and discovery.
The Markdown format includes figure description, tables, and a YAML header with some metadata like original license. It is filtered to… See the full description on the dataset page: https://huggingface.co/datasets/casperhansen/pmc-oa-markdown.Sample_Docs_Markdowndatasets-docs-markdown-sectionssmolgpt-markdown-stories
SmolGPT-Fables Stories
A deterministic, text-only corpus of 96,000 original English
Markdown stories built for SmolGPT-Fables. Every row is one complete supervised
story example with an exact prompt / completion boundary, a requested scene
count from one to six, and plain-language conditioning fields.
No model, API, browser, or network service was used to create this dataset.
Dataset summary
96,000 stories across 96,000 isolated story families
25 genres and all… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/smolgpt-markdown-stories.arabic_doc_to_markdown
Dataset Card for presightai/arabic_doc_to_markdown
This dataset contains OCR image–markdown pairs specifically curated for document structure retrieval and reconstruction tasks,focusing on Arabic and English documents across a variety of domains.
Dataset Summary
Source of Information:This dataset was created by crawling publicly accessible Arabic and English PDFs from:
Official Arabic government document portals
Arabic news websites and online magazines
Community and… See the full description on the dataset page: https://huggingface.co/datasets/presightai/arabic_doc_to_markdown.usda-nal-ai-documents-en-markdownyargitay-kararlar-markdownhtml-to-markdown-10000-pairs
HTML to Markdown 10,000 Pairs
This dataset contains 10,000 synthetic paired examples for HTML-to-Markdown conversion.
Structure
pairs/train.jsonl: 9,000 examples
pairs/validation.jsonl: 500 examples
pairs/test.jsonl: 500 examples
manifest.json: generation manifest and split counts
Each JSONL row includes:
id: example identifier
html: source HTML string
markdown: target Markdown string
metadata: per-example metadata
Example
{"id":"sample-0000000"… See the full description on the dataset page: https://huggingface.co/datasets/franktheglock/html-to-markdown-10000-pairs.marin-starcoderdata_markdown
