datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sec-10k-markdown-uncompressed
📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents)
Dataset Summary
This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025).
The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.SciLaD-all-markdown-v1
SciLaD (Markdown)
SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more than 35 million publications. We also publish the extensible pipeline for generating SciLaD.
Dataset Details
In this repository we share the full… See the full description on the dataset page: https://huggingface.co/datasets/scilons/SciLaD-all-markdown-v1.markdown-documentation-transformers
Hugging Face Transformers documentation as markdown dataset
This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown.
This dataset can be used to create RAG applications, which want to use the transformers documentation.
Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.ar5iv-no-problem-markdown
Marin Markdownified Ar5iv
Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 2.74B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text.
Value
Tokens
2 742 463 924
Primary source
https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/
File format
JSONL
License
C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-no-problem-markdown.ar5iv-warning-markdown
Marin Markdownified Ar5iv
Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 22.34B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text.
Value
Tokens
19 552 307 274
Primary source
https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/
File format
JSONL
License
C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-warning-markdown.safedocs-markdown-200k-unifiedQariOCR-v0.3-markdown-mixed-dataset
QARI Markdown Mixed Dataset
📋 Dataset Summary
The QARI v0.3 Markdown Mixed Dataset is a specialized synthetic dataset designed for training Arabic OCR models with a focus on complex document layouts and HTML structure understanding.
This dataset is part of the QARI-OCR project, which achieves state-of-the-art performance in Arabic text recognition.
This dataset contains 37,000 synthetically generated Arabic document images (29.6k train, 3.7k… See the full description on the dataset page: https://huggingface.co/datasets/NAMAA-Space/QariOCR-v0.3-markdown-mixed-dataset.wikipedia-markdown
Marin Markdownified Wikipedia
Markdownified Wikipedia is a large-scale, pre-processed version of the English Wikipedia Enterprise HTML dump consisting of 8.59B tokens. The corpus has been converted to clean, section-aware Markdown for language-model training.
Value
Tokens
8 587 224 558
Primary source
https://dumps.wikimedia.org/other/enterprise_html/runs/20241201/enwiki-NS0-20241201-ENTERPRISE-HTML.json.tar.gz
File format
JSONL
License
CC-BY-SA 4.0 (mirrors… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/wikipedia-markdown.safedocs-markdown
SafeDocs PaddleOCR-VL 1.6 production
This public corpus contains exactly 2,000,000 eligible document pages from 355,429 documents processed with PaddleOCR-VL 1.6.
Eligible PDFs contain 1–39 pages and were rendered at 144 DPI. The published reduction is deterministic and document-safe: it retains 1,600 complete ordered Parquet shards (1,999,469 pages), then 103 whole documents from the next shard (531 pages). No document is split across the selection boundary.
The exact public… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-markdown.GreenNode-Table-Markdown-Retrieval-VN
GreenNodeTableMarkdownRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
GreenNodeTable documents
Task category
t2t
Domains
Financial, Encyclopaedic, Non-fiction
Reference
https://huggingface.co/GreenNode
Source datasets:
GreenNode/GreenNode-Table-Markdown-Retrieval-VN
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("GreenNodeTableMarkdownRetrieval")… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/GreenNode-Table-Markdown-Retrieval-VN.Tauri-RL-Markdown-System-V2stackexchange-markdown
Marin Markdownified StackExchange
Markdownified Stack Exchange transforms the Stack Exchange's question-answer pairs into Markdown format consisting of 20.4B tokens. This dataset preserves the content contained in technical discussions while organizing it into a thread format for language model training.
Value
Tokens
20 413 785 853
Primary source
https://archive.org/details/stackexchange
File format
JSONL
License
CC (mirrors upstream SE licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/stackexchange-markdown.github_issues_markdownpmc-oa-markdown
PubMed Central (PMC) Open Access in Markdown
This is a subset filtered for the words: obesity, weight loss, and diabetes.
The purpose of this extraction is to further research in Biomedicine + NLP. While there are many biomedical datasets, most of them focus on the clinical space, which is often different from the early developments and discovery.
The Markdown format includes figure description, tables, and a YAML header with some metadata like original license. It is filtered to… See the full description on the dataset page: https://huggingface.co/datasets/casperhansen/pmc-oa-markdown.datasets-docs-markdown-sectionsarabic_doc_to_markdown
Dataset Card for presightai/arabic_doc_to_markdown
This dataset contains OCR image–markdown pairs specifically curated for document structure retrieval and reconstruction tasks,focusing on Arabic and English documents across a variety of domains.
Dataset Summary
Source of Information:This dataset was created by crawling publicly accessible Arabic and English PDFs from:
Official Arabic government document portals
Arabic news websites and online magazines
Community and… See the full description on the dataset page: https://huggingface.co/datasets/presightai/arabic_doc_to_markdown.smolgpt-markdown-stories
SmolGPT-Fables Stories
A deterministic, text-only corpus of 96,000 original English
Markdown stories built for SmolGPT-Fables. Every row is one complete supervised
story example with an exact prompt / completion boundary, a requested scene
count from one to six, and plain-language conditioning fields.
No model, API, browser, or network service was used to create this dataset.
Dataset summary
96,000 stories across 96,000 isolated story families
25 genres and all… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/smolgpt-markdown-stories.yargitay-kararlar-markdownpmc-oa-markdown-qa-results
A hard biological benchmark that requires search
The goal of a model in drug discovery is to find the right information and make a solid conclusion. This benchmark was synthesized from a total of 6 semantically similar papers, and if a model is able to rediscover these papers on its own through search or retrieval, it will have a good shot at being able to answer the question correctly.
The evaluation dataset was adversarially constructed so that DeepSeek R1:
could correctly… See the full description on the dataset page: https://huggingface.co/datasets/casperhansen/pmc-oa-markdown-qa-results.html-to-markdown-10000-pairs
HTML to Markdown 10,000 Pairs
This dataset contains 10,000 synthetic paired examples for HTML-to-Markdown conversion.
Structure
pairs/train.jsonl: 9,000 examples
pairs/validation.jsonl: 500 examples
pairs/test.jsonl: 500 examples
manifest.json: generation manifest and split counts
Each JSONL row includes:
id: example identifier
html: source HTML string
markdown: target Markdown string
metadata: per-example metadata
Example
{"id":"sample-0000000"… See the full description on the dataset page: https://huggingface.co/datasets/franktheglock/html-to-markdown-10000-pairs.chinese-markdown
Chinese Markdown
from datasets import load_dataset
chinese_markdown = load_dataset("rojas-diego/chinese-markdown", split="train")
Dataset({
features: ['code', 'size', 'license'],
num_rows: 187258
})
markdown-table-qa
Markdown Table QA Dataset
A synthetic dataset of 11,000 (instruction, input, response) triples (10,000 train + 1,000 validation) for training and evaluating language models on structured table understanding and computational reasoning.
What's in it
Each sample contains a markdown table paired with a natural language question and a conversational answer:
Field
Description
instruction
Natural language question about the table
input
The markdown table… See the full description on the dataset page: https://huggingface.co/datasets/cetusian/markdown-table-qa.arxiv-markdown
Dataset card for arxiv-markdown
[!Note]
[25/04/2025] Images are now hosted on Cloudflare R2 and referenced in the markdowns as external URLs rather than embedded base64. This significantly reduces storage requirements while maintaining full image access. Images are referenced in the dataset as . Community mirrors are welcome!
This dataset contains open-access papers retrieved from arXiv and converted to markdown format using docling; specifically, we use the following… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/arxiv-markdown.Viet-Table-MarkdownKITAB_pdf_to_markdown_reviewed
KITAB_pdf_to_markdown_reviewed (Corrected KITAB-Bench PDF→Markdown)
Short description. A carefully reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic document OCR evaluation. We fixed ground-truth errors (hallucinated text, missing page numbers, omissions of small-font text) and standardized formatting to provide a reliable benchmark for model comparison.
TL;DR
✅ Human-verified ground truth for Arabic PDF→Markdown
✅ Removes hallucinations and fills… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/KITAB_pdf_to_markdown_reviewed.khmer-doc-markdown-pdf-kh3pmc-oa-markdown-embeddings
PubMed Central (PMC) Open Access in Markdown (With Embeddings)
This is a subset of casperhansen/pmc-oa-markdown with embeddings generated.
Cost to generate: 7 hours of 8xH100 (about $100).
Details:
Embeddings are created with Qwen3-Embedding-4B
Embeddings have a dimension of 2560
These are full-text embeddings of the entire text
A few thousand examples were discarded that were longer than 32k tokens
python.sft.markdown.python.mbpp.markdownImage-2-Markdown
Image-to-Markdown
100,000 (page image, OCR markdown) pairs for training/evaluating document OCR
and image-to-markdown models.
Columns
image: a rendered image of a single PDF page (grayscale, longest side 1536px).
text: the OCR transcription in olmOCR's markdown dialect (LaTeX for equations,
HTML for tables,  tags for figures/diagrams).
pdf_relpath, url, page_number, primary_language, is_table, is_diagram:
carried-over metadata.
How it was… See the full description on the dataset page: https://huggingface.co/datasets/DBlake-BoxedLogic/Image-2-Markdown.enwiki-articles-markdown-cleaned-2410
