datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
markdown-documentation-transformers
Hugging Face Transformers documentation as markdown dataset
This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown.
This dataset can be used to create RAG applications, which want to use the transformers documentation.
Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.ar5iv-no-problem-markdown
Marin Markdownified Ar5iv
Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 2.74B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text.
Value
Tokens
2 742 463 924
Primary source
https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/
File format
JSONL
License
C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-no-problem-markdown.ar5iv-warning-markdown
Marin Markdownified Ar5iv
Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 22.34B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text.
Value
Tokens
19 552 307 274
Primary source
https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/
File format
JSONL
License
C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-warning-markdown.wikipedia-markdown
Marin Markdownified Wikipedia
Markdownified Wikipedia is a large-scale, pre-processed version of the English Wikipedia Enterprise HTML dump consisting of 8.59B tokens. The corpus has been converted to clean, section-aware Markdown for language-model training.
Value
Tokens
8 587 224 558
Primary source
https://dumps.wikimedia.org/other/enterprise_html/runs/20241201/enwiki-NS0-20241201-ENTERPRISE-HTML.json.tar.gz
File format
JSONL
License
CC-BY-SA 4.0 (mirrors… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/wikipedia-markdown.Tauri-RL-Markdown-System-V2stackexchange-markdown
Marin Markdownified StackExchange
Markdownified Stack Exchange transforms the Stack Exchange's question-answer pairs into Markdown format consisting of 20.4B tokens. This dataset preserves the content contained in technical discussions while organizing it into a thread format for language model training.
Value
Tokens
20 413 785 853
Primary source
https://archive.org/details/stackexchange
File format
JSONL
License
CC (mirrors upstream SE licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/stackexchange-markdown.yargitay-kararlar-markdownpmc-oa-markdown-qa-results
A hard biological benchmark that requires search
The goal of a model in drug discovery is to find the right information and make a solid conclusion. This benchmark was synthesized from a total of 6 semantically similar papers, and if a model is able to rediscover these papers on its own through search or retrieval, it will have a good shot at being able to answer the question correctly.
The evaluation dataset was adversarially constructed so that DeepSeek R1:
could correctly… See the full description on the dataset page: https://huggingface.co/datasets/casperhansen/pmc-oa-markdown-qa-results.html-to-markdown-10000-pairs
HTML to Markdown 10,000 Pairs
This dataset contains 10,000 synthetic paired examples for HTML-to-Markdown conversion.
Structure
pairs/train.jsonl: 9,000 examples
pairs/validation.jsonl: 500 examples
pairs/test.jsonl: 500 examples
manifest.json: generation manifest and split counts
Each JSONL row includes:
id: example identifier
html: source HTML string
markdown: target Markdown string
metadata: per-example metadata
Example
{"id":"sample-0000000"… See the full description on the dataset page: https://huggingface.co/datasets/franktheglock/html-to-markdown-10000-pairs.chinese-markdown
Chinese Markdown
from datasets import load_dataset
chinese_markdown = load_dataset("rojas-diego/chinese-markdown", split="train")
Dataset({
features: ['code', 'size', 'license'],
num_rows: 187258
})
ACL-23-Paper-OCR-Markdown
ACL 2023 Paper in Markdown after OCR
This dataset contains 2150 papers from Association for Computational Linguistics (ACL) 2023:
Long Papers (912 papers)
Short Papers (185 papers)
System Demonstrations (59 paper)
Student Research Workshop (35 papers)
Industry Track (77 papers)
Tutorial Abstracts (7 papers)
Findings (902 papers)
This dataset is processed and compiled by @hu_yifei as part of open-source effort from the Open Research Assistant Project.
OCR process
The… See the full description on the dataset page: https://huggingface.co/datasets/yifeihu/ACL-23-Paper-OCR-Markdown.telelogs_markdown
TeleLogs Dataset (Processed MCQ Format)
This dataset has been extracted from the original netop/TeleLogs dataset and processed into multiple-choice question (MCQ) format for easier evaluation.
Dataset Description
TeleLogs is a telecommunications log analysis benchmark where models must identify the root cause of network issues from 5G wireless network drive-test data and engineering parameters.
Processed Format
This version has been restructured for MCQ… See the full description on the dataset page: https://huggingface.co/datasets/emolero/telelogs_markdown.bank_statement_markdown_to_csvTauri-RL-Markdown-Systemjunos-user-manual-markdown-raw
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Coldyuja/junos-user-manual-markdown-raw.markdownCreatorDatasetv2Markdown_BIBLESmarkdownCreatorDatasetprolewiki-library-markdownified
