CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01astr010 /sec-10k-markdown-uncompressed 📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents) Dataset Summary This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025). The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.texttext-generation1K<n<10K0 likes3.3k downloads2mo agoHugging Face02scilons /SciLaD-all-markdown-v1 SciLaD (Markdown) SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more than 35 million publications. We also publish the extensible pipeline for generating SciLaD. Dataset Details In this repository we share the full… See the full description on the dataset page: https://huggingface.co/datasets/scilons/SciLaD-all-markdown-v1.text10M<n<100M0 likes1.9k downloads1mo agoHugging Face03philschmid /markdown-documentation-transformers Hugging Face Transformers documentation as markdown dataset This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown. This dataset can be used to create RAG applications, which want to use the transformers documentation. Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.textn<1K11 likes943 downloads3y agoHugging Face04marin-community /ar5iv-no-problem-markdown Marin Markdownified Ar5iv Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 2.74B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text. Value Tokens 2 742 463 924 Primary source https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/ File format JSONL License C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-no-problem-markdown.texttext-generation100K<n<1M5 likes817 downloads1y agoHugging Face05marin-community /ar5iv-warning-markdown Marin Markdownified Ar5iv Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 22.34B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text. Value Tokens 19 552 307 274 Primary source https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/ File format JSONL License C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-warning-markdown.texttext-generation1M<n<10M0 likes612 downloads1y agoHugging Face06albertklorer /safedocs-markdown-200k-unifiedtext10K<n<100K0 likes607 downloads1mo agoHugging Face07NAMAA-Space /QariOCR-v0.3-markdown-mixed-dataset QARI Markdown Mixed Dataset 📋 Dataset Summary The QARI v0.3 Markdown Mixed Dataset is a specialized synthetic dataset designed for training Arabic OCR models with a focus on complex document layouts and HTML structure understanding. This dataset is part of the QARI-OCR project, which achieves state-of-the-art performance in Arabic text recognition. This dataset contains 37,000 synthetically generated Arabic document images (29.6k train, 3.7k… See the full description on the dataset page: https://huggingface.co/datasets/NAMAA-Space/QariOCR-v0.3-markdown-mixed-dataset.imagetext-to-image10K<n<100K13 likes540 downloads1y agoHugging Face08marin-community /wikipedia-markdown Marin Markdownified Wikipedia Markdownified Wikipedia is a large-scale, pre-processed version of the English Wikipedia Enterprise HTML dump consisting of 8.59B tokens. The corpus has been converted to clean, section-aware Markdown for language-model training. Value Tokens 8 587 224 558 Primary source https://dumps.wikimedia.org/other/enterprise_html/runs/20241201/enwiki-NS0-20241201-ENTERPRISE-HTML.json.tar.gz File format JSONL License CC-BY-SA 4.0 (mirrors… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/wikipedia-markdown.texttext-generation1M<n<10M7 likes339 downloads1y agoHugging Face09albertklorer /safedocs-markdown SafeDocs PaddleOCR-VL 1.6 production This public corpus contains exactly 2,000,000 eligible document pages from 355,429 documents processed with PaddleOCR-VL 1.6. Eligible PDFs contain 1–39 pages and were rendered at 144 DPI. The published reduction is deterministic and document-safe: it retains 1,600 complete ordered Parquet shards (1,999,469 pages), then 103 whole documents from the next shard (531 pages). No document is split across the selection boundary. The exact public… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-markdown.image100K<n<1M0 likes333 downloads1mo agoHugging Face10GreenNode /GreenNode-Table-Markdown-Retrieval-VN GreenNodeTableMarkdownRetrieval An MTEB dataset Massive Text Embedding Benchmark GreenNodeTable documents Task category t2t Domains Financial, Encyclopaedic, Non-fiction Reference https://huggingface.co/GreenNode Source datasets: GreenNode/GreenNode-Table-Markdown-Retrieval-VN How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("GreenNodeTableMarkdownRetrieval")… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/GreenNode-Table-Markdown-Retrieval-VN.texttext-retrieval100K<n<1M5 likes260 downloads9mo agoHugging Face11Delta-Vector /Tauri-RL-Markdown-System-V2textn<1K0 likes240 downloads1mo agoHugging Face12marin-community /stackexchange-markdown Marin Markdownified StackExchange Markdownified Stack Exchange transforms the Stack Exchange's question-answer pairs into Markdown format consisting of 20.4B tokens. This dataset preserves the content contained in technical discussions while organizing it into a thread format for language model training. Value Tokens 20 413 785 853 Primary source https://archive.org/details/stackexchange File format JSONL License CC (mirrors upstream SE licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/stackexchange-markdown.texttext-generation10M<n<100M6 likes236 downloads1y agoHugging Face13andersonbcdefg /github_issues_markdowntext10K<n<100K4 likes200 downloads3y agoHugging Face14casperhansen /pmc-oa-markdown PubMed Central (PMC) Open Access in Markdown This is a subset filtered for the words: obesity, weight loss, and diabetes. The purpose of this extraction is to further research in Biomedicine + NLP. While there are many biomedical datasets, most of them focus on the clinical space, which is often different from the early developments and discovery. The Markdown format includes figure description, tables, and a YAML header with some metadata like original license. It is filtered to… See the full description on the dataset page: https://huggingface.co/datasets/casperhansen/pmc-oa-markdown.text100K<n<1M8 likes182 downloads1y agoHugging Face15nielsr /datasets-docs-markdown-sectionstextn<1K0 likes156 downloads2y agoHugging Face16presightai /arabic_doc_to_markdown Dataset Card for presightai/arabic_doc_to_markdown This dataset contains OCR image–markdown pairs specifically curated for document structure retrieval and reconstruction tasks,focusing on Arabic and English documents across a variety of domains. Dataset Summary Source of Information:This dataset was created by crawling publicly accessible Arabic and English PDFs from: Official Arabic government document portals Arabic news websites and online magazines Community and… See the full description on the dataset page: https://huggingface.co/datasets/presightai/arabic_doc_to_markdown.imageimage-to-text10K<n<100K5 likes143 downloads1y agoHugging Face17neonforestmist /smolgpt-markdown-stories SmolGPT-Fables Stories A deterministic, text-only corpus of 96,000 original English Markdown stories built for SmolGPT-Fables. Every row is one complete supervised story example with an exact prompt / completion boundary, a requested scene count from one to six, and plain-language conditioning fields. No model, API, browser, or network service was used to create this dataset. Dataset summary 96,000 stories across 96,000 isolated story families 25 genres and all… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/smolgpt-markdown-stories.tabulartext-generation10K<n<100K0 likes140 downloads2mo agoHugging Face18altaidevorg /yargitay-kararlar-markdowntext1M<n<10M1 likes121 downloads2y agoHugging Face19casperhansen /pmc-oa-markdown-qa-results A hard biological benchmark that requires search The goal of a model in drug discovery is to find the right information and make a solid conclusion. This benchmark was synthesized from a total of 6 semantically similar papers, and if a model is able to rediscover these papers on its own through search or retrieval, it will have a good shot at being able to answer the question correctly. The evaluation dataset was adversarially constructed so that DeepSeek R1: could correctly… See the full description on the dataset page: https://huggingface.co/datasets/casperhansen/pmc-oa-markdown-qa-results.textquestion-answering1K<n<10K0 likes118 downloads10mo agoHugging Face20franktheglock /html-to-markdown-10000-pairs HTML to Markdown 10,000 Pairs This dataset contains 10,000 synthetic paired examples for HTML-to-Markdown conversion. Structure pairs/train.jsonl: 9,000 examples pairs/validation.jsonl: 500 examples pairs/test.jsonl: 500 examples manifest.json: generation manifest and split counts Each JSONL row includes: id: example identifier html: source HTML string markdown: target Markdown string metadata: per-example metadata Example {"id":"sample-0000000"… See the full description on the dataset page: https://huggingface.co/datasets/franktheglock/html-to-markdown-10000-pairs.texttranslation10K<n<100K0 likes117 downloads6mo agoHugging Face21rojasdiego /chinese-markdown Chinese Markdown from datasets import load_dataset chinese_markdown = load_dataset("rojas-diego/chinese-markdown", split="train") Dataset({ features: ['code', 'size', 'license'], num_rows: 187258 }) text100K<n<1M4 likes104 downloads3y agoHugging Face22cetusian /markdown-table-qa Markdown Table QA Dataset A synthetic dataset of 11,000 (instruction, input, response) triples (10,000 train + 1,000 validation) for training and evaluating language models on structured table understanding and computational reasoning. What's in it Each sample contains a markdown table paired with a natural language question and a conversational answer: Field Description instruction Natural language question about the table input The markdown table… See the full description on the dataset page: https://huggingface.co/datasets/cetusian/markdown-table-qa.tabular10K<n<100K0 likes102 downloads6mo agoHugging Face23marcodsn /arxiv-markdown Dataset card for arxiv-markdown [!Note] [25/04/2025] Images are now hosted on Cloudflare R2 and referenced in the markdowns as external URLs rather than embedded base64. This significantly reduces storage requirements while maintaining full image access. Images are referenced in the dataset as ![Image](url). Community mirrors are welcome! This dataset contains open-access papers retrieved from arXiv and converted to markdown format using docling; specifically, we use the following… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/arxiv-markdown.text1K<n<10K9 likes94 downloads1y agoHugging Face245CD-AI /Viet-Table-Markdownimage10K<n<100K18 likes91 downloads2y agoHugging Face25Misraj /KITAB_pdf_to_markdown_reviewed KITAB_pdf_to_markdown_reviewed (Corrected KITAB-Bench PDF→Markdown) Short description. A carefully reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic document OCR evaluation. We fixed ground-truth errors (hallucinated text, missing page numbers, omissions of small-font text) and standardized formatting to provide a reliable benchmark for model comparison. TL;DR ✅ Human-verified ground truth for Arabic PDF→Markdown ✅ Removes hallucinations and fills… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/KITAB_pdf_to_markdown_reviewed.imagen<1K4 likes87 downloads1y agoHugging Face26rinabuoy /khmer-doc-markdown-pdf-kh3image1K<n<10K0 likes66 downloads1mo agoHugging Face27casperhansen /pmc-oa-markdown-embeddings PubMed Central (PMC) Open Access in Markdown (With Embeddings) This is a subset of casperhansen/pmc-oa-markdown with embeddings generated. Cost to generate: 7 hours of 8xH100 (about $100). Details: Embeddings are created with Qwen3-Embedding-4B Embeddings have a dimension of 2560 These are full-text embeddings of the entire text A few thousand examples were discarded that were longer than 32k tokens text100K<n<1M4 likes57 downloads1y agoHugging Face28stojchet /python.sft.markdown.python.mbpp.markdowntextn<1K0 likes52 downloads2y agoHugging Face29DBlake-BoxedLogic /Image-2-Markdown Image-to-Markdown 100,000 (page image, OCR markdown) pairs for training/evaluating document OCR and image-to-markdown models. Columns image: a rendered image of a single PDF page (grayscale, longest side 1536px). text: the OCR transcription in olmOCR's markdown dialect (LaTeX for equations, HTML for tables, ![alt](...) tags for figures/diagrams). pdf_relpath, url, page_number, primary_language, is_table, is_diagram: carried-over metadata. How it was… See the full description on the dataset page: https://huggingface.co/datasets/DBlake-BoxedLogic/Image-2-Markdown.imageimage-to-text100K<n<1M0 likes50 downloads4mo agoHugging Face30tom-010 /enwiki-articles-markdown-cleaned-2410tabular1M<n<10M0 likes47 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.