CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01astr010 /sec-10k-markdown-uncompressed 📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents) Dataset Summary This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025). The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.texttext-generation1K<n<10K0 likes3.6k downloads2mo agoHugging Face02opendatalab /awesome-markdown-ebooks Awesome-markdown-ebooks Your GitHub PDFs, Now AI-Ready. Project repo: https://github.com/OpenDataLab/awesome-markdown-ebooks text-generation100K<n<1M7 likes2.9k downloads1y agoHugging Face03scilons /SciLaD-all-markdown-v1 SciLaD (Markdown) SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more than 35 million publications. We also publish the extensible pipeline for generating SciLaD. Dataset Details In this repository we share the full… See the full description on the dataset page: https://huggingface.co/datasets/scilons/SciLaD-all-markdown-v1.text10M<n<100M0 likes1.9k downloads1mo agoHugging Face04CGIAR /Embrapa-ai-documents-markdown1 likes1.7k downloads2mo agoHugging Face05philschmid /markdown-documentation-transformers Hugging Face Transformers documentation as markdown dataset This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown. This dataset can be used to create RAG applications, which want to use the transformers documentation. Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.textn<1K11 likes942 downloads3y agoHugging Face06marin-community /ar5iv-no-problem-markdown Marin Markdownified Ar5iv Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 2.74B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text. Value Tokens 2 742 463 924 Primary source https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/ File format JSONL License C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-no-problem-markdown.texttext-generation100K<n<1M5 likes803 downloads1y agoHugging Face07CGIAR /ifpri-ai-documents-markdown GAIA / GARDIAN-CIGI Agricultural Research Corpus This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline. Dataset Overview Property Value Total Documents 21,726 Total Size 623.27 MB Total Tokens 85,359,442 Total Pages 0 Languages 25 Unique Keywords 7,127 Resource Types 20 Date Generated 2026-07-31 02:55:20 Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents-markdown.summarization10K<n<100K1 likes798 downloads2mo agoHugging Face08albertklorer /safedocs-markdown-200k-unifiedtext10K<n<100K0 likes636 downloads1mo agoHugging Face09CGIAR /gardian-cigi-ai-documents-markdown1 likes632 downloads2mo agoHugging Face10marin-community /ar5iv-warning-markdown Marin Markdownified Ar5iv Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 22.34B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text. Value Tokens 19 552 307 274 Primary source https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/ File format JSONL License C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-warning-markdown.texttext-generation1M<n<10M0 likes609 downloads1y agoHugging Face11AIMindLink /AlphaPrompt-QuantumLullaby-Markdown Quantum Lullaby Books (Markdown - AI Optimized) 🤖 44 books in markdown format for AI training, RAG, and processing. 🤖 For AI Systems Clean markdown formatting Optimized for tokenization Easy parsing for RAG/training No formatting artifacts 🔗 Also Available PDF (Human-readable): AlphaPrompt-QuantumLullaby-PDF Extracted SFT Datasets: AlphaPrompt-Metatron-SFT 💡 Use Cases RAG systems: Load as context documents Training data:… See the full description on the dataset page: https://huggingface.co/datasets/AIMindLink/AlphaPrompt-QuantumLullaby-Markdown.0 likes600 downloads3d agoHugging Face12NAMAA-Space /QariOCR-v0.3-markdown-mixed-dataset QARI Markdown Mixed Dataset 📋 Dataset Summary The QARI v0.3 Markdown Mixed Dataset is a specialized synthetic dataset designed for training Arabic OCR models with a focus on complex document layouts and HTML structure understanding. This dataset is part of the QARI-OCR project, which achieves state-of-the-art performance in Arabic text recognition. This dataset contains 37,000 synthetically generated Arabic document images (29.6k train, 3.7k… See the full description on the dataset page: https://huggingface.co/datasets/NAMAA-Space/QariOCR-v0.3-markdown-mixed-dataset.imagetext-to-image10K<n<100K13 likes558 downloads1y agoHugging Face13open-index /open-wikipedia-markdown Open Wikipedia (Markdown) Every Wikipedia article converted to clean Markdown, organized by language and updated from the latest Wikimedia dumps What is it? This dataset contains every article from every language edition of Wikipedia, converted from raw MediaWiki markup into clean, readable Markdown. Headings, bold, italic, code blocks, and internal links are all preserved as proper Markdown syntax, while templates, infoboxes, references, tables, categories, and… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-wikipedia-markdown.text-generation10M<n<100M17 likes365 downloads4mo agoHugging Face14marin-community /wikipedia-markdown Marin Markdownified Wikipedia Markdownified Wikipedia is a large-scale, pre-processed version of the English Wikipedia Enterprise HTML dump consisting of 8.59B tokens. The corpus has been converted to clean, section-aware Markdown for language-model training. Value Tokens 8 587 224 558 Primary source https://dumps.wikimedia.org/other/enterprise_html/runs/20241201/enwiki-NS0-20241201-ENTERPRISE-HTML.json.tar.gz File format JSONL License CC-BY-SA 4.0 (mirrors… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/wikipedia-markdown.texttext-generation1M<n<10M7 likes348 downloads1y agoHugging Face15albertklorer /safedocs-markdown SafeDocs PaddleOCR-VL 1.6 production This public corpus contains exactly 2,000,000 eligible document pages from 355,429 documents processed with PaddleOCR-VL 1.6. Eligible PDFs contain 1–39 pages and were rendered at 144 DPI. The published reduction is deterministic and document-safe: it retains 1,600 complete ordered Parquet shards (1,999,469 pages), then 103 whole documents from the next shard (531 pages). No document is split across the selection boundary. The exact public… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-markdown.image100K<n<1M0 likes337 downloads1mo agoHugging Face16astr010 /sec-10k-markdown-compressed 📦 SEC 10-K Full Markdown Filings (Compressed Archives) Dataset Summary This dataset contains 12,361 full-length SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025). The filings are packaged as 1,379 compressed .tar.gz archives (1 archive per company ticker). Each archive contains the un-chunked annual 10-K Markdown reports and a JSON metadata manifest (manifest.json).… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-compressed.text-generation10M<n<100M0 likes303 downloads2mo agoHugging Face17GreenNode /GreenNode-Table-Markdown-Retrieval-VN GreenNodeTableMarkdownRetrieval An MTEB dataset Massive Text Embedding Benchmark GreenNodeTable documents Task category t2t Domains Financial, Encyclopaedic, Non-fiction Reference https://huggingface.co/GreenNode Source datasets: GreenNode/GreenNode-Table-Markdown-Retrieval-VN How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("GreenNodeTableMarkdownRetrieval")… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/GreenNode-Table-Markdown-Retrieval-VN.texttext-retrieval100K<n<1M5 likes252 downloads9mo agoHugging Face18marin-community /stackexchange-markdown Marin Markdownified StackExchange Markdownified Stack Exchange transforms the Stack Exchange's question-answer pairs into Markdown format consisting of 20.4B tokens. This dataset preserves the content contained in technical discussions while organizing it into a thread format for language model training. Value Tokens 20 413 785 853 Primary source https://archive.org/details/stackexchange File format JSONL License CC (mirrors upstream SE licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/stackexchange-markdown.texttext-generation10M<n<100M6 likes237 downloads1y agoHugging Face19samwaugh /artefact-markdown ArteFact Markdown Dataset A comprehensive collection of scholarly art historical texts and associated images, organized by individual artworks and their corresponding scholarly documentation. �� Dataset Overview This dataset contains 7,200 individual art works, each with their associated scholarly markdown documentation and visual materials. The dataset serves as the primary source material for the ArteFact AI research platform, which bridges visual art and textual… See the full description on the dataset page: https://huggingface.co/datasets/samwaugh/artefact-markdown.1 likes233 downloads1y agoHugging Face20Delta-Vector /Tauri-RL-Markdown-System-V2textn<1K0 likes230 downloads1mo agoHugging Face21andersonbcdefg /github_issues_markdowntext10K<n<100K4 likes184 downloads3y agoHugging Face22casperhansen /pmc-oa-markdown PubMed Central (PMC) Open Access in Markdown This is a subset filtered for the words: obesity, weight loss, and diabetes. The purpose of this extraction is to further research in Biomedicine + NLP. While there are many biomedical datasets, most of them focus on the clinical space, which is often different from the early developments and discovery. The Markdown format includes figure description, tables, and a YAML header with some metadata like original license. It is filtered to… See the full description on the dataset page: https://huggingface.co/datasets/casperhansen/pmc-oa-markdown.text100K<n<1M8 likes178 downloads1y agoHugging Face23vibrantlabsai /Sample_Docs_Markdown4 likes169 downloads2y agoHugging Face24nielsr /datasets-docs-markdown-sectionstextn<1K0 likes156 downloads2y agoHugging Face25neonforestmist /smolgpt-markdown-stories SmolGPT-Fables Stories A deterministic, text-only corpus of 96,000 original English Markdown stories built for SmolGPT-Fables. Every row is one complete supervised story example with an exact prompt / completion boundary, a requested scene count from one to six, and plain-language conditioning fields. No model, API, browser, or network service was used to create this dataset. Dataset summary 96,000 stories across 96,000 isolated story families 25 genres and all… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/smolgpt-markdown-stories.tabulartext-generation10K<n<100K0 likes147 downloads2mo agoHugging Face26presightai /arabic_doc_to_markdown Dataset Card for presightai/arabic_doc_to_markdown This dataset contains OCR image–markdown pairs specifically curated for document structure retrieval and reconstruction tasks,focusing on Arabic and English documents across a variety of domains. Dataset Summary Source of Information:This dataset was created by crawling publicly accessible Arabic and English PDFs from: Official Arabic government document portals Arabic news websites and online magazines Community and… See the full description on the dataset page: https://huggingface.co/datasets/presightai/arabic_doc_to_markdown.imageimage-to-text10K<n<100K5 likes142 downloads1y agoHugging Face27CGIAR /usda-nal-ai-documents-en-markdown1 likes140 downloads2mo agoHugging Face28altaidevorg /yargitay-kararlar-markdowntext1M<n<10M1 likes119 downloads2y agoHugging Face29franktheglock /html-to-markdown-10000-pairs HTML to Markdown 10,000 Pairs This dataset contains 10,000 synthetic paired examples for HTML-to-Markdown conversion. Structure pairs/train.jsonl: 9,000 examples pairs/validation.jsonl: 500 examples pairs/test.jsonl: 500 examples manifest.json: generation manifest and split counts Each JSONL row includes: id: example identifier html: source HTML string markdown: target Markdown string metadata: per-example metadata Example {"id":"sample-0000000"… See the full description on the dataset page: https://huggingface.co/datasets/franktheglock/html-to-markdown-10000-pairs.texttranslation10K<n<100K0 likes117 downloads6mo agoHugging Face30verify-ppt /marin-starcoderdata_markdown0 likes113 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.