markdown
Datasets
All datasets matching “markdown”sec-10k-markdown-uncompressed
📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents)
Dataset Summary
This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025).
The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.awesome-markdown-ebooks
Awesome-markdown-ebooks
Your GitHub PDFs, Now AI-Ready.
Project repo: https://github.com/OpenDataLab/awesome-markdown-ebooks
Embrapa-ai-documents-markdownSciLaD-all-markdown-v1
SciLaD (Markdown)
SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more than 35 million publications. We also publish the extensible pipeline for generating SciLaD.
Dataset Details
In this repository we share the full… See the full description on the dataset page: https://huggingface.co/datasets/scilons/SciLaD-all-markdown-v1.markdown-documentation-transformers
Hugging Face Transformers documentation as markdown dataset
This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown.
This dataset can be used to create RAG applications, which want to use the transformers documentation.
Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.ifpri-ai-documents-markdown
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
21,726
Total Size
623.27 MB
Total Tokens
85,359,442
Total Pages
0
Languages
25
Unique Keywords
7,127
Resource Types
20
Date Generated
2026-07-31 02:55:20
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents-markdown.
