CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01philschmid /markdown-documentation-transformers Hugging Face Transformers documentation as markdown dataset This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown. This dataset can be used to create RAG applications, which want to use the transformers documentation. Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.textn<1K11 likes943 downloads3y agoHugging Face02marin-community /ar5iv-no-problem-markdown Marin Markdownified Ar5iv Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 2.74B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text. Value Tokens 2 742 463 924 Primary source https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/ File format JSONL License C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-no-problem-markdown.texttext-generation100K<n<1M5 likes817 downloads1y agoHugging Face03marin-community /ar5iv-warning-markdown Marin Markdownified Ar5iv Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 22.34B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text. Value Tokens 19 552 307 274 Primary source https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/ File format JSONL License C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-warning-markdown.texttext-generation1M<n<10M0 likes612 downloads1y agoHugging Face04marin-community /wikipedia-markdown Marin Markdownified Wikipedia Markdownified Wikipedia is a large-scale, pre-processed version of the English Wikipedia Enterprise HTML dump consisting of 8.59B tokens. The corpus has been converted to clean, section-aware Markdown for language-model training. Value Tokens 8 587 224 558 Primary source https://dumps.wikimedia.org/other/enterprise_html/runs/20241201/enwiki-NS0-20241201-ENTERPRISE-HTML.json.tar.gz File format JSONL License CC-BY-SA 4.0 (mirrors… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/wikipedia-markdown.texttext-generation1M<n<10M7 likes339 downloads1y agoHugging Face05Delta-Vector /Tauri-RL-Markdown-System-V2textn<1K0 likes240 downloads1mo agoHugging Face06marin-community /stackexchange-markdown Marin Markdownified StackExchange Markdownified Stack Exchange transforms the Stack Exchange's question-answer pairs into Markdown format consisting of 20.4B tokens. This dataset preserves the content contained in technical discussions while organizing it into a thread format for language model training. Value Tokens 20 413 785 853 Primary source https://archive.org/details/stackexchange File format JSONL License CC (mirrors upstream SE licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/stackexchange-markdown.texttext-generation10M<n<100M6 likes236 downloads1y agoHugging Face07altaidevorg /yargitay-kararlar-markdowntext1M<n<10M1 likes121 downloads2y agoHugging Face08casperhansen /pmc-oa-markdown-qa-results A hard biological benchmark that requires search The goal of a model in drug discovery is to find the right information and make a solid conclusion. This benchmark was synthesized from a total of 6 semantically similar papers, and if a model is able to rediscover these papers on its own through search or retrieval, it will have a good shot at being able to answer the question correctly. The evaluation dataset was adversarially constructed so that DeepSeek R1: could correctly… See the full description on the dataset page: https://huggingface.co/datasets/casperhansen/pmc-oa-markdown-qa-results.textquestion-answering1K<n<10K0 likes118 downloads10mo agoHugging Face09franktheglock /html-to-markdown-10000-pairs HTML to Markdown 10,000 Pairs This dataset contains 10,000 synthetic paired examples for HTML-to-Markdown conversion. Structure pairs/train.jsonl: 9,000 examples pairs/validation.jsonl: 500 examples pairs/test.jsonl: 500 examples manifest.json: generation manifest and split counts Each JSONL row includes: id: example identifier html: source HTML string markdown: target Markdown string metadata: per-example metadata Example {"id":"sample-0000000"… See the full description on the dataset page: https://huggingface.co/datasets/franktheglock/html-to-markdown-10000-pairs.texttranslation10K<n<100K0 likes117 downloads6mo agoHugging Face10rojasdiego /chinese-markdown Chinese Markdown from datasets import load_dataset chinese_markdown = load_dataset("rojas-diego/chinese-markdown", split="train") Dataset({ features: ['code', 'size', 'license'], num_rows: 187258 }) text100K<n<1M4 likes104 downloads3y agoHugging Face11yifeihu /ACL-23-Paper-OCR-Markdown ACL 2023 Paper in Markdown after OCR This dataset contains 2150 papers from Association for Computational Linguistics (ACL) 2023: Long Papers (912 papers) Short Papers (185 papers) System Demonstrations (59 paper) Student Research Workshop (35 papers) Industry Track (77 papers) Tutorial Abstracts (7 papers) Findings (902 papers) This dataset is processed and compiled by @hu_yifei as part of open-source effort from the Open Research Assistant Project. OCR process The… See the full description on the dataset page: https://huggingface.co/datasets/yifeihu/ACL-23-Paper-OCR-Markdown.textsummarization1K<n<10K19 likes36 downloads2y agoHugging Face12emolero /telelogs_markdown TeleLogs Dataset (Processed MCQ Format) This dataset has been extracted from the original netop/TeleLogs dataset and processed into multiple-choice question (MCQ) format for easier evaluation. Dataset Description TeleLogs is a telecommunications log analysis benchmark where models must identify the root cause of network issues from 5G wireless network drive-test data and engineering parameters. Processed Format This version has been restructured for MCQ… See the full description on the dataset page: https://huggingface.co/datasets/emolero/telelogs_markdown.textmultiple-choicen<1K0 likes30 downloads10mo agoHugging Face13ANIKET5555555 /bank_statement_markdown_to_csvtextn<1K0 likes30 downloads3mo agoHugging Face14Delta-Vector /Tauri-RL-Markdown-Systemtextn<1K0 likes17 downloads5mo agoHugging Face15Coldyuja /junos-user-manual-markdown-raw Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Coldyuja/junos-user-manual-markdown-raw.text1K<n<10K0 likes13 downloads6mo agoHugging Face16dohyung97022 /markdownCreatorDatasetv2text1K<n<10K0 likes6 downloads2y agoHugging Face17LeroyDyer /Markdown_BIBLEStextn<1K1 likes4 downloads2y agoHugging Face18dohyung97022 /markdownCreatorDatasettextn<1K0 likes4 downloads2y agoHugging Face19red-ml /prolewiki-library-markdownifiedtext1K<n<10K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.