CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-community /ar5iv-no-problem-markdown Marin Markdownified Ar5iv Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 2.74B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text. Value Tokens 2 742 463 924 Primary source https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/ File format JSONL License C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-no-problem-markdown.texttext-generation100K<n<1M5 likes821 downloads1y agoHugging Face02marin-community /ar5iv-warning-markdown Marin Markdownified Ar5iv Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 22.34B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text. Value Tokens 19 552 307 274 Primary source https://sigmathling.kwarc.info/resources/ar5iv-dataset-2024/ File format JSONL License C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/ar5iv-warning-markdown.texttext-generation1M<n<10M0 likes613 downloads1y agoHugging Face03marin-community /wikipedia-markdown Marin Markdownified Wikipedia Markdownified Wikipedia is a large-scale, pre-processed version of the English Wikipedia Enterprise HTML dump consisting of 8.59B tokens. The corpus has been converted to clean, section-aware Markdown for language-model training. Value Tokens 8 587 224 558 Primary source https://dumps.wikimedia.org/other/enterprise_html/runs/20241201/enwiki-NS0-20241201-ENTERPRISE-HTML.json.tar.gz File format JSONL License CC-BY-SA 4.0 (mirrors… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/wikipedia-markdown.texttext-generation1M<n<10M7 likes314 downloads1y agoHugging Face04marin-community /stackexchange-markdown Marin Markdownified StackExchange Markdownified Stack Exchange transforms the Stack Exchange's question-answer pairs into Markdown format consisting of 20.4B tokens. This dataset preserves the content contained in technical discussions while organizing it into a thread format for language model training. Value Tokens 20 413 785 853 Primary source https://archive.org/details/stackexchange File format JSONL License CC (mirrors upstream SE licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/stackexchange-markdown.texttext-generation10M<n<100M6 likes312 downloads1y agoHugging Face05yifeihu /ACL-23-Paper-OCR-Markdown ACL 2023 Paper in Markdown after OCR This dataset contains 2150 papers from Association for Computational Linguistics (ACL) 2023: Long Papers (912 papers) Short Papers (185 papers) System Demonstrations (59 paper) Student Research Workshop (35 papers) Industry Track (77 papers) Tutorial Abstracts (7 papers) Findings (902 papers) This dataset is processed and compiled by @hu_yifei as part of open-source effort from the Open Research Assistant Project. OCR process The… See the full description on the dataset page: https://huggingface.co/datasets/yifeihu/ACL-23-Paper-OCR-Markdown.textsummarization1K<n<10K19 likes35 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.