CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CGIAR /ifpri-ai-documents-markdown GAIA / GARDIAN-CIGI Agricultural Research Corpus This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline. Dataset Overview Property Value Total Documents 21,726 Total Size 623.27 MB Total Tokens 85,359,442 Total Pages 0 Languages 25 Unique Keywords 7,127 Resource Types 20 Date Generated 2026-07-31 02:55:20 Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents-markdown.summarization10K<n<100K1 likes903 downloads2mo agoHugging Face02open-index /open-wikipedia-markdown Open Wikipedia (Markdown) Every Wikipedia article converted to clean Markdown, organized by language and updated from the latest Wikimedia dumps What is it? This dataset contains every article from every language edition of Wikipedia, converted from raw MediaWiki markup into clean, readable Markdown. Headings, bold, italic, code blocks, and internal links are all preserved as proper Markdown syntax, while templates, infoboxes, references, tables, categories, and… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-wikipedia-markdown.text-generation10M<n<100M17 likes364 downloads4mo agoHugging Face03casperhansen /pmc-oa-markdown-qa-results A hard biological benchmark that requires search The goal of a model in drug discovery is to find the right information and make a solid conclusion. This benchmark was synthesized from a total of 6 semantically similar papers, and if a model is able to rediscover these papers on its own through search or retrieval, it will have a good shot at being able to answer the question correctly. The evaluation dataset was adversarially constructed so that DeepSeek R1: could correctly… See the full description on the dataset page: https://huggingface.co/datasets/casperhansen/pmc-oa-markdown-qa-results.textquestion-answering1K<n<10K0 likes121 downloads10mo agoHugging Face04cetusian /markdown-table-expert Markdown Table Expert A large-scale dataset for teaching language models to read, understand, and reason over markdown tables. Contains 44,000 samples (40,000 train + 4,000 validation) spanning 35 real-world domains with detailed step-by-step reasoning traces. Why This Dataset Markdown tables are everywhere — in documentation, reports, READMEs, financial statements, and web content. Yet most LLMs struggle with structured tabular data, especially when asked to perform… See the full description on the dataset page: https://huggingface.co/datasets/cetusian/markdown-table-expert.tabularquestion-answering10K<n<100K0 likes33 downloads6mo agoHugging Face05emolero /telelogs_markdown TeleLogs Dataset (Processed MCQ Format) This dataset has been extracted from the original netop/TeleLogs dataset and processed into multiple-choice question (MCQ) format for easier evaluation. Dataset Description TeleLogs is a telecommunications log analysis benchmark where models must identify the root cause of network issues from 5G wireless network drive-test data and engineering parameters. Processed Format This version has been restructured for MCQ… See the full description on the dataset page: https://huggingface.co/datasets/emolero/telelogs_markdown.textmultiple-choicen<1K0 likes29 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.