datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ifpri-ai-documents-markdown
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
21,726
Total Size
623.27 MB
Total Tokens
85,359,442
Total Pages
0
Languages
25
Unique Keywords
7,127
Resource Types
20
Date Generated
2026-07-31 02:55:20
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents-markdown.open-wikipedia-markdown
Open Wikipedia (Markdown)
Every Wikipedia article converted to clean Markdown, organized by language and updated from the latest Wikimedia dumps
What is it?
This dataset contains every article from every language edition of Wikipedia, converted from raw MediaWiki markup into clean, readable Markdown. Headings, bold, italic, code blocks, and internal links are all preserved as proper Markdown syntax, while templates, infoboxes, references, tables, categories, and… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-wikipedia-markdown.pmc-oa-markdown-qa-results
A hard biological benchmark that requires search
The goal of a model in drug discovery is to find the right information and make a solid conclusion. This benchmark was synthesized from a total of 6 semantically similar papers, and if a model is able to rediscover these papers on its own through search or retrieval, it will have a good shot at being able to answer the question correctly.
The evaluation dataset was adversarially constructed so that DeepSeek R1:
could correctly… See the full description on the dataset page: https://huggingface.co/datasets/casperhansen/pmc-oa-markdown-qa-results.markdown-table-expert
Markdown Table Expert
A large-scale dataset for teaching language models to read, understand, and reason over markdown tables. Contains 44,000 samples (40,000 train + 4,000 validation) spanning 35 real-world domains with detailed step-by-step reasoning traces.
Why This Dataset
Markdown tables are everywhere — in documentation, reports, READMEs, financial statements, and web content. Yet most LLMs struggle with structured tabular data, especially when asked to perform… See the full description on the dataset page: https://huggingface.co/datasets/cetusian/markdown-table-expert.telelogs_markdown
TeleLogs Dataset (Processed MCQ Format)
This dataset has been extracted from the original netop/TeleLogs dataset and processed into multiple-choice question (MCQ) format for easier evaluation.
Dataset Description
TeleLogs is a telecommunications log analysis benchmark where models must identify the root cause of network issues from 5G wireless network drive-test data and engineering parameters.
Processed Format
This version has been restructured for MCQ… See the full description on the dataset page: https://huggingface.co/datasets/emolero/telelogs_markdown.
