datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bne-hemeroteca-ocr-xix
BNE Hemeroteca OCR Dataset (19th Century)
This dataset provides full-text OCR and high-resolution page images for 19th-century Spanish publications sourced from the Biblioteca Nacional de España (BNE) - Hemeroteca Digital. It consists of over 830,000 pages from approximately 40,000 documents, totaling roughly 800 million text tokens.
Note on Temporal Coverage: Although the selected publications originated in the 19th century, some long-running titles extend into the early 20th… See the full description on the dataset page: https://huggingface.co/datasets/ferjorosa/bne-hemeroteca-ocr-xix.perception-mcp-benchmark
Perception MCP Workflow Benchmark
A side-by-side benchmark of AI assistants doing real digital-asset research workflows, with and without the Perception MCP connected.
Question: does connecting an industry-specific data corpus to a frontier AI assistant produce measurably better research work than the same assistant with its native web search?
Answer, across 48 scored runs: yes. Blind-judged mean score 11.4 → 16.4 (max 25, +44%), with the widest gains in recency (+74%) and… See the full description on the dataset page: https://huggingface.co/datasets/ferniko/perception-mcp-benchmark.SauerkrautLM-Fermented-Irrelevance-GER-DPO
SauerkrautLM-Fermented-Irrelevance-GER-DPO Dataset
Overview
SauerkrautLM-Fermented-Irrelevance-GER-DPO is a specialized dataset designed for training language models in function calling irrelevance detection using Direct Preference Optimization (DPO). The dataset consists of 2,000 carefully evaluated instruction-response pairs, specifically curated to help models recognize situations where function calls are unnecessary and direct responses are more appropriate.… See the full description on the dataset page: https://huggingface.co/datasets/VAGOsolutions/SauerkrautLM-Fermented-Irrelevance-GER-DPO.SauerkrautLM-Fermented-GER-DPO
SauerkrautLM-Fermented-GER-DPO Dataset
Overview
SauerkrautLM-Fermented-GER-DPO is a high-quality German instruction-response dataset specifically designed for Direct Preference Optimization (DPO) training. The dataset consists of 3,305 instruction-response pairs. Rather than being merged from existing German datasets, it was carefully created through a sophisticated augmentation process, transforming curated English instructions and responses into culturally adapted… See the full description on the dataset page: https://huggingface.co/datasets/VAGOsolutions/SauerkrautLM-Fermented-GER-DPO.bnqmark-20
BNqMark-20
BNqMark-20 is a benchmark dataset for evaluating Large Language Models (LLMs) on exact probabilistic inference in discrete Bayesian Networks. It isolates probabilistic computation from linguistic interpretation by giving models complete conditional probability table (CPT) specifications and asking them to answer conditional probability queries.
The dataset includes 78 Bayesian networks with 4-20 binary variables, 434 conditional probability queries, and 7,812 LLM… See the full description on the dataset page: https://huggingface.co/datasets/ferjorosa/bnqmark-20.chatgptprompts
📚 Custom ChatGPT Prompts Collection
This dataset is based on the fka/awesome-chatgpt-prompts and includes custom additions.
✨ Added custom prompts related to:
Islamic reflections
Biotechnology tutoring
Startup and productivity coaching
Ethical, creative, and futuristic prompts
Perfect for:
Students
Developers
Educa
fertility-medical-reasoning
Tanit Fertility Medical Reasoning Dataset
Dataset Description
High-quality medical reasoning dataset specialized for fertility care.
Total: 45,183 samples (42,923 train / 2,260 validation)
Dataset Composition
Source
Samples
Percentage
Purpose
MedReason
27,711
61.3%
Expert-verified clinical reasoning
Medical-O1
14,272
31.6%
O1-style systematic thinking
Synthetic-Fertility
3,200
7.1%
Domain-specific fertility cases
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/MohamedISSAOUI/fertility-medical-reasoning.parity-fertility-atlas
Parity fertility atlas
How many tokens each tokenizer charges for the same meaning, measured on a
parallel corpus (opus100).
column
meaning
tokenizer_id
the tokenizer measured
lang
ISO code
tokens_per_char
tokens per NFC character, excluding whitespace
tokens_per_word
tokens per whitespace word; null for scripts without word spaces
parity_ratio
tokens(target) / tokens(aligned English) — the headline
parity_ratio_median
median of the per-sentence ratios… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/parity-fertility-atlas.medease
