datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kat57-ground-truth
Kat57 ground truth
Hugging Face conversion of Lund University Library's
Kat57 ground-truth release: 10,695 scanned catalogue
cards with manually corrected PAGE XML transcriptions.
The cards come from Catalogue -1957, Lund University Library's alphabetical
catalogue of holdings published through 1957. They contain a mixture of
typewritten and handwritten text in several languages.
Fields
image: original PNG card scan
reference: line transcriptions joined in PAGE… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ground-truth.groundtruth-dynamic-benchmarking
Groundtruth Dynamic Benchmarking — Geology
Question sets and grading rubrics for evaluating LLMs on real-world geological
reasoning. Every question is authored from a real source corpus, and every
claim in the grading key carries an evidence locator back to that corpus —
nothing is synthetic. Licensing/redistribution status varies by corpus — see
License.
This dataset holds the questions, grading rubrics, and source corpora.
Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.ground-truth-mmmu-pro-visionmultiple_samples_ground_truth_numina_aimemultichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.bnl_ground_truth_newspapers_before_1878
Dataset description
33.000 transcribed text lines from historical newspapers (before 1878) along with the cropped images of the original scans
Text line based OCR
19.000 text lines in Antiqua
14.000 text lines in Fraktur
Transcribed using double-keying (99.95% accuracy)
Public Domain, CC0 (See copyright notice)
Best for training an OCR engine
The newspapers used are:
Le Gratis luxembourgeois (1857-1858)
Luxemburger Volks-Freund (1869-1876)
L'Arlequin (1848-1848)
Courrier du… See the full description on the dataset page: https://huggingface.co/datasets/biglam/bnl_ground_truth_newspapers_before_1878.ground-truth-mmmu-pro-standard-10bfsi-bench
BFSI-Bench
BFSI-Bench is a benchmark for testing how well language models answer questions about India’s banking, financial services, and insurance (BFSI) rules.
In this domain, the correct answer often depends on circulars and regulations that change frequently, and the official sources (sites like RBI, SEBI, and IRDAI) can be hard to find, parse, and keep current. BFSI-Bench measures five capability areas:
Jurisdiction-Aware Compliance: Disambiguate to the Indian context, or… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/bfsi-bench.groundtruthmath_ground_truth_minerva_styleClinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft
Dataset details:-
This dataset is basically mapping of final anchor-positive pair data with their refernce answer.
The given input data considered because:-
(i) it had the had purest anchor-positive pairs with semantically bound anchors with context/positive.
(ii) gave us the best result on final embedding fine tuning model.
The anchor-context(positive)-reference_answer data has been generated via Qwen-2.5-7B teacher model with temperature 0.1 and a strict system prompt.… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft.kat57-ground-truth-500
Kat57 500-card benchmark sample
A deterministic 500-card sample of
tadad/kat57-ground-truth for comparing OCR systems
against Kat57's human-corrected transcriptions.
The sample is drawn from all 10,695 source rows by ranking each stable id with
SHA-256 over seed + NUL + id, selecting the lowest 500 digests, and restoring
source order. Sampling seed: 57. Pinned source revision: 2f4b7e6a8f8746631c0628280dd3f40be2b997f6.
Every selected row has a non-empty reference; all original… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ground-truth-500.llm-ground-truth-reasoningClinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA-junk_handled_ft
Dataset details:-
This dataset is the 2nd iteration following bugs in 1st dataset.
The initial data suffered with followoing cases:-
(i) The failed reference_answers generation(due error totalling 23) primarly because of 2 reasons/exceptions:- (a) There was normal limit(300 in 1st request) and worst case limit(450 in 3rd request) number of tokens for consolidated 4 refernce_answers per chunk and its 4 corresponding answers. However certain answers breached this higher… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA-junk_handled_ft.multiple_samples_ground_truth_openr1_llm_verifiermultiple_samples_ground_truth_openr1_llm_verifier_cleanNuminaMath-groundtrutheuropeana-newspapers-ground-truth
Europeana Newspapers — Historical Newspapers Ground Truth
50 pages of digitised historical German newspapers from the Berlin State Library
(Staatsbibliothek zu Berlin), with PAGE XML ground truth produced for the EU
Europeana Newspapers project.
Each row pairs three things: the page image, the human-corrected ground truth
(regions, polygons, reading order, text), and the ALTO OCR output that ABBYY FineReader
actually produced. That last column is what makes this an OCR… See the full description on the dataset page: https://huggingface.co/datasets/biglam/europeana-newspapers-ground-truth.groundtruth-hallucination-bench-sample
Groundtruth Data Hallucination Benchmark Sample
This public teaser contains 180 representative, source-backed examples from Groundtruth Data products.
Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is:
Find where a model fails.
Prove the failure with a larger verified evaluation.
Provide targeted remediation/training data.
Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.multiple_samples_ground_truth_numina_aime_w_openthoughtsllm-ground-truth-generalkat57-ground-truth-smoke
Kat57 ground-truth smoke subset
A 50-card integration subset of tadad/kat57-ground-truth, drawn from the first published Parquet shard.
This subset exists to test OCR pipelines and exact CER/WER scoring without downloading the full 36 GB collection. It is ordered by source identifier and is not a representative benchmark sample; substantive Kat57 claims should use a documented sample across the full collection.
All fields are preserved from the full conversion, including the… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ground-truth-smoke.prompts_with_constraints_for_ground_truthmetamathqa_ground_truth
Dataset Card for "metamathqa_ground_truth"
More Information needed
ground-truth-mmmu-pro-vision-sampling-500gsm8k_math_ground_truth
Dataset Card for "gsm8k_math_ground_truth"
More Information needed
math_ground_truthOpenVul_Ground_Truth_Vulnerability_InformationThis dataset provides ground truth vulnerability information (CWE ID, CVE description, commit message, and patch diff) for all samples in the OpenVul_Vulnerability_Query_Dataset_for_RL collection, enabling multi-granular model reward evaluation and performance evaluation.
math_ground_truth_zsgwalther-handwriting-ground-truth
Gwalther's Latin Handwriting Ground Truth
Dataset Description
The Gwalther's Latin Handwriting Ground Truth dataset provides resources for Handwritten Text Recognition (HTR) systems, specifically focused on a single hand from the Reformation era.
The dataset contains ground truth for the handwriting of Ruolph Gwalther (1519-1586), extracted from his work Lateinische Gedichte, which accumulated his writings between 1540 and 1580.
Key Figures
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/SSamDav/gwalther-handwriting-ground-truth.
