datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
groundtruth-dynamic-benchmarking
Groundtruth Dynamic Benchmarking — Geology
Question sets and grading rubrics for evaluating LLMs on real-world geological
reasoning. Every question is authored from a real source corpus, and every
claim in the grading key carries an evidence locator back to that corpus —
nothing is synthetic. Licensing/redistribution status varies by corpus — see
License.
This dataset holds the questions, grading rubrics, and source corpora.
Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.multichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.groundtruthgroundtruth-hallucination-bench-sample
Groundtruth Data Hallucination Benchmark Sample
This public teaser contains 180 representative, source-backed examples from Groundtruth Data products.
Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is:
Find where a model fails.
Prove the failure with a larger verified evaluation.
Provide targeted remediation/training data.
Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.dbpedia-hindi-benchie-ground-truth
DBpedia Hindi — BenchIE Ground Truth
The first DBpedia property ground truth for the Hindi BenchIE benchmark — 139 canonical triples across 112 sentences, built for the DBpedia Hindi Chapter (Google Summer of Code 2026).
Why This Was Needed
BenchIE contains human-verified gold subject/relation/object spans, but was designed for open information extraction evaluation, not DBpedia alignment — it had no mapping to DBpedia properties before this work.… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-benchie-ground-truth.womens-health-benchmark-ground-truth
Womens Health Benchmark Ground Truth
A curated evaluation dataset for assessing large language models (LLMs) in womens health-related tasks. The dataset consists of model stumps paired with expert-written justifications describing observed errors.
Research Focus
This dataset supports structured evaluation of LLM behavior in clinically relevant womens health contexts, with emphasis on safety, reasoning quality, and evidence alignment.
Methodology
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/therubricai/womens-health-benchmark-ground-truth.english_ground_truthmore_diverse_ground_truth_if_traingroundtruthtelco-gaia-groundtruth
Telco-GAIA — Ground Truth (gated)
Answers + gold reasoning steps for the 100 Telco-GAIA tasks, plus the scorer.
ground_truth.json — task_id, category, language, question, final_answer, answer_type, steps, tools
evaluate.py, answer_matching.py — GAIA-style exact-match scorer
Pair with the open dataset kaust-generative-ai/telco-gaia for the questions, website,
and harness.
python evaluate.py --submission submission.json --ground-truth ground_truth.json --output results.json
gsm8k_math_ground_truth_zero_shotDocStream_Ground_Truth_Compare
Feature
Type
Description
event_idx
int
Event index within the session (aligns with HF and local JSONL).
prompt
string
The full prompt used for the model prediction (from local JSONL).
pred_thinking
string
Model-generated chain-of-thought from the local JSONL.
thinking
string
Gold/reference chain-of-thought from HF.
pred_depth
int
Predicted depth value.
pred_annotation
string
Predicted one-sentence event annotation.
gt_depth
int
Gold/reference depth label.
gt_annotation… See the full description on the dataset page: https://huggingface.co/datasets/VictorShea/DocStream_Ground_Truth_Compare.
