ground-truth
Audreygyj_-_pythia-160m-online-dpo-ground-truth-lead-merge-ggufAudreygyj_-_pythia-1b-online-dpo-ground-truth-lead-merge-ggufrethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step150rethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step50rethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step100gpt-oss-20b-olympiads-ground-truth-false-on-policy-1e5-1multiwoz_with_ground_truth_actpythia-160m-online-dpo-ground-truth-lead-merge
groundtruth-dynamic-benchmarking-submissions
Groundtruth Dynamic Benchmarking — Geology — Submissions
Community-submitted evaluation runs against the groundtruth-dynamic-benchmarking geology rubrics, feeding the leaderboard. We are currently running two tracks: model benchmarking (comparing different models with no special harness) and harness benchmarking (comparing different harnesses using a single standard model - GLM 4.7).
Each submission is a pointwise rubric score: one model, scored 0–10 per question against a… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking-submissions.kat57-ground-truth
Kat57 ground truth
Hugging Face conversion of Lund University Library's
Kat57 ground-truth release: 10,695 scanned catalogue
cards with manually corrected PAGE XML transcriptions.
The cards come from Catalogue -1957, Lund University Library's alphabetical
catalogue of holdings published through 1957. They contain a mixture of
typewritten and handwritten text in several languages.
Fields
image: original PNG card scan
reference: line transcriptions joined in PAGE… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ground-truth.groundtruth-dynamic-benchmarking
Groundtruth Dynamic Benchmarking — Geology
Question sets and grading rubrics for evaluating LLMs on real-world geological
reasoning. Every question is authored from a real source corpus, and every
claim in the grading key carries an evidence locator back to that corpus —
nothing is synthetic. Licensing/redistribution status varies by corpus — see
License.
This dataset holds the questions, grading rubrics, and source corpora.
Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.ground-truth-mmmu-pro-visionmultiple_samples_ground_truth_numina_aimemultichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.
