datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
groundtruth-dynamic-benchmarking
Groundtruth Dynamic Benchmarking — Geology
Question sets and grading rubrics for evaluating LLMs on real-world geological
reasoning. Every question is authored from a real source corpus, and every
claim in the grading key carries an evidence locator back to that corpus —
nothing is synthetic. Licensing/redistribution status varies by corpus — see
License.
This dataset holds the questions, grading rubrics, and source corpora.
Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.aya101-benchmarking
Dataset Card for Evaluation run of CohereForAI/aya-101
Dataset automatically created during the evaluation run of model CohereForAI/aya-101
The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya101-benchmarking.building-energy-benchmarking-requirements
Building Energy Benchmarking and Performance Standard Requirements by Jurisdiction
Canonical, always-current version: https://referencesource.org/building-energy-benchmarking-requirements/
Machine-readable: https://referencesource.org/building-energy-benchmarking-requirements/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-15
Stale after: 2027-02-11 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records:… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/building-energy-benchmarking-requirements.llama-south-africa-benchmarking
Dataset Card for Evaluation run of chad-brouze/llama-8b-south-africa
Dataset automatically created during the evaluation run of model chad-brouze/llama-8b-south-africa
The dataset is composed of 17 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 14 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/llama-south-africa-benchmarking.InkubaLM-benchmarking
Dataset Card for Evaluation run of lelapa/InkubaLM-0.4B
Dataset automatically created during the evaluation run of model lelapa/InkubaLM-0.4B
The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/InkubaLM-benchmarking.Treebank-Benchmarking
Turkish Treebank Benchmarking
This is the repo for Turkish treebank benchmarking, namely evaluating Tranformer models on POS-Dep-Morph task.
For the data, we used two treebank, IMST and BOUN. We converted conllu format to json lines for being compatible to HF dataset formats.
Here are treebank sizes at a glance:
Dataset
train lines
dev lines
test lines
BOUN
7803
979
979
IMST
3435
1100
1100
A typical instance from the dataset looks like:
{
"id": "ins_1267"… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Treebank-Benchmarking.aya23-benchmarking
Dataset Card for Evaluation run of CohereForAI/aya-23-8B
Dataset automatically created during the evaluation run of model CohereForAI/aya-23-8B
The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya23-benchmarking.math_precision_benchmarking
🧮 Math Precision — Benchmarking
A Formal Framework for High-Precision Arithmetic Evaluation in Large Language Models
Math Precision — Benchmarking, developed by Sapiens Technology®, is a rigorous framework for evaluating the true arithmetic capabilities of large language models by generating fully stochastic, high-precision mathematical problems that eliminate memorization and heuristic guessing; operating in a 100-digit floating-point field, it forces extreme numerical precision… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/math_precision_benchmarking.Hindi_Benchmarking_questions
Hindi Language Benchmarking Dataset
Overview
This repository contains the first comprehensive Hindi language benchmarking dataset designed to evaluate the intelligence of language models (LLMs) in Hindi. The dataset includes 1000 questions across various topics and difficulty levels, providing a robust tool for assessing the capabilities of LLMs in understanding and processing the Hindi language.
Dataset Structure
The dataset is meticulously curated to… See the full description on the dataset page: https://huggingface.co/datasets/DrDrek/Hindi_Benchmarking_questions.aya-benchmarking
Dataset Card for Evaluation run of CohereForAI/aya-23-8B
Dataset automatically created during the evaluation run of model CohereForAI/aya-23-8B
The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya-benchmarking.Rethinking_Benchmarking_model_editing
