datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chess-position-evaluations
Dataset Card for the Lichess Evaluations dataset
Dataset Description
394,669,566 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 957,860,115 rows.
This dataset is updated monthly, and was last updated on July 8th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.user-evaluationsnyush-galaxea-a1-lingbot-va-real-world-evaluations
LingBot-VA on Galaxea A1 — Real-World Evaluations
Fruit-placement rollouts and open-loop diagnostics of object grounding,
layout generalization, and predicted robot motion.
Fruit step-1000: lemon-to-plate rollout in the Official layout.
Evidence
Scale
Real closed-loop rollouts
61 archived; 60 scored
Matched base-model controls
9 predictions
Post-trained diagnostics
48 full-horizon predictions; 1,211 rolling futures
Controlled OOD studies
558 predictions… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-lingbot-va-real-world-evaluations.chess-evaluations
Chess Evaluations Dataset
This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations:
tactics: Includes chess positions, their evaluations, and the best move in the position.
randoms: Contains random chess positions and their evaluations.
chess_data: General chess positions with evaluations.
This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.patient-evaluations
Patient Evaluations Dataset
This dataset contains clinician evaluations of AI-generated patient summaries from MIMIC-III data.
Dataset Description
The dataset includes expert clinician assessments of AI-generated patient summaries, with detailed ratings across multiple dimensions including clinical accuracy, completeness, relevance, and identification of hallucinations or critical omissions.
Dataset Structure
The dataset contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/JesseLiu/patient-evaluations.evaluation_singularity_imagesThis dataset repository holds the singularity images for the shared OELLM CLI workflows in OpenEuroLLM/oellm-cli.
This singularity images are updated automatically using a GitHub Actions workflow if there is a change to the container definition files or the workflow itself is updated.
The oellm-cli tool will detect changes to the respective singularity image file in this repo and download it to the cluster the user is launching the workflow from, before the workflow task is scheduled.
circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations
CircuitLens & WeightLens: Transcoder Descriptions and Evaluations
This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods.
Methods
CircuitLens: https://github.com/egolimblevskaia/CircuitLens
WeightLens: https://github.com/egolimblevskaia/WeightLens
Dataset Structure
The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.model_response_evaluationsThis dataset contains the evaluation results for the responses provided by different models to the INTIMA prompts.
The classification follows a two-level taxonomy.
We predict one label for the high-level category, and a relevance level for each of the sub-categories (in ["null", "low", "medium", "high"]).
A sub-category can have relevance even when it is not from the predicted top-level category.
The toxonomy is as follows:
{
"companionship_reinforcing": {
"classification_code":… See the full description on the dataset page: https://huggingface.co/datasets/AI-companionship/model_response_evaluations.nyush-galaxea-a1-lingbot-va-plug-insertion-evaluations
LingBot-VA on Galaxea A1: plug-insertion Evaluations
Real-robot closed-loop evaluation of a LingBot-VA checkpoint trained to pick up
a charger and insert it into the leftmost socket of a power strip.
Result
Valid rollouts
Successes
Failures
Success rate
12
1
11
8.3%
Five additional runs stopped at the initial grasp and are excluded because they
did not evaluate insertion.
Representative rollouts
SuccessFailure
Stable insertion… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-lingbot-va-plug-insertion-evaluations.turkish-brand-bias-evaluations
Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri
Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü
incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir.
An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in
Turkish product and service recommendations.
Veri seti özeti
300 tamamlanmış ve judge edilmiş yanıt
Domainler: VPN (150) ve kozmetik (150)
Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.extrinsic-evaluations
Extrinsic evaluations — the union view
One tidy long-format table of every extrinsic (downstream, task-level) evaluation
produced across the 2026-08-26 mergeability workstreams, so that a single file answers
"how did model X score on benchmark Y" regardless of which experiment produced it.
The per-experiment datasets remain the authoritative record of their own methods,
figures and caveats. This is the union view, not a replacement, and it deliberately
carries no analysis of its… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/extrinsic-evaluations.high_quality_private_evaluationsHigh-quality question-answer pairs, from private versions of datasets designed to mimic ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/.
Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC-like questions), and a question quality score.
turblimp-evaluations
TurBLiMP Evaluations
This dataset hosts the TurBLiMP evaluation results on my Turkish Model Zoo.
More about the TurBLiMP benchmark:
TurBLiMP is the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs).
This benchmark covers 16 core grammatical phenomena in Turkish, with 1,000 minimal pairs per phenomenon.
Additionally, it incorporates experimental paradigms that examine model… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/turblimp-evaluations.high_quality_public_evaluationsHigh-quality question-answer pairs, originally from ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/.
Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC questions), and a question quality score.
advent_of_code_evaluations
Advent of Code Evaluation
This evaluation is conducted on the advent of code dataset on several models including Qwen2.5-Coder-32B-Instruct, DeepSeek-V3-fp8, Llama-3.3-70B-Instruct, GPT-4o-mini, DeepSeek-R1.The aim is to to see how well these models can handle real-world puzzle prompts, generate correct Python code, and ultimately shed light on which LLM truly excels at reasoning and problem-solving.We used pass@1 to measure the functional correctness.
Results… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_evaluations.cardio_evaluationsasr-evaluationsasr-evaluationssentinel-evaluations
Sentinel Evaluations
Evaluation results for multiple alignment seeds across various AI safety benchmarks.
Overview
This dataset contains:
Seeds: Alignment prompts from different sources (Sentinel, FAS, Safyte xAI)
Results: Evaluation results across HarmBench, JailbreakBench, GDS-12, and more
Quick Start
from datasets import load_dataset
# Load seeds
seeds = load_dataset("sentinelseed/sentinel-evaluations", "seeds", split="train")
# Load results
results =… See the full description on the dataset page: https://huggingface.co/datasets/sentinelseed/sentinel-evaluations.chess_position_evaluationsbrand-bias-evaluations
Brand Bias in LLM Recommendations
Evaluation dataset measuring how 4 frontier LLMs recommend brands/products with and without web search, across 4 consumer domains.
Paper: PDF (source)Code: github.com/ThreeRiversAINexus/brand-bias-evaluationsDataset: huggingface.co/datasets/3RAIN/brand-bias-evaluationsContact: Three Rivers AI Nexus LLC — threeriversainexus@gmail.com — for custom evaluations and prompt optimization
Quick Start
from datasets import load_dataset
# Load one… See the full description on the dataset page: https://huggingface.co/datasets/3RAIN/brand-bias-evaluations.Evaluation-STAR-41K-Distillation-DeepSeek-R1-Distill-Qwen-7B-Size-16-BlockwiseCompressionclae-full-evaluations2mlops-repo-evaluationsclae-full-evaluations3memefact-llm-evaluations
MemeFact LLM Evaluations Dataset
This dataset contains 7,680 evaluation records where state-of-the-art Large Language Models (LLMs) assessed fact-checking memes according to specific quality criteria. The dataset provides comprehensive insights into how different AI models evaluate visual-textual content and how these evaluations compare to human judgments.
Dataset Description
Overview
The "MemeFact LLM Evaluations" dataset documents a systematic… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-llm-evaluations.hellora-olmoe-gsm8k-evaluations-seed42
OLMoE HELLoRA replication: held-out GSM8K generations
Reference archive containing every held-out GSM8K generation used to
compare the pinned pretrained base, selective HELLoRA, and full LoRA in the
single-GPU replication.
Source: https://github.com/toolanzyhhh1234/HELLoRA-replication
HELLoRA checkpoint:
https://huggingface.co/toolazyhhh123/hellora-olmoe-1b-7b-gsm8k-seed42
Full LoRA checkpoint:
https://huggingface.co/toolazyhhh123/lora-olmoe-1b-7b-gsm8k-seed42… See the full description on the dataset page: https://huggingface.co/datasets/toolazyhhh123/hellora-olmoe-gsm8k-evaluations-seed42.Math-LLM-Evaluationsfleurs-reducedbaseline-model-evaluationsenergy_D_eval_evaluations_v6
