datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chess-position-evaluations
Dataset Card for the Lichess Evaluations dataset
Dataset Description
394,669,566 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 957,860,115 rows.
This dataset is updated monthly, and was last updated on July 8th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.chess-evaluations
Chess Evaluations Dataset
This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations:
tactics: Includes chess positions, their evaluations, and the best move in the position.
randoms: Contains random chess positions and their evaluations.
chess_data: General chess positions with evaluations.
This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations
CircuitLens & WeightLens: Transcoder Descriptions and Evaluations
This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods.
Methods
CircuitLens: https://github.com/egolimblevskaia/CircuitLens
WeightLens: https://github.com/egolimblevskaia/WeightLens
Dataset Structure
The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.turkish-brand-bias-evaluations
Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri
Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü
incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir.
An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in
Turkish product and service recommendations.
Veri seti özeti
300 tamamlanmış ve judge edilmiş yanıt
Domainler: VPN (150) ve kozmetik (150)
Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.extrinsic-evaluations
Extrinsic evaluations — the union view
One tidy long-format table of every extrinsic (downstream, task-level) evaluation
produced across the 2026-08-26 mergeability workstreams, so that a single file answers
"how did model X score on benchmark Y" regardless of which experiment produced it.
The per-experiment datasets remain the authoritative record of their own methods,
figures and caveats. This is the union view, not a replacement, and it deliberately
carries no analysis of its… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/extrinsic-evaluations.cardio_evaluationsasr-evaluationssentinel-evaluations
Sentinel Evaluations
Evaluation results for multiple alignment seeds across various AI safety benchmarks.
Overview
This dataset contains:
Seeds: Alignment prompts from different sources (Sentinel, FAS, Safyte xAI)
Results: Evaluation results across HarmBench, JailbreakBench, GDS-12, and more
Quick Start
from datasets import load_dataset
# Load seeds
seeds = load_dataset("sentinelseed/sentinel-evaluations", "seeds", split="train")
# Load results
results =… See the full description on the dataset page: https://huggingface.co/datasets/sentinelseed/sentinel-evaluations.chess_position_evaluationsbrand-bias-evaluations
Brand Bias in LLM Recommendations
Evaluation dataset measuring how 4 frontier LLMs recommend brands/products with and without web search, across 4 consumer domains.
Paper: PDF (source)Code: github.com/ThreeRiversAINexus/brand-bias-evaluationsDataset: huggingface.co/datasets/3RAIN/brand-bias-evaluationsContact: Three Rivers AI Nexus LLC — threeriversainexus@gmail.com — for custom evaluations and prompt optimization
Quick Start
from datasets import load_dataset
# Load one… See the full description on the dataset page: https://huggingface.co/datasets/3RAIN/brand-bias-evaluations.mlops-repo-evaluationsmemefact-llm-evaluations
MemeFact LLM Evaluations Dataset
This dataset contains 7,680 evaluation records where state-of-the-art Large Language Models (LLMs) assessed fact-checking memes according to specific quality criteria. The dataset provides comprehensive insights into how different AI models evaluate visual-textual content and how these evaluations compare to human judgments.
Dataset Description
Overview
The "MemeFact LLM Evaluations" dataset documents a systematic… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-llm-evaluations.energy_D_eval_evaluations_v6em-gemma-2-9b-it-layer-16-evaluationsenergy-eval-filtered_evaluations_v3tinyllama_sql_evaluationsprobe-evaluations-gemma-2-9b-layer20benchmark-energy-mcq-harder_evaluations_easy2reproscreener_manual_evaluationsenergy-eval-filtered_evaluations_v4_cleantinydolphin_sql_evaluationsenergy-eval-filtered_evaluations_v3benchmark-energy-mcq-harder_evaluations_hard2_basechess-position-evaluations
Dataset Card for the Lichess Evaluations dataset
Dataset Description
342,059,879 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 844,812,067 rows.
This dataset is updated monthly, and was last updated on January 6th, 2026.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/coryvegan/chess-position-evaluations.benchmark-energy-mcq-harder_evaluations_hard2OpenHerms7B_Q4_k_m_sql_evaluationsadaption-marketing-fit-evaluations
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-marketing_fit_evaluations
This dataset contains prompt-completion pairs where an AI evaluates marketing communications for specific regional markets, focusing on trust, cultural context, and decision friction. Each completion provides a market fit score, strategic diagnosis, and a rewritten copy optimized for the target audience and platform. The entries cover diverse sectors and… See the full description on the dataset page: https://huggingface.co/datasets/Debbyjaye001/adaption-marketing-fit-evaluations.energy_D_eval_evaluations_v3energy_D_eval_evaluations_v4benchmark-energy-mcq-harder_evaluations_hard
