datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chess-position-evaluations
Dataset Card for the Lichess Evaluations dataset
Dataset Description
394,669,566 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 957,860,115 rows.
This dataset is updated monthly, and was last updated on July 8th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.chess-evaluations
Chess Evaluations Dataset
This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations:
tactics: Includes chess positions, their evaluations, and the best move in the position.
randoms: Contains random chess positions and their evaluations.
chess_data: General chess positions with evaluations.
This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.turkish-brand-bias-evaluations
Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri
Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü
incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir.
An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in
Turkish product and service recommendations.
Veri seti özeti
300 tamamlanmış ve judge edilmiş yanıt
Domainler: VPN (150) ve kozmetik (150)
Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.high_quality_private_evaluationsHigh-quality question-answer pairs, from private versions of datasets designed to mimic ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/.
Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC-like questions), and a question quality score.
advent_of_code_evaluations
Advent of Code Evaluation
This evaluation is conducted on the advent of code dataset on several models including Qwen2.5-Coder-32B-Instruct, DeepSeek-V3-fp8, Llama-3.3-70B-Instruct, GPT-4o-mini, DeepSeek-R1.The aim is to to see how well these models can handle real-world puzzle prompts, generate correct Python code, and ultimately shed light on which LLM truly excels at reasoning and problem-solving.We used pass@1 to measure the functional correctness.
Results… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_evaluations.high_quality_public_evaluationsHigh-quality question-answer pairs, originally from ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/.
Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC questions), and a question quality score.
cardio_evaluationsasr-evaluationsEvaluation-STAR-41K-Distillation-DeepSeek-R1-Distill-Qwen-7B-Size-16-BlockwiseCompressionmlops-repo-evaluationsbrand-bias-evaluations
Brand Bias in LLM Recommendations
Evaluation dataset measuring how 4 frontier LLMs recommend brands/products with and without web search, across 4 consumer domains.
Paper: PDF (source)Code: github.com/ThreeRiversAINexus/brand-bias-evaluationsDataset: huggingface.co/datasets/3RAIN/brand-bias-evaluationsContact: Three Rivers AI Nexus LLC — threeriversainexus@gmail.com — for custom evaluations and prompt optimization
Quick Start
from datasets import load_dataset
# Load one… See the full description on the dataset page: https://huggingface.co/datasets/3RAIN/brand-bias-evaluations.Math-LLM-Evaluationsenergy_D_eval_evaluations_v6em-gemma-2-9b-it-layer-16-evaluationsfleurs-reducedbaseline-model-evaluationsenergy-eval-filtered_evaluations_v3evaluation-setevaluation-super-sneaky-highlevelprobe-evaluations-gemma-2-9b-layer20benchmark-energy-mcq-harder_evaluations_easy2tinydolphin_sql_evaluationstinyllama_sql_evaluationsenergy-eval-filtered_evaluations_v3benchmark-energy-mcq-harder_evaluations_hard2_basechess-position-evaluations
Dataset Card for the Lichess Evaluations dataset
Dataset Description
342,059,879 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 844,812,067 rows.
This dataset is updated monthly, and was last updated on January 6th, 2026.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/coryvegan/chess-position-evaluations.energy-eval-filtered_evaluations_v4_cleanOpenHerms7B_Q4_k_m_sql_evaluationsevaluation-super-sneaky-lowlevelbenchmark-energy-mcq-harder_evaluations_hard2quality-evaluations
