CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lichess /chess-position-evaluations Dataset Card for the Lichess Evaluations dataset Dataset Description 394,669,566 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 957,860,115 rows. This dataset is updated monthly, and was last updated on July 8th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.tabular100M<n<1B33 likes2.8k downloads3mo agoHugging Face02ssingh22 /chess-evaluations Chess Evaluations Dataset This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations: tactics: Includes chess positions, their evaluations, and the best move in the position. randoms: Contains random chess positions and their evaluations. chess_data: General chess positions with evaluations. This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.tabularquestion-answering10M<n<100M2 likes1k downloads2y agoHugging Face03furkankarli /turkish-brand-bias-evaluations Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir. An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in Turkish product and service recommendations. Veri seti özeti 300 tamamlanmış ve judge edilmiş yanıt Domainler: VPN (150) ve kozmetik (150) Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.tabulartext-generationn<1K1 likes101 downloads20d agoHugging Face04imbue /high_quality_private_evaluationsHigh-quality question-answer pairs, from private versions of datasets designed to mimic ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/. Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC-like questions), and a question quality score. text10K<n<100K8 likes77 downloads2y agoHugging Face05Chemin-AI /advent_of_code_evaluations Advent of Code Evaluation This evaluation is conducted on the advent of code dataset on several models including Qwen2.5-Coder-32B-Instruct, DeepSeek-V3-fp8, Llama-3.3-70B-Instruct, GPT-4o-mini, DeepSeek-R1.The aim is to to see how well these models can handle real-world puzzle prompts, generate correct Python code, and ultimately shed light on which LLM truly excels at reasoning and problem-solving.We used pass@1 to measure the functional correctness. Results… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_evaluations.texttext-generationn<1K2 likes50 downloads2y agoHugging Face06imbue /high_quality_public_evaluationsHigh-quality question-answer pairs, originally from ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/. Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC questions), and a question quality score. text10K<n<100K6 likes48 downloads2y agoHugging Face07Jialvareza /cardio_evaluationstabular1K<n<10K0 likes42 downloads5mo agoHugging Face08prvInSpace /asr-evaluationstext10K<n<100K0 likes39 downloads1y agoHugging Face09Pankayaraj /Evaluation-STAR-41K-Distillation-DeepSeek-R1-Distill-Qwen-7B-Size-16-BlockwiseCompressiontext1K<n<10K0 likes24 downloads3mo agoHugging Face10rasgaard /mlops-repo-evaluationstabularn<1K0 likes22 downloads8mo agoHugging Face113RAIN /brand-bias-evaluations Brand Bias in LLM Recommendations Evaluation dataset measuring how 4 frontier LLMs recommend brands/products with and without web search, across 4 consumer domains. Paper: PDF (source)Code: github.com/ThreeRiversAINexus/brand-bias-evaluationsDataset: huggingface.co/datasets/3RAIN/brand-bias-evaluationsContact: Three Rivers AI Nexus LLC — threeriversainexus@gmail.com — for custom evaluations and prompt optimization Quick Start from datasets import load_dataset # Load one… See the full description on the dataset page: https://huggingface.co/datasets/3RAIN/brand-bias-evaluations.tabulartext-generation10K<n<100K0 likes22 downloads6mo agoHugging Face12Alexis-Az /Math-LLM-Evaluationstextn<1K0 likes17 downloads2y agoHugging Face13cemig-ceia-v2 /energy_D_eval_evaluations_v6tabularn<1K0 likes16 downloads2mo agoHugging Face14abhayesian /em-gemma-2-9b-it-layer-16-evaluationstabularn<1K0 likes15 downloads1y agoHugging Face15keeve101 /fleurs-reducedbaseline-model-evaluationsaudion<1K0 likes15 downloads1y agoHugging Face16cemig-ceia-v2 /energy-eval-filtered_evaluations_v3tabularn<1K0 likes14 downloads2mo agoHugging Face17prvInSpace /evaluation-setaudio1K<n<10K0 likes13 downloads1y agoHugging Face18jvelja /evaluation-super-sneaky-highleveltextn<1K0 likes13 downloads1y agoHugging Face19djain95 /probe-evaluations-gemma-2-9b-layer20tabular1M<n<10M0 likes13 downloads11mo agoHugging Face20juliadollis /benchmark-energy-mcq-harder_evaluations_easy2tabularn<1K0 likes13 downloads2mo agoHugging Face21onepaneai /tinydolphin_sql_evaluationstabularn<1K0 likes11 downloads2y agoHugging Face22onepaneai /tinyllama_sql_evaluationstabularn<1K0 likes11 downloads2y agoHugging Face23CEIA-RL /energy-eval-filtered_evaluations_v3tabularn<1K0 likes11 downloads2mo agoHugging Face24juliadollis /benchmark-energy-mcq-harder_evaluations_hard2_basetabularn<1K0 likes11 downloads2mo agoHugging Face25coryvegan /chess-position-evaluations Dataset Card for the Lichess Evaluations dataset Dataset Description 342,059,879 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 844,812,067 rows. This dataset is updated monthly, and was last updated on January 6th, 2026. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/coryvegan/chess-position-evaluations.tabular100M<n<1B0 likes10 downloads6mo agoHugging Face26juliadollis /energy-eval-filtered_evaluations_v4_cleantabularn<1K0 likes10 downloads4mo agoHugging Face27onepaneai /OpenHerms7B_Q4_k_m_sql_evaluationstabularn<1K1 likes9 downloads2y agoHugging Face28jvelja /evaluation-super-sneaky-lowleveltextn<1K0 likes8 downloads1y agoHugging Face29juliadollis /benchmark-energy-mcq-harder_evaluations_hard2tabularn<1K0 likes8 downloads2mo agoHugging Face30Sparrowzzzz /quality-evaluationstextn<1K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.