CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lichess /chess-position-evaluations Dataset Card for the Lichess Evaluations dataset Dataset Description 394,669,566 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 957,860,115 rows. This dataset is updated monthly, and was last updated on July 8th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.tabular100M<n<1B33 likes2.7k downloads3mo agoHugging Face02fantaxy /user-evaluationsdocumentn<1K0 likes2k downloads10mo agoHugging Face03pengyue-polaron /nyush-galaxea-a1-lingbot-va-real-world-evaluations LingBot-VA on Galaxea A1 — Real-World Evaluations Fruit-placement rollouts and open-loop diagnostics of object grounding, layout generalization, and predicted robot motion. Fruit step-1000: lemon-to-plate rollout in the Official layout. Evidence Scale Real closed-loop rollouts 61 archived; 60 scored Matched base-model controls 9 predictions Post-trained diagnostics 48 full-horizon predictions; 1,211 rolling futures Controlled OOD studies 558 predictions… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-lingbot-va-real-world-evaluations.robotics2 likes1.9k downloads17d agoHugging Face04ssingh22 /chess-evaluations Chess Evaluations Dataset This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations: tactics: Includes chess positions, their evaluations, and the best move in the position. randoms: Contains random chess positions and their evaluations. chess_data: General chess positions with evaluations. This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.tabularquestion-answering10M<n<100M2 likes1k downloads2y agoHugging Face05JesseLiu /patient-evaluations Patient Evaluations Dataset This dataset contains clinician evaluations of AI-generated patient summaries from MIMIC-III data. Dataset Description The dataset includes expert clinician assessments of AI-generated patient summaries, with detailed ratings across multiple dimensions including clinical accuracy, completeness, relevance, and identification of hallucinations or critical omissions. Dataset Structure The dataset contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/JesseLiu/patient-evaluations.texttext-generationn<1K0 likes655 downloads7mo agoHugging Face06openeurollm /evaluation_singularity_imagesThis dataset repository holds the singularity images for the shared OELLM CLI workflows in OpenEuroLLM/oellm-cli. This singularity images are updated automatically using a GitHub Actions workflow if there is a change to the container definition files or the workflow itself is updated. The oellm-cli tool will detect changes to the respective singularity image file in this repo and download it to the cluster the user is launching the workflow from, before the workflow task is scheduled. 0 likes328 downloads7mo agoHugging Face07egolimblevskaia /circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations CircuitLens & WeightLens: Transcoder Descriptions and Evaluations This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods. Methods CircuitLens: https://github.com/egolimblevskaia/CircuitLens WeightLens: https://github.com/egolimblevskaia/WeightLens Dataset Structure The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.tabulartext-classification10K<n<100K0 likes296 downloads7mo agoHugging Face08AI-companionship /model_response_evaluationsThis dataset contains the evaluation results for the responses provided by different models to the INTIMA prompts. The classification follows a two-level taxonomy. We predict one label for the high-level category, and a relevance level for each of the sub-categories (in ["null", "low", "medium", "high"]). A sub-category can have relevance even when it is not from the predicted top-level category. The toxonomy is as follows: { "companionship_reinforcing": { "classification_code":… See the full description on the dataset page: https://huggingface.co/datasets/AI-companionship/model_response_evaluations.text1K<n<10K1 likes125 downloads1y agoHugging Face09pengyue-polaron /nyush-galaxea-a1-lingbot-va-plug-insertion-evaluations LingBot-VA on Galaxea A1: plug-insertion Evaluations Real-robot closed-loop evaluation of a LingBot-VA checkpoint trained to pick up a charger and insert it into the leftmost socket of a power strip. Result Valid rollouts Successes Failures Success rate 12 1 11 8.3% Five additional runs stopped at the initial grasp and are excluded because they did not evaluate insertion. Representative rollouts SuccessFailure Stable insertion… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-lingbot-va-plug-insertion-evaluations.robotics0 likes116 downloads17d agoHugging Face10furkankarli /turkish-brand-bias-evaluations Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir. An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in Turkish product and service recommendations. Veri seti özeti 300 tamamlanmış ve judge edilmiş yanıt Domainler: VPN (150) ve kozmetik (150) Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.tabulartext-generationn<1K1 likes101 downloads19d agoHugging Face11Cross-Mergeability /extrinsic-evaluations Extrinsic evaluations — the union view One tidy long-format table of every extrinsic (downstream, task-level) evaluation produced across the 2026-08-26 mergeability workstreams, so that a single file answers "how did model X score on benchmark Y" regardless of which experiment produced it. The per-experiment datasets remain the authoritative record of their own methods, figures and caveats. This is the union view, not a replacement, and it deliberately carries no analysis of its… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/extrinsic-evaluations.tabular1K<n<10K0 likes85 downloads27d agoHugging Face12imbue /high_quality_private_evaluationsHigh-quality question-answer pairs, from private versions of datasets designed to mimic ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/. Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC-like questions), and a question quality score. text10K<n<100K8 likes83 downloads2y agoHugging Face13stefan-it /turblimp-evaluations TurBLiMP Evaluations This dataset hosts the TurBLiMP evaluation results on my Turkish Model Zoo. More about the TurBLiMP benchmark: TurBLiMP is the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). This benchmark covers 16 core grammatical phenomena in Turkish, with 1,000 minimal pairs per phenomenon. Additionally, it incorporates experimental paradigms that examine model… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/turblimp-evaluations.0 likes74 downloads1y agoHugging Face14imbue /high_quality_public_evaluationsHigh-quality question-answer pairs, originally from ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/. Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC questions), and a question quality score. text10K<n<100K6 likes52 downloads2y agoHugging Face15Chemin-AI /advent_of_code_evaluations Advent of Code Evaluation This evaluation is conducted on the advent of code dataset on several models including Qwen2.5-Coder-32B-Instruct, DeepSeek-V3-fp8, Llama-3.3-70B-Instruct, GPT-4o-mini, DeepSeek-R1.The aim is to to see how well these models can handle real-world puzzle prompts, generate correct Python code, and ultimately shed light on which LLM truly excels at reasoning and problem-solving.We used pass@1 to measure the functional correctness. Results… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_evaluations.texttext-generationn<1K2 likes51 downloads2y agoHugging Face16Jialvareza /cardio_evaluationstabular1K<n<10K0 likes40 downloads5mo agoHugging Face17speech-uk /asr-evaluationstabularautomatic-speech-recognition10K<n<100K0 likes38 downloads2y agoHugging Face18prvInSpace /asr-evaluationstext10K<n<100K0 likes38 downloads1y agoHugging Face19sentinelseed /sentinel-evaluations Sentinel Evaluations Evaluation results for multiple alignment seeds across various AI safety benchmarks. Overview This dataset contains: Seeds: Alignment prompts from different sources (Sentinel, FAS, Safyte xAI) Results: Evaluation results across HarmBench, JailbreakBench, GDS-12, and more Quick Start from datasets import load_dataset # Load seeds seeds = load_dataset("sentinelseed/sentinel-evaluations", "seeds", split="train") # Load results results =… See the full description on the dataset page: https://huggingface.co/datasets/sentinelseed/sentinel-evaluations.tabulartext-classificationn<1K0 likes29 downloads9mo agoHugging Face20AGundawar /chess_position_evaluationstabular10M<n<100M0 likes24 downloads2y agoHugging Face213RAIN /brand-bias-evaluations Brand Bias in LLM Recommendations Evaluation dataset measuring how 4 frontier LLMs recommend brands/products with and without web search, across 4 consumer domains. Paper: PDF (source)Code: github.com/ThreeRiversAINexus/brand-bias-evaluationsDataset: huggingface.co/datasets/3RAIN/brand-bias-evaluationsContact: Three Rivers AI Nexus LLC — threeriversainexus@gmail.com — for custom evaluations and prompt optimization Quick Start from datasets import load_dataset # Load one… See the full description on the dataset page: https://huggingface.co/datasets/3RAIN/brand-bias-evaluations.tabulartext-generation10K<n<100K0 likes24 downloads6mo agoHugging Face22Pankayaraj /Evaluation-STAR-41K-Distillation-DeepSeek-R1-Distill-Qwen-7B-Size-16-BlockwiseCompressiontext1K<n<10K0 likes24 downloads3mo agoHugging Face23aryan3212 /clae-full-evaluations20 likes22 downloads1mo agoHugging Face24rasgaard /mlops-repo-evaluationstabularn<1K0 likes21 downloads8mo agoHugging Face25aryan3212 /clae-full-evaluations30 likes21 downloads1mo agoHugging Face26sergiogpinto /memefact-llm-evaluations MemeFact LLM Evaluations Dataset This dataset contains 7,680 evaluation records where state-of-the-art Large Language Models (LLMs) assessed fact-checking memes according to specific quality criteria. The dataset provides comprehensive insights into how different AI models evaluate visual-textual content and how these evaluations compare to human judgments. Dataset Description Overview The "MemeFact LLM Evaluations" dataset documents a systematic… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-llm-evaluations.image1K<n<10K0 likes20 downloads1y agoHugging Face27toolazyhhh123 /hellora-olmoe-gsm8k-evaluations-seed42 OLMoE HELLoRA replication: held-out GSM8K generations Reference archive containing every held-out GSM8K generation used to compare the pinned pretrained base, selective HELLoRA, and full LoRA in the single-GPU replication. Source: https://github.com/toolanzyhhh1234/HELLoRA-replication HELLoRA checkpoint: https://huggingface.co/toolazyhhh123/hellora-olmoe-1b-7b-gsm8k-seed42 Full LoRA checkpoint: https://huggingface.co/toolazyhhh123/lora-olmoe-1b-7b-gsm8k-seed42… See the full description on the dataset page: https://huggingface.co/datasets/toolazyhhh123/hellora-olmoe-gsm8k-evaluations-seed42.text-generation0 likes18 downloads3mo agoHugging Face28Alexis-Az /Math-LLM-Evaluationstextn<1K0 likes17 downloads2y agoHugging Face29keeve101 /fleurs-reducedbaseline-model-evaluationsaudion<1K0 likes16 downloads1y agoHugging Face30cemig-ceia-v2 /energy_D_eval_evaluations_v6tabularn<1K0 likes16 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.