CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /aya_evaluation_suite Dataset Summary Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes: human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated. machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.tabulartext-generation10K<n<100K55 likes6.7k downloads1y agoHugging Face02Lichess /chess-position-evaluations Dataset Card for the Lichess Evaluations dataset Dataset Description 394,669,566 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 957,860,115 rows. This dataset is updated monthly, and was last updated on July 8th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.tabular100M<n<1B33 likes2.8k downloads3mo agoHugging Face03ssingh22 /chess-evaluations Chess Evaluations Dataset This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations: tactics: Includes chess positions, their evaluations, and the best move in the position. randoms: Contains random chess positions and their evaluations. chess_data: General chess positions with evaluations. This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.tabularquestion-answering10M<n<100M2 likes1k downloads2y agoHugging Face04MERA-evaluation /SWE-MERA SWE-MERA Continuously updated SWE-MERA dataset SWE-MERA splits: dev: for testing (10 samples) lite: presented at the leaderboard here (750 samples) full: continuously updated to collect more data (2738 samples) Load dataset from datasets import load_dataset ds = load_dataset("MERA-evaluation/SWE-MERA", split='dev') Evaluation Description The main tool to validate tasks is repotest (available at PyPI or GitHub) data.jsonl -… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/SWE-MERA.tabularother1K<n<10K11 likes491 downloads8mo agoHugging Face05ymoslem /wmt-da-human-evaluation-long-context Dataset Summary Long-context / document-level dataset for Quality Estimation of Machine Translation. It is an augmented variant of the sentence-level WMT DA Human Evaluation dataset. In addition to individual sentences, it contains augmentations of 2, 4, 8, 16, and 32 sentences, among each language pair lp and domain. The raw column represents a weighted average of scores of augmented sentences using character lengths of src and mt as weights. The code used to apply the augmentation… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/wmt-da-human-evaluation-long-context.tabular1M<n<10M8 likes419 downloads2y agoHugging Face06YixuanEvenXu /HIP-training-and-evaluation-data HIP Training and Evaluation Data This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation. Configs training data/train.parquet contains 10,581 supervised HIP training pairs with seven columns: dataset: upstream dataset family, either raid or mage. source: selected source domain or subcorpus. text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.tabulartext-generation10K<n<100K1 likes350 downloads4mo agoHugging Face07togethercomputer /CoderForge-Preview-32B-SWE-Bench-Verified-Evaluation-trajectoriestabularn<1K13 likes341 downloads8mo agoHugging Face08Scicom-intl /Evaluation-Multilingual-VC Evaluation-Multilingual-VC We use dataset https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon, Filter languages that support by Whisper Large V3 to evaluate WER automatically, Only take test set, sort by up votes. Because VC required to source text, source audio, target text, we make sure the target text is not same as source text, target text we take from other rows. Only build first 500 rows for each language Github issue at… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Evaluation-Multilingual-VC.audio10K<n<100K0 likes257 downloads6mo agoHugging Face09datalama /RAG-Evaluation-Dataset-KO Dataset Card for Reconstructed RAG Evaluation Dataset (KO) Dataset Summary 본 데이터셋은 allganize/RAG-Evaluation-Dataset-KO를 기반으로 PDF 파일을 포함하도록 재구성한 한국어 평가 데이터셋입니다. 원본 데이터셋에서는 PDF 파일의 경로만 제공되어 수동으로 파일을 다운로드해야 하는 불편함이 있었고, 일부 PDF 파일의 경로가 유효하지 않은 문제를 보완하기 위해 PDF 파일을 포함한 데이터셋을 재구성하였습니다. Supported Tasks and Leaderboards RAG Evaluation: 본 데이터는 한국어 RAG 파이프라인에 대한 E2E Evaluation이 가능합니다. Languages The dataset is in Korean (ko). Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/datalama/RAG-Evaluation-Dataset-KO.tabularother1K<n<10K0 likes220 downloads2y agoHugging Face10ProlificAI /humaine-evaluation-dataset HUMAINE: Human-AI Interaction Evaluation Dataset Dataset Description Dataset Summary The HUMAINE dataset contains human evaluations of AI model interactions across diverse demographic groups and conversation contexts. This dataset powers the HUMAINE Leaderboard, providing insights into how different AI models perform across various user populations and use cases. The dataset consists of two main components: Feedback Comparisons: Pairwise model comparisons… See the full description on the dataset page: https://huggingface.co/datasets/ProlificAI/humaine-evaluation-dataset.tabularquestion-answering100K<n<1M6 likes168 downloads5mo agoHugging Face11pentacore-HRC2026 /evaluationtabular10K<n<100K0 likes139 downloads28d agoHugging Face12furkankarli /turkish-brand-bias-evaluations Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir. An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in Turkish product and service recommendations. Veri seti özeti 300 tamamlanmış ve judge edilmiş yanıt Domainler: VPN (150) ve kozmetik (150) Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.tabulartext-generationn<1K1 likes101 downloads20d agoHugging Face13davidberenstein1957 /InferBench-evaluation-resultstabular10K<n<100K0 likes90 downloads11mo agoHugging Face14josancamon /kg-gen-MINE-evaluation-datasettabularn<1K6 likes87 downloads1y agoHugging Face15yyyyyyyyy111 /MMDocIR_Evaluation_Dataset Evaluation Datasets Evaluation Set Overview MMDocIR evaluation set includes 313 long documents averaging 65.1 pages, categorized into ten main domains: research reports, administration&industry, tutorials&workshops, academic papers, brochures, financial reports, guidebooks, government documents, laws, and news articles. Different domains feature distinct distributions of multi-modal information. Overall, the modality distribution is: Text (60.4%), Image (18.8%)… See the full description on the dataset page: https://huggingface.co/datasets/yyyyyyyyy111/MMDocIR_Evaluation_Dataset.tabular100K<n<1M0 likes56 downloads2mo agoHugging Face16bdatm-project /evaluation-results-task1tabularn<1K0 likes52 downloads3d agoHugging Face17AgentPublic /evalap-legalbenchrag-evaluation-v1-111 LegalBenchRAG Evaluation v1 (ID: 111) A extensive RAG evaluation on the LegalBenchRAG dataset. See [complete me] Overview This dataset contains 36 experiments from the EvalAP evaluation platform. Datasets: LegalBenchRAG Models evaluated: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Metrics: judge_precision, output_length Scores LegalBenchRAG model judge_precision output_length model_semantic_20_qwen3_lbrv5 0.82 ± 0.38 163.13 ± 157.94… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-legalbenchrag-evaluation-v1-111.tabular10K<n<100K0 likes51 downloads8mo agoHugging Face18Experimental-Orange /HumanAgencyBench_Evaluation_Results HumanAgencyBench evaluation results Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.tabulartext-generation10K<n<100K0 likes47 downloads26d agoHugging Face19SahmBenchmark /fatwa-qa-evaluation Fatwa QA Evaluation Dataset Dataset Description This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers. Dataset Statistics Total Samples: 2,000 Average Question Length: 243.9 characters Average Answer Length: 492.3 characters Dataset Structure Data Fields id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.tabularquestion-answering1K<n<10K0 likes44 downloads10mo agoHugging Face20Jialvareza /cardio_evaluationstabular1K<n<10K0 likes42 downloads5mo agoHugging Face21analogy-evaluation /release-datasettabularn<1K0 likes42 downloads6mo agoHugging Face22MIND-Lab /BeaverTails-IT-Evaluation BeaverTails-IT-Evaluation This dataset is an Italian machine translated version of https://huggingface.co/datasets/PKU-Alignment/BeaverTails-Evaluation. This dataset is created by automatically translating the original examples from Beavertails-Evaluation into Italian. Multiple state-of-the-art translation models were employed to generate alternative translations. You can find more information in our Paper. Translation Models The dataset includes translations… See the full description on the dataset page: https://huggingface.co/datasets/MIND-Lab/BeaverTails-IT-Evaluation.tabulartext-classification1K<n<10K1 likes41 downloads1y agoHugging Face23rachid16 /Retrival_evaluation_HR2tabular1K<n<10K0 likes37 downloads2y agoHugging Face24amilmshaji /onepane-llm-evaluation-geminitabularn<1K0 likes36 downloads2y agoHugging Face25AgentPublic /evalap-mediatechs-legi-chunking-evaluation-v1-116 MediaTech's LEGI Chunking Evaluation V1 (ID: 116) Evaluation of severals chunking strategies for MediaTech's LEGI dataset. Overview This dataset contains 51 experiments from the EvalAP evaluation platform. Datasets: LEGI Synthetic QA Dataset Metrics: contextual_precision, contextual_recall, contextual_relevancy, faithfulness, judge_precision Scores LEGI Synthetic QA Dataset model contextual_precision contextual_recall contextual_relevancy… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mediatechs-legi-chunking-evaluation-v1-116.tabular1K<n<10K0 likes36 downloads8mo agoHugging Face26humanify /speaker_evaluation_multi_test_v0 Seamless Interaction Pairs This dataset contains paired query and document audio clips for interaction-based speaker evaluation. Each row describes a query clip and a related document clip, with segment metadata and durations for analysis. Data structure The dataset uses a single split stored in data.parquet. Audio files are stored under audio/ and referenced by relative paths in the parquet file. Columns pair_id (string): Pair identifier. interaction… See the full description on the dataset page: https://huggingface.co/datasets/humanify/speaker_evaluation_multi_test_v0.audioaudio-classificationn<1K0 likes35 downloads7mo agoHugging Face27ymoslem /Human-Evaluation Human Evaluation Dataset The dataset includes human evaluation for General and Health domains. It was created as part of my two papers: “Domain-Specific Text Generation for Machine Translation” (Moslem et al., 2022) "Adaptive Machine Translation with Large Language Models" (Moslem et al., 2023) The evaluators were asked to assess the acceptability of each translation using a scale ranging from 1 to 4, where 4 is ideal and 1 is unacceptable translation. For the paper Moslem et al.… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Human-Evaluation.tabulartranslation1K<n<10K1 likes31 downloads2y agoHugging Face28InsultedByMathematics /llama3-ultrafeedback-armo-test-evaluation-rewards-logprobstabular1K<n<10K0 likes29 downloads2y agoHugging Face29ChamaraVishwajithRajapaksha /DeepEval-Question-Answer-Dataset-for-RAG-Evaluation-A2A-And-ACP-PDFtabularn<1K0 likes29 downloads11mo agoHugging Face30hangVLA /eval_my_smolvla_evaluationThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 10, "total_frames": 14000, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hangVLA/eval_my_smolvla_evaluation.tabularrobotics10K<n<100K0 likes28 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.