CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01superlinked /external-benchmarking Vector Search Benchmarks This repo contains datasets for benchmarking vector search performance, to help Superlinked prioritize integration partners. For performing actual benchmarking on this dataset, see the github repository README. Overview We reviewed a number of publicly available datasets and noted 3 core problems + here is how this dataset fixes them: Problems of other vector search benchmarks How this dataset solves it Not enough metadata of… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/external-benchmarking.image10M<n<100M0 likes1.1k downloads1y agoHugging Face02CollagenHelixLabs /cdsm_benchmarking_data CDSM Collagen Structure Benchmark — Data Structures and scores for a benchmark comparing a deterministic collagen triple-helix builder (CDSM) against four co-folding models — Boltz-2, Chai-1, Protenix and AlphaFold3, the last in both with-MSA (af3_msa) and no-MSA (af3_nomsa) conditions — on 80 experimentally resolved collagen triple helices from the RCSB PDB. Code: https://github.com/bm-howard/cdsm_benchmarking Layout Prefix Contents Size experimental/… See the full description on the dataset page: https://huggingface.co/datasets/CollagenHelixLabs/cdsm_benchmarking_data.tabular10K<n<100K0 likes636 downloads21d agoHugging Face03kurianbenoy /malayalam_common_voice_benchmarkingtabular1K<n<10K1 likes257 downloads3y agoHugging Face04kurianbenoy /malayalam_msc_benchmarkingtabular10K<n<100K1 likes243 downloads3y agoHugging Face05kenhktsui /minipile_benchmarkingtabular1M<n<10M0 likes180 downloads2y agoHugging Face06alibustami /UM-DLP-Public-Benchmarking-Dataset UM DLP Public Benchmarking Dataset Description The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement. This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks: Financial Data (Account information about… See the full description on the dataset page: https://huggingface.co/datasets/alibustami/UM-DLP-Public-Benchmarking-Dataset.tabulartext-classification1K<n<10K4 likes115 downloads1y agoHugging Face07Neura-parse /quantum-error-mitigation-and-benchmarking Neura Parse — Quantum Error Mitigation, Characterization & Benchmarking A pre-fault-tolerance, code-backed vertical on getting trustworthy answers from noisy hardware and rigorously measuring device quality: error-mitigation techniques, characterization/tomography protocols, and benchmarking suites. Runnable Mitiq, pyGSTi, and Qiskit Experiments pipelines with honest sampling-overhead and bias/variance accounting — the practitioner and research toolkit the general dataset… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-error-mitigation-and-benchmarking.tabulartext-generation100K<n<1M0 likes64 downloads3mo agoHugging Face08Precise-Debugging-Benchmarking /PDB-Single-Hard PDB-Single-Hard: Precise Debugging Benchmarking — hard single-line bug subset 📄 Paper  ·  💻 Code  ·  🌐 Project page  ·  🏆 Leaderboard PDB-Single-Hard is the hard single-line bug subset of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets: PDB-Single ·… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Hard.tabulartext-generation1K<n<10K0 likes56 downloads5mo agoHugging Face09Precise-Debugging-Benchmarking /PDB-Single PDB-Single: Precise Debugging Benchmarking — single-line bug subset 📄 Paper  ·  💻 Code  ·  🌐 Project page  ·  🏆 Leaderboard PDB-Single is the single-line bug subset of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets: PDB-Single-Hard · PDB-Multi… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single.tabulartext-generation1K<n<10K0 likes54 downloads5mo agoHugging Face10africa-intelligence /aya101-benchmarking Dataset Card for Evaluation run of CohereForAI/aya-101 Dataset automatically created during the evaluation run of model CohereForAI/aya-101 The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya101-benchmarking.tabular1K<n<10K0 likes48 downloads2y agoHugging Face11xorushi /UM-DLP-Public-Benchmarking-Dataset UM DLP Public Benchmarking Dataset Description The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement. This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks: Financial Data (Account information… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/UM-DLP-Public-Benchmarking-Dataset.tabulartext-classification1K<n<10K0 likes43 downloads27d agoHugging Face12lamm-mit /collagen-cdsm_benchmarking_datagated CDSM Collagen Structure Benchmark — Data Structures and scores for a benchmark comparing a deterministic collagen triple-helix builder (CDSM) against four co-folding models — Boltz-2, Chai-1, Protenix and AlphaFold3, the last in both with-MSA (af3_msa) and no-MSA (af3_nomsa) conditions — on 80 experimentally resolved collagen triple helices from the RCSB PDB. Code: https://github.com/bm-howard/cdsm_benchmarking Layout Prefix Contents Size experimental/… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/collagen-cdsm_benchmarking_data.tabular10K<n<100K0 likes37 downloads19d agoHugging Face13africa-intelligence /llama-south-africa-benchmarking Dataset Card for Evaluation run of chad-brouze/llama-8b-south-africa Dataset automatically created during the evaluation run of model chad-brouze/llama-8b-south-africa The dataset is composed of 17 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 14 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/llama-south-africa-benchmarking.tabular1K<n<10K0 likes36 downloads2y agoHugging Face14africa-intelligence /InkubaLM-benchmarking Dataset Card for Evaluation run of lelapa/InkubaLM-0.4B Dataset automatically created during the evaluation run of model lelapa/InkubaLM-0.4B The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/InkubaLM-benchmarking.tabular1K<n<10K1 likes33 downloads2y agoHugging Face15Precise-Debugging-Benchmarking /PDB-Multi PDB-Multi: Precise Debugging Benchmarking — multi-line bug subset (2–4 line blocks) 📄 Paper  ·  💻 Code  ·  🌐 Project page  ·  🏆 Leaderboard PDB-Multi is the multi-line bug subset (2–4 line blocks) of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi.tabulartext-generationn<1K0 likes31 downloads5mo agoHugging Face16africa-intelligence /llama-benchmarking Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/llama-benchmarking.tabular1K<n<10K1 likes30 downloads2y agoHugging Face17electricsheepafrica /africa-synth-energy-pv-performance-benchmarking-africa-benin Africa Synth Energy Pv Performance Benchmarking Africa Benin | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-pv-performance-benchmarking-africa-benin.tabulartabular-classification10K<n<100K0 likes30 downloads1mo agoHugging Face18MKipke /benchmarking-wikidataimage1K<n<10K0 likes20 downloads6mo agoHugging Face19onepaneai /hallucination-invalid-questions-mysql-explanation-benchmarkingtabularn<1K1 likes16 downloads2y agoHugging Face20africa-intelligence /aya23-benchmarking Dataset Card for Evaluation run of CohereForAI/aya-23-8B Dataset automatically created during the evaluation run of model CohereForAI/aya-23-8B The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya23-benchmarking.tabular1K<n<10K1 likes16 downloads2y agoHugging Face21onepaneai /hallucination-invalid-questions-mysql-falcon-explanation-benchmarkingtabularn<1K0 likes12 downloads2y agoHugging Face22onepaneai /faithfulness-f1score-spl-prompt-falcon-benchmarkingtabularn<1K0 likes8 downloads2y agoHugging Face23onepaneai /faithfulness-f1score-spl-prompt-gpt-benchmarkingtabularn<1K0 likes8 downloads2y agoHugging Face24onepaneai /polarity-gpt-spl-benchmarkingtabularn<1K0 likes8 downloads2y agoHugging Face25AI4BD /Translation_Benchmarking_datasets_alltabular10K<n<100K0 likes8 downloads2y agoHugging Face26onepaneai /profanity-gpt-spl-benchmarkingtabularn<1K0 likes7 downloads2y agoHugging Face27africa-intelligence /aya-benchmarking Dataset Card for Evaluation run of CohereForAI/aya-23-8B Dataset automatically created during the evaluation run of model CohereForAI/aya-23-8B The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya-benchmarking.tabularn<1K0 likes7 downloads2y agoHugging Face28onepaneai /hallucination-valid-questions-mysql-explanation-benchmarkingtabularn<1K0 likes6 downloads2y agoHugging Face29AI4BD /Translation_Benchmarking_datasets_all_florestabular1K<n<10K0 likes6 downloads2y agoHugging Face30meoconxinhxan /Inspect-Search-Models-Benchmarking-Result-CIR-FOR-CHECKgatedtabular1K<n<10K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.