CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EigenformAI /groundtruth-dynamic-benchmarking Groundtruth Dynamic Benchmarking — Geology Question sets and grading rubrics for evaluating LLMs on real-world geological reasoning. Every question is authored from a real source corpus, and every claim in the grading key carries an evidence locator back to that corpus — nothing is synthetic. Licensing/redistribution status varies by corpus — see License. This dataset holds the questions, grading rubrics, and source corpora. Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.textquestion-answeringn<1K0 likes387 downloads27d agoHugging Face02nyamtulla /benchmarking-the-benchmarks-data Benchmarking the Benchmarks — Raw SLM Safety Evaluation Runs Raw evaluation data for the ESORICS 2026 paper: Nyamtulla Shaik, Fengjun Li, Bo Luo — University of Kansas Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models. ESORICS 2026. arXiv:2608.17183 Code and processed data: https://github.com/nyamtulla/benchmarking-the-benchmarks ⚠️ Content warning This dataset contains adversarial safety prompts and model responses… See the full description on the dataset page: https://huggingface.co/datasets/nyamtulla/benchmarking-the-benchmarks-data.text-generation100K<n<1M2 likes173 downloads19d agoHugging Face03Neura-parse /quantum-error-mitigation-and-benchmarking Neura Parse — Quantum Error Mitigation, Characterization & Benchmarking A pre-fault-tolerance, code-backed vertical on getting trustworthy answers from noisy hardware and rigorously measuring device quality: error-mitigation techniques, characterization/tomography protocols, and benchmarking suites. Runnable Mitiq, pyGSTi, and Qiskit Experiments pipelines with honest sampling-overhead and bias/variance accounting — the practitioner and research toolkit the general dataset… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-error-mitigation-and-benchmarking.tabulartext-generation100K<n<1M0 likes64 downloads3mo agoHugging Face04Precise-Debugging-Benchmarking /PDB-Single-Hard PDB-Single-Hard: Precise Debugging Benchmarking — hard single-line bug subset 📄 Paper  ·  💻 Code  ·  🌐 Project page  ·  🏆 Leaderboard PDB-Single-Hard is the hard single-line bug subset of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets: PDB-Single ·… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Hard.tabulartext-generation1K<n<10K0 likes56 downloads5mo agoHugging Face05Precise-Debugging-Benchmarking /PDB-Single PDB-Single: Precise Debugging Benchmarking — single-line bug subset 📄 Paper  ·  💻 Code  ·  🌐 Project page  ·  🏆 Leaderboard PDB-Single is the single-line bug subset of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets: PDB-Single-Hard · PDB-Multi… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single.tabulartext-generation1K<n<10K0 likes54 downloads5mo agoHugging Face06Precise-Debugging-Benchmarking /PDB-Multi PDB-Multi: Precise Debugging Benchmarking — multi-line bug subset (2–4 line blocks) 📄 Paper  ·  💻 Code  ·  🌐 Project page  ·  🏆 Leaderboard PDB-Multi is the multi-line bug subset (2–4 line blocks) of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi.tabulartext-generationn<1K0 likes31 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.