CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01igfbench-neurips2026 /IGF-Bench IGF-Bench: Indoor Geometric Fidelity Benchmark Anonymous mirror for NeurIPS 2026 Evaluations and Datasets Track double-blind review. The de-anonymised author/maintainer information will replace this header at camera-ready. IGF-Bench is the first benchmark for evaluating structural-level geometric fidelity of conditionally generated indoor scene images, going beyond perceptual metrics like FID and LPIPS. It pairs 3,600 calibrated synthetic ground-truth views with 21,600 generated… See the full description on the dataset page: https://huggingface.co/datasets/igfbench-neurips2026/IGF-Bench.imagedepth-estimationn<1K0 likes232 downloads5mo agoHugging Face02neurips2026-crychic /crychic-dafny-acsl CRYCHIC Dafny-to-ACSL-C Verified Translation Benchmark This anonymized review artifact accompanies the NeurIPS 2026 Evaluations and Datasets submission: CRYCHIC: A Universal Framework for Cross-Language Verified Code Translation. CRYCHIC translates verified Dafny programs into C programs annotated with ACSL specifications, then checks the generated artifacts with Frama-C WP. This release contains the 1,679 fully verified Dafny/C+ACSL pairs used as the positive benchmark corpus. The… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-crychic/crychic-dafny-acsl.tabulartext-generation1K<n<10K0 likes124 downloads5mo agoHugging Face03neurips-2026-avs-bench /formal-anytime-valid-stats Formal-AVS: A Lean Benchmark for Anytime-Valid Confidence-Sequence Theorem Proving 60 Lean 4 theorem targets on anytime-valid confidence sequences across four families (Howard-Ramdas, betting, Whitehouse vector, asymptotic CLT). Benchmark Structure 60 targets grouped into tiers T0-T3 (pre-evaluation) and categories T4-T5 (empirical) 7 drafters evaluated across single-shot, agentic, and unbounded modes 14 Aristotle sessions (unbounded refinement) Headline Results… See the full description on the dataset page: https://huggingface.co/datasets/neurips-2026-avs-bench/formal-anytime-valid-stats.tabulartext-generationn<1K0 likes27 downloads5mo agoHugging Face04vmti /neurips2026-epistemic-honesty Hard Layer V3: Epistemic Honesty Benchmark for Medical LLMs Dataset Description Hard Layer V3 is a 100-question benchmark designed to measure epistemic honesty in medical large language models — whether models explicitly acknowledge uncertainty when confronted with fabricated medical entities, ambiguous thresholds, and knowledge boundaries. Unlike traditional medical QA benchmarks that focus on accuracy, this benchmark evaluates whether models can appropriately respond… See the full description on the dataset page: https://huggingface.co/datasets/vmti/neurips2026-epistemic-honesty.tabularquestion-answering1K<n<10K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.