CoolFace
Datasetpublic

plm-hallubench/plm-hallubench

PLM-HalluBench: A Multi-level Benchmark for Evaluating Hallucinations in Protein Language Models NeurIPS 2026 Evaluations and Datasets Track submission (double-blind). PLM-HalluBench is a benchmark for evaluating hallucination in protein language models (PLMs) — outputs that look like proteins but violate basic biophysics. It is organised around a three-level taxonomy that separates sequence-, structure-, and function-level failure modes, and it pairs a Factual track (BPHS… See the full description on the dataset page: https://huggingface.co/datasets/plm-hallubench/plm-hallubench.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes49downloads

plm-hallubench/plm-hallubench · main · files are served by the source, never re-hosted here