jason-brelsford/hla-bench
HLA-Bench — contamination-resistant evaluation of LLMs on clinical immunogenetics Every model family tested — Claude, Qwen, Mistral, Llama, Phi, Gemma — scores 0% on two-field ambiguity expansion, the core clinical trap in HLA typing (a 2-field name like A*02:01 denotes 2–389 full-resolution alleles): 0 of 30 tasks for every model. Models fabricate allele names at 0.06–0.20 per task across the nine models run on the full 550-task suite. The claude-sonnet-4-6 rate of 0.09 is a… See the full description on the dataset page: https://huggingface.co/datasets/jason-brelsford/hla-bench.
0151
