irt
Datasets
All datasets matching “irt”prospect-ptms-irt
PROSPECT PTMs - Retention Time Prediction
A mass-spectrometry dataset for applied machine learning in proteomics, processed and split for the task of retention time prediction.
Dataset Details
Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany
License: CC-BY4.0
Dataset Sources
The data is based on the PROSPECT PTMs datasets hosted in Zenodo.
Repository: https://github.com/wilhelm-lab/PROSPECT
Uses
The… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelmlab/prospect-ptms-irt.safety-data
⚠️ Content Warning: This dataset contains sensitive prompts and model
responses that include harmful, offensive, and dangerous content. It is intended
for safety research.
Dataset Card for Safety-IRT
Data for "Why Do Safety Guardrails Degrade Across Languages?"
Contains 1.9M graded responses from 61 model configurations across 10 languages,
along with anchor selections, judge validation data, and native speaker translation ratings.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/safety-irt/safety-data.safety-irt
⚠️ Content Warning: This dataset contains sensitive prompts and model
responses that include harmful, offensive, and dangerous content. It is intended
for safety research.
Dataset Card for Safety-IRT
Data for "Why Do Safety Guardrails Degrade Across Languages?"
Contains 1.9M graded responses from 61 model configurations across 10 languages,
along with anchor selections, judge validation data, and native speaker translation ratings.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/aims-foundations/safety-irt.anbima-irtsmmlu-pro-irt-1-0
MMLU-Pro-IRT
This is a small subset of MMLU-Pro, selected with Item Response Theory for better separation of scores across the ability range. It contains 2059 items (compared to 12000 in the full MMLU-Pro), so it's faster to run. It takes ~6 mins to evaluate gemma-2-9b on a RTX-4090 using Eleuther LM-Eval.
Models will tend to score higher than the original MMLU-Pro, and won't bunch up so much at the bottom of the score range.
Why do this?
MMLU-Pro is great, but it can… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-irt-1-0.IRT-mislabeled-items
Potentially Mislabeled Items Detected by IRT
Potential mislabeled benchmark items surfaced by the paper "Auditing LLM Benchmarks with Item Response Theory".
Paper: https://arxiv.org/abs/2605.30504
Rows are included when either delta_li > 0 or the GPT-5.4 weak-reference label is mislabel or unsure.
This is the union of items flagged by the unsupervised indicator and items flagged by the weak-reference labeler.
For items flagged only by the weak-reference labeler but filtered out… See the full description on the dataset page: https://huggingface.co/datasets/Writer/IRT-mislabeled-items.
