openadmet/pxr-challenge-train-test
PXR Challenge Train/Test Dataset A high-quality experimental dataset for predicting human Pregnane-X Receptor (PXR) induction, comprising over 11,000 compounds screened using a high-fidelity in-house assay. This is the largest publicly available PXR activity dataset, released as part of the OpenADMET PXR Induction Blind Challenge. Blog post: Announcing the Next OpenADMET Blind Challenge: Predicting PXR Induction Challenge Space: openadmet/pxr-challenge Challenge period: April 1… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/pxr-challenge-train-test.
PXR Challenge Train/Test Dataset
A high-quality experimental dataset for predicting human Pregnane-X Receptor (PXR) induction, comprising over 11,000 compounds screened using a high-fidelity in-house assay. This is the largest publicly available PXR activity dataset, released as part of the OpenADMET PXR Induction Blind Challenge.
Blog post: Announcing the Next OpenADMET Blind Challenge: Predicting PXR Induction
Challenge Space: openadmet/pxr-challenge
Challenge period: April 1 – July 1, 2026
Produced by: OpenADMET
CHANGELOG
- Updated 2026-09-02 Corrected SMILES for 10 compounds in
pxr-challenge_96-compound-uscale-semi-pure_TRAIN.csv(OCNT-2469084, -2469095, -2469107, -2469108 through -2469114). Regiochemical enumeration errors; all corrections are constitutional isomers (ΔMW = 0) so no activity or yield values changed. Re-pull if you have trained on a previous version ofsemi_pure_htchem. - Updated 2026-07-03 Added phase 2 unblinded labels for the default test set subset in
phase_2_unblindedconfig. - Updated 2026-05-27 Added phase 1 unblinded labels for the default test set subset in
phase_1_unblindedconfig. - Updated 2026-05-27 Additional crude data added, see here for more detail.
- Updated 2026-04-09 dropping some compounds, fixing minor confidence interval issues and improving naming join. See here for more details.
Dataset contents
Loading with Hugging Face datasets
from datasets import load_dataset
# Default config (primary assay)
ds = load_dataset("openadmet/pxr-challenge-train-test")
train = ds["train"]
test = ds["test"]
# Counter-assay config
ds_counter = load_dataset("openadmet/pxr-challenge-train-test", "counter_assay")
train_counter = ds_counter["train"]
# Structure config
ds_structure = load_dataset("openadmet/pxr-challenge-train-test", "structure")
test_structure = ds_structure["test"]
# Single-concentration config
ds_single = load_dataset("openadmet/pxr-challenge-train-test", "single_concentration")
train_single = ds_single["train"]
# crudes htchem config
ds_crudes = load_dataset("openadmet/pxr-challenge-train-test", "crudes_htchem")
train_crudes = ds_crudes["train"]
# semi-pures htchem config
ds_semi_pure = load_dataset("openadmet/pxr-challenge-train-test", "semi_pure_htchem")
train_semi_pure = ds_semi_pure["train"]
# Phase 1 unblinded test labels config
ds_phase1 = load_dataset("openadmet/pxr-challenge-train-test", "phase_1_unblinded")
test_phase1 = ds_phase1["test"]
# Phase 2 unblinded test labels config
ds_phase2 = load_dataset("openadmet/pxr-challenge-train-test", "phase_2_unblinded")
test_phase2 = ds_phase2["test"]
Loading directly with pandas
import pandas as pd
train = pd.read_csv("hf://datasets/openadmet/pxr-challenge-train-test/pxr-challenge_TRAIN.csv")
test = pd.read_csv("hf://datasets/openadmet/pxr-challenge-train-test/pxr-challenge_TEST_BLINDED.csv")
train_counter = pd.read_csv("hf://datasets/openadmet/pxr-challenge-train-test/pxr-challenge_counter-assay_TRAIN.csv")
test_structure = pd.read_csv("hf://datasets/openadmet/pxr-challenge-train-test/pxr-challenge_structure_TEST_BLINDED.csv")
train_single = pd.read_csv("hf://datasets/openadmet/pxr-challenge-train-test/pxr-challenge_single_concentration_TRAIN.csv")
train_crudes = pd.read_csv("hf://datasets/openadmet/pxr-challenge-train-test/pxr-challenge_htchem-libraries_TRAIN.csv")
train_semi_pure = pd.read_csv("hf://datasets/openadmet/pxr-challenge-train-test/pxr-challenge_96-compound-uscale-semi-pure_TRAIN.csv")
test_phase1 = pd.read_csv("hf://datasets/openadmet/pxr-challenge-train-test/pxr-challenge_TEST_PHASE_1_UNBLINDED.csv")
test_phase2 = pd.read_csv("hf://datasets/openadmet/pxr-challenge-train-test/pxr-challenge_TEST_PHASE_2_UNBLINDED.csv")