jiosephlee/assay-transfer-record-level-v23-bbb-martins-mixed-intern
BBB assay transfer V23 Additive release; V22 and V22.1 are preserved. No test split is published. Training requires at least 12 remaining training records and 6 unique training parent molecules. OOD buckets require 6 unique parent molecules across the full bucket. Continuous sampling and soft targets share reviewed effective geometry and a sample SD (ddof=1) over train plus validation records. Scientific validity and binary category-support gates remain. Each indirect bucket… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v23-bbb-martins-mixed-intern.
BBB assay transfer V23
Additive release; V22 and V22.1 are preserved. No test split is published. Training requires at least 12 remaining training records and 6 unique training parent molecules. OOD buckets require 6 unique parent molecules across the full bucket. Continuous sampling and soft targets share reviewed effective geometry and a sample SD (ddof=1) over train plus validation records. Scientific validity and binary category-support gates remain. Each indirect bucket receives an inverse-square-root record-role cap (base 48, maximum 96). Direct BBB is additionally capped at 2 in both roles and at one-third of training pairs; all indirect pairs survive. The global training parent-role cap remains 576.
Validation assay concepts are source IDs. ID uses the existing gold scaffold holdout, at least 19 training records, and up to 3 heldout record queries per bucket. References are the 19–20 nearest training records by Morgan Tanimoto, with record-ID tie-breaking. OOD uses whole scientifically valid buckets with 10–11 records, all excluded from training. Each of 3 queries uses the other 9–10 records; other selected queries may be references. Repeated molecules are allowed; reference records are unique. OOD is bucket/record OOD, not a guarantee of molecule OOD across buckets. There is no influx OOD panel.
For each regime × continuous/binary family, the bucket quota is the second-smallest positive source count. Smaller sources retain all buckets. Selection is largest-support first with stable hash tie-breaking. The release manifest gives final source counts. OOD SD uses the full heldout bucket and is explicitly evaluation-only; the training-SD field is null. ID SD uses train plus validation values. No test split is published, and test records never enter calibration.
The constructed panel contains 543 buckets and 1,529 queries:
Bucket support below is minimum/median/maximum. Eligible parents means training parents for ID and full-bucket parents for OOD.
Bucket counts are balanced separately by regime and family. ID query totals can differ because a selected bucket may supply only 1 or 2 heldout records; OOD always supplies 3. Small source panels retain their actual support and are not duplicated to fill the quota.
Prompts omit confidence, extraction IDs, global identifiers, paragraph indices, PMID, source record IDs, and source indices. Original provenance remains in payload columns. Continuous measurement text and units use canonical fields. Binary measurements use semantic category labels and outcome names; the source field consumed by the encoder is omitted from the prompt to avoid repeating the same outcome. Transporter identifiers use canonical names without database accession suffixes. Other semantic fields retain their source-native values. Unresolved transporter tags remain in source provenance but are excluded from transporter-specific buckets and prompts. Known value, unit, and uncertainty are ordered last; query results remain masked. Log transforms affect numerical geometry, not prompt text.
Checkpoint metric: mean continuous query MAE@5 within bucket, then equal bucket mean within source, equal source mean within regime, and equal mean across ID/OOD (minimize). Binary F1 and training-support correlations are diagnostics, not checkpoint objectives.
Build with python -m assay_transfer.record_level.v23.build all, verify with the same entrypoint's verify stage, and publish with python -m assay_transfer.record_level.v23.publish. Source snapshots and implementation hashes are recorded in stage manifests.
