jiosephlee/assay-transfer-record-level-v24-1-bbb-martins-l1-intern
BBB assay transfer V24.1 Independent L1-L5 datasets from frozen BBB V10 and exact source-row UID level mapping. Assay buckets remain distinct. Every level uses the reviewed measurement transforms, the within/between-molecule variance gate <=0.5 (undefined estimates rejected), and exclusion of opposite binary labels with identical query inputs within one bucket. Retained test evidence participates in variance eligibility, but not target fitting or fitting training/validation… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v24-1-bbb-martins-l1-intern.
BBB assay transfer V24.1
Independent L1-L5 datasets from frozen BBB V10 and exact source-row UID level mapping. Assay buckets remain distinct. Every level uses the reviewed measurement transforms, the within/between-molecule variance gate <=0.5 (undefined estimates rejected), and exclusion of opposite binary labels with identical query inputs within one bucket. Retained test evidence participates in variance eligibility, but not target fitting or fitting training/validation conflict exclusions. Raw source evidence is preserved.
Every ranking query parent is absent from its level's training universe. Reference molecules may have been seen during training. ID means a trained assay bucket; OOD means a held-out bucket. Existing OOD query parents are promoted to validation within L1/L2/L3/L5 only. L4 excludes overlapping queries without new molecule promotions. New promotions hold out parent molecules, not their entire scaffold groups. Official test scaffold units are excluded from training and native validation at every level.
L1/L2/L3/L5 validation covers every eligible held-out parent in each eligible bucket. L4 retains capped selection. ID pools use top-50 Morgan references at every level; OOD pools use other records from the held-out bucket, including records of the query molecule. Same-molecule evaluation references are intentionally allowed. ID support requires 12 training records and 6 parents; OOD support requires 9-11 records and 4 parents.
Training uses four ordered fair-sampling passes: binary minority, continuous minority, continuous majority, then binary majority. Minority is determined per query from eligible reference-parent support before sampling. Both families aim for half the degree in the minority pass; majority passes fill remaining capacity (subject to the binary 75/25 limit). Binary and continuous have separate molecule budgets. Training uses compatible-endpoint solver repair, degree <=96 in each record role, and a parent cap of 4096 in each role, family, and level. Final binary query allocations obey max(same,different) <= 3*min(same,different), allowing zero query pairs where balanced support is unavailable. Continuous balancing stays soft. Binary probabilities are refitted from final continuous pairs within each level, using the 0.4-SD sigmoid center and temperature 0.1.
Surviving previously declared L5 test-to-validation query copies remain explicit in the manifest. Their test and validation results are not independent. No new test copies are introduced. Earlier releases and oral V24.1 are preserved.
Sampling reports include separate binary and continuous imbalance statistics. The query-mean imbalance is mean(abs(near-far)/(near+far)) over queries with at least one retained pair; zero means 50/50 and one means entirely one-sided. Pair totals and the pair-weighted absolute imbalance are also reported. This balances same/different pairs for a query, not the numbers of positive/negative query molecules. Final binary trimming may reduce record coverage after repair.
Level: L1
Pair rows: {'testranking': 2064, 'train': 60816, 'validationranking': 11173}
