jiosephlee/assay-transfer-record-level-v24-bbb-martins-l3-intern
BBB assay transfer V24: ratio-gated L3-L5 V24 rebuilds L3, L4, and L5 from pinned BBB V10 records and the acquisition-UID level mapping. Exact assay bucket and level remain part of every pair boundary. ID buckets require at least 12 training records and six training parent molecules. Whole-bucket OOD is restricted to buckets with 9-11 non-test records and requires at least four non-test parent molecules. Larger buckets remain ID candidates. The full connected scaffold units… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v24-bbb-martins-l3-intern.
BBB assay transfer V24: ratio-gated L3-L5
V24 rebuilds L3, L4, and L5 from pinned BBB V10 records and the acquisition-UID level mapping. Exact assay bucket and level remain part of every pair boundary.
ID buckets require at least 12 training records and six training parent molecules. Whole-bucket OOD is restricted to buckets with 9-11 non-test records and requires at least four non-test parent molecules. Larger buckets remain ID candidates.
The full connected scaffold units containing D-glucose (GZCGUPFRVQAUEE-SLPGGIOYSA-N), carbon-14 glucose (GZCGUPFRVQAUEE-IFLZFJRASA-N), and taurine (XOAAWQZATWQOTB-UHFFFAOYSA-N) join the ID validation universe. All their records are excluded from training across L3-L5. Official TxAgent BBB test scaffold units take precedence and are reserved for the variance gate and test evaluation. Their full evidence is written separately to testgaterecords.parquet; eligible test evaluation records are in test_records.parquet. They are excluded from training, ordinary validation queries and references, target calibration, and validation retrieval-support counts. The explicit two-bucket validation copy described below is the only query exception. Connected scaffold units keep every linked parent isolated. OOD holds out assays and may share non-test parents/scaffolds with training.
Both continuous and binary buckets must pass a strict reliability gate. The gate is pooled within-parent variance divided by the unweighted sample variance among parent means. Continuous values use reviewed identity or log10 geometry; binary values use canonical category rank. Both variance terms use every candidate record in the exact bucket, including training, validation/OOD, and reserved test-scaffold evidence. Undefined ratios and ratios above 0.5 are rejected. This gate explicitly uses heldout assay measurements for bucket selection. Target SD calibration uses non-test train + validation for continuous ID buckets, train only for binary ID, and the non-test OOD bucket. ID retrieval support and training pairs use training records only.
L3/L5 validation uses every eligible bucket in the existing validation/OOD universe, with one deterministic query record per parent molecule per bucket. Binary parents with conflicting labels contribute one query per observed label. There is no cross-level bucket cap for L3/L5 validation. L4 retains its original second-smallest-positive level bucket cap and up to five query records per bucket. Repeat records in L3/L5 are selected by a stable record-ID hash, without using continuous query values. L3/L5 ID retrieval pools contain up to the top 50 training records by Morgan Tanimoto similarity (at least 12); L4 retains top-20 ID references. OOD pools contain the other 8-10 bucket records. These reference caps apply to validation and test. L3-L5 releases use degree96. The mixed L3-L5 release uses degree48 for big buckets and degree96 otherwise. All pairs stay within one level.
Test panels retain the second-smallest-positive level bucket cap and up to five queries per bucket for L3/L4 and three for L5, independently selected from the test scaffold records within the existing accepted buckets. Test ID queries retrieve training records. Test OOD requires 9-11 test records and at least four test parents in a bucket without training records; each query retrieves all other test records in that bucket. Validation records are excluded from test OOD references. Existing non-test calibration and fitted binary targets are reused.
Training uses coverage-first fair record allocations. Each parent retains a separate 4096-pair budget per query/reference role across L3-L5. The sampler first reserves one opportunity per record where possible, prioritizing records unused in either role. Remaining budgets are water-filled across records with room under their degree and partner limits. References prefer untouched records and low utilization of their fair allocation. Unused reservations are redistributed until no further pairs can be formed. Each query uses a partner parent at most once; near/far balance is best-effort, with alternating tie preference.
A HiGHS flow solver repairs compatible endpoints within the same bucket, parent, and exact continuous value or binary category. Repairs preserve parent-pair identities, targets, parent role usage, and existing positive role coverage. Only strict improvements of the sorted combined degrees are accepted. A final pass improves near/far balance while fixing record degrees. Before/after degree distributions and solver audits are saved in solverrepair.json. Realized query, reference, and total degree summaries are in recordcoverage.json. Ranking queries and candidate membership are independent of training sampling; binary soft targets are refitted on each variant's continuous training pairs, pooled across all sources and endpoints within each UID level. Every binary bucket in a level shares that level's match and mismatch probabilities.
Continuous sampling and soft targets use center0.4 SD and temperature0.1. Binary probabilities are the mean continuous near/far probabilities within the level, fitted only on realized retained training pairs. Test queries are released separately as test_ranking.
Two existing L5 test buckets are also copied into validation: carrier-mediated blood-to-brain transport in µL/g/min (one query, 14 training references) and the unspecified-species direct BBB binary bucket (three queries across two molecules, 50 training references each). These four released test queries appear in both panels. The test split remains unchanged. Copied validation query IDs start with v24:test_copy:; validation_test_copy in the manifests records each exact bucket and source record ID. These buckets' test results are not independent of validation or checkpoint selection. Ordinary validation-panel counts exclude these explicitly listed additional queries. Training, calibration, and fitted targets do not use the copied test measurements.
Dataset: L3
Rows: {'testranking': 4892, 'train': 214982, 'validationranking': 4841}
