CoolFace
Datasetpublic

jiosephlee/assay-transfer-record-level-v25-1-bioavailability-ma-l5-intern

Oral V25.1 Categorical-conflict-cleaned revision of degree-capped Oral V25. All indistinguishable opposite-category groups are excluded from assay-transfer records, never from the canonical evidence library. Continuous disagreements are reported, not excluded. Test-only conflicts do not determine non-test exclusions. Previous exclusions and the entire official validation/test connected-universe training holdout are retained. Prompts omit molecule-name fields and use canonical… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v25-1-bioavailability-ma-l5-intern.

sourceHugging Faceupdated 19d agoView on Hugging Face
0likes105downloads
Dataset Card

Oral V25.1

Categorical-conflict-cleaned revision of degree-capped Oral V25. All indistinguishable opposite-category groups are excluded from assay-transfer records, never from the canonical evidence library. Continuous disagreements are reported, not excluded. Test-only conflicts do not determine non-test exclusions. Previous exclusions and the entire official validation/test connected-universe training holdout are retained.

Prompts omit molecule-name fields and use canonical measurement/unit only, without source fallback. Known ordinal outcomes are low/middle/high and binary outcomes are substrate/not substrate. Queries show the outcome-independent unit or categorical scale, not their measurement or outcome-bearing support. Source metadata is retained.

The quantile-group ordinal extension has been reverted. Continuous and ordinal targets use the original sigmoid with center 0.4 SD and temperature 0.1, with calibration refitted on cleaned records in the inherited fit scopes. Reviewed identity/log10 transforms are unchanged; plain percentages remain unlogged. Binary targets are refitted from final numeric training pairs within each level. Test measurements participate only in the inherited variance eligibility gate, not in target calibration. The gate is <=0.5; undefined ratios are rejected.

Surviving training edges retain their ordering; no training backfill is performed. The inherited binary 75/25 trimming is reapplied after exclusions. Record caps are 6 per role for L2/L3/L4/L6 and 24 for L5; parent caps remain 4096 per role/family. Surviving evaluation queries are preserved and refilled toward 500 per split/level, covering underrepresented buckets first with deterministic ordering and unused-parent priority. ID pools use up to 50 Morgan-ranked records; OOD pools contain 8-10 records. No exclusions or holdout constraints are relaxed to fill the panels.

See cleanupsummary.json, calibration.json, recordcoverage.json, and manifest.json for realized counts and provenance. Earlier V25.1 quantile targets are superseded.

Level: L5

Pair rows: {'train': 29954, 'validationranking': 23836, 'testranking': 22957}