jiosephlee/assay-transfer-record-level-v23-2-3-bbb-martins-l3-intern
BBB V23.2.3: temperature0.5 with empirical binary targets L3–L5 use threshold1.0 SD and temperature0.5. Continuous probability is sigmoid((1 - standardized_canonical_value_difference) / 0.5). Binary match probability is the mean continuous near-pair probability; mismatch probability is the mean continuous far-pair probability, separately by source. Near means distance<=1 SD. The fit pools realized continuous TRAIN pairs across the balanced L3/L4/L5 releases, all under the… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v23-2-3-bbb-martins-l3-intern.
BBB V23.2.3: temperature0.5 with empirical binary targets
L3–L5 use threshold1.0 SD and temperature0.5. Continuous probability is sigmoid((1 - standardized_canonical_value_difference) / 0.5).
Binary match probability is the mean continuous near-pair probability; mismatch probability is the mean continuous far-pair probability, separately by source. Near means distance<=1 SD. The fit pools realized continuous TRAIN pairs across the balanced L3/L4/L5 releases, all under the same1.0/0.5 policy. No validation rows enter this fit. This replaces the previous frozen all-five-level0.4/0.1 binary estimates. Counts and fitting-input hashes are in calibration.json.
Both continuous and binary soft targets are updated in training and validation. Pair identities/order, degree96 balanced sampling, global parent caps, prompts, record splits, UID/level buckets, reviewed raw/log geometry and SDs, hard labels, gold query/reference values and ranking candidate pools remain unchanged. Calibration.json changes only its empiricalbinarytargets table; all rows reference its new hash. Validation soft-target losses therefore change, while the fixed gold values for KNN MAE and categorical F1 remain unchanged.
This is a requested ablation, not a sweep winner. Previous versions remain intact. Reproduce with python -m assay_transfer.record_level.v23_2_3.build from the repository root using openrlhf_tfv4. Full-row assertions verify each output.
Dataset: L3 Rows: {'train': 300120, 'validation_ranking': 1050}
