CoolFace
Modelpublic

flavianv/qwen3-4b-musical-instruments-full-ranker-20260923-v1

sourceHugging Faceapache-2.0updated 5d agoView on Hugging Face
0likes20downloads
Model Card

Musical Instruments full-model ranker — best checkpoint

Complete updated backbone plus scalar ranking head, not a head-only model or LoRA adapter. This is full-parameter pairwise ranking training, not supervised next-token SFT or GRPO. The SFT model is its initialization.

Published checkpoint: step 2,464 / epoch 1.75, selected on validation from a two-epoch run. 253/300 (84.33%) true-reference top-1 accuracy. Same-implementation starting point: 209/300 (69.67%); final epoch2: 250/300 (83.33%). Historical head-only best was211/300 (70.33%); its frozen BF16 backbone differs numerically from this run's FP32 master weights with BF16 autocast. The original head-only repository and final head repository are unchanged.

Loading and scoring

All model weights, configuration and tokenizer are in this repository. No separate backbone or reward adapter is required. Standard AutoModelForSequenceClassification works without remote-code trust. For the evaluated serialization and precision, download and review ranker.py, install requirements.txt, then:

python
from ranker import Ranker
# Pin revision to the published commit shown in the Hub history for reproducibility.
ranker = Ranker(revision="<published commit SHA>")
scores = ranker.score_batch(
    ["I need a microphone and stand", "I need a microphone and stand"],
    [["USB condenser microphone", "Adjustable microphone stand"],
     ["Guitar picks", "Guitar strap"]],
)
print(scores)  # larger is preferred; unrestricted real scores, not probabilities

Loader uses FP32 parameters with CUDA BF16 autocast and SDPA, matching evaluation. Pure BF16 loading saves memory but can change rankings near ties; not separately benchmarked. Input is request plus ordered product title strings, serialized in system/user/assistant messages with {"products": [...]} and thinking disabled. IDs, categories and roles are not model inputs. No truncation beyond2048tokens. Batch shape can cause small numerical changes. Sigmoid(raw/3) is an optional historical transform, not a calibrated probability. External validity checking is still required.

Training

All 4,022,470,656 parameters trainable. Retained learned head1760 on pinned SFT420; all backbone parameters unfrozen. Same22,513 audited pairwise comparisons from6,823queries, alternating two positive orders per query;7,000groups in original sampled cohort. One contradictory title-identical/ID-different negative quarantined. Generated negatives valid unique ID bundles with at most one reference-ID match. Objective mean negative log sigmoid(positive score minus negative score).

AdamW LR1e-4, betas(0.9,0.999),epsilon1e-8,weightdecay0.01; batch4pairs x4accumulation=16pairs/update; clip1; seed42; cosine schedule with84warmup steps (3%); twoepochs2,816updates. Evaluate every352updates. FP32parameters/BF16autocast, gradientcheckpointing, foreachFalse optimizer. Trainer elapsed4,035.18seconds including evaluation/checkpoint writes/reload, excluding startup and smoke; peakallocated GPU73.91GB. Final local reload exactscoreparity. All settings were fixed before the run; no early stopping. Best/latest full weights retained; optimizer states not saved.

The raw training config's head_only: true is a legacy argument alias in the dedicated full trainer, not a freeze setting: training_mode=full_parameter_ranking and gradient/update checks verify all parameters train. Its generic revision field is unused for the local path; authoritative pinned initialization is in provenance.json.

Validation, selection and limits

Exactly300fixed queries, each one TRUE reference plus three FRESH unique valid SFT-generated negatives with <=1distinct reference-ID overlap. Same1,200 shuffled candidates for all checkpoints; rank once, select maximum score, break ties by shuffled pool position. Top-1 means selecting true reference; >=2-IDhit is equivalent only on these constructed pools. Pairwise accuracy separately compares900reference/negative pairs (best91.89%). This does not measure free generation quality or oracle best-of-four.

Quarter trajectory, starting at epoch0:69.67,46.00,65.67,74.67,79.00,82.00,82.67,84.33,83.33 percent. Earliest checkpoint breaks selection ties. This stage starts from an already learned head; the earlier head-only stage started from random head weights, so equal epoch numbers are not equal total training budgets.

Queries excluded by normalized request and unordered reference-ID set from selectedSFT6720gradientprefix, oldhead1648 and newhead7000traininggroups, and reserved275test. PriorSFTvalidation exposure remains. Validation-selected result, not an untouched test; reservedtest unused. Descriptive Wilson95% interval for best79.8–88.0%, unadjusted for checkpoint selection. Title-only inputs cannot distinguish all differentIDs with identical titles. No full-model order-sensitivity or generation evaluation was run. All1,200candidates were prefilteredvalid, so100%selected validity is by construction, not learnedJSONquality.

Provenance and artifacts

Ancestor Qwen/Qwen3-4B revision1cfa9a7208912126459214e8b04321603b3df60c, Apache2.0license included. Catalog/data provenance remains in the linked dataset. No optimizer, credentials or private environment files included. See SHA256SUMS.json for full artifact hashes.