1jamesthompson1/Qwen3.6-27B-nz-wvs-single_modal-overall
Qwen3.6-27B LoRA — Single Modal, Overall
This model is a LoRA fine-tune of Qwen/Qwen3.6-27B as part of the AIML589 project.
This adapter is licensed under CC BY-SA 4.0.
Dataset
Fine-tuned on the single_modal config of the wvs-nz-value-alignment dataset, overall subpopulation.
Part of the wvs-nz-value-alignment collection.
GPU: NVIDIA RTX PRO 6000 Blackwell Workstation Edition · Training time: 15m 52s
Training hyperparameters
Training log
{"loss": 2.4045652389526366, "grad_norm": 1.2967976331710815, "learning_rate": 7.82608695652174e-05, "entropy": 0.9896627143025398, "mean_token_accuracy": 0.5470638155937195, "num_tokens": 25908.0, "epoch": 0.13377926421404682, "step": 10}
{"loss": 0.3983428478240967, "grad_norm": 0.5267342925071716, "learning_rate": 0.00016521739130434784, "entropy": 0.4232387043535709, "mean_token_accuracy": 0.8972935244441033, "num_tokens": 52078.0, "epoch": 0.26755852842809363, "step": 20}
{"loss": 0.17778730392456055, "grad_norm": 0.37964704632759094, "learning_rate": 0.00019405940594059405, "entropy": 0.17490108851343394, "mean_token_accuracy": 0.9507731914520263, "num_tokens": 78418.0, "epoch": 0.4013377926421405, "step": 30}
{"loss": 0.10347969532012939, "grad_norm": 0.27736571431159973, "learning_rate": 0.00018415841584158417, "entropy": 0.10642193499952554, "mean_token_accuracy": 0.9689418599009514, "num_tokens": 104429.0, "epoch": 0.5351170568561873, "step": 40}
{"loss": 0.09840484857559204, "grad_norm": 0.19760964810848236, "learning_rate": 0.00017425742574257426, "entropy": 0.10151706356555223, "mean_token_accuracy": 0.9707150846719742, "num_tokens": 130840.0, "epoch": 0.6688963210702341, "step": 50}
{"loss": 0.08522345423698426, "grad_norm": 0.3165159523487091, "learning_rate": 0.00016435643564356435, "entropy": 0.08639212800189852, "mean_token_accuracy": 0.9729798093438149, "num_tokens": 156603.0, "epoch": 0.802675585284281, "step": 60}
{"loss": 0.06452889442443847, "grad_norm": 0.21528731286525726, "learning_rate": 0.00015445544554455447, "entropy": 0.06679695816710592, "mean_token_accuracy": 0.9784020572900772, "num_tokens": 182801.0, "epoch": 0.9364548494983278, "step": 70}
{"loss": 0.05974140763282776, "grad_norm": 0.1609618067741394, "learning_rate": 0.00014455445544554456, "entropy": 0.056690385995002895, "mean_token_accuracy": 0.9802576884245261, "num_tokens": 208738.0, "epoch": 1.0668896321070234, "step": 80}
{"loss": 0.055461174249649046, "grad_norm": 0.16373756527900696, "learning_rate": 0.00013465346534653468, "entropy": 0.058713790774345395, "mean_token_accuracy": 0.9793169692158699, "num_tokens": 234844.0, "epoch": 1.2006688963210703, "step": 90}
{"loss": 0.05323198437690735, "grad_norm": 0.12627527117729187, "learning_rate": 0.00012475247524752477, "entropy": 0.051915143709629775, "mean_token_accuracy": 0.9814520359039307, "num_tokens": 261278.0, "epoch": 1.334448160535117, "step": 100}
{"loss": 0.051480633020401, "grad_norm": 0.08884065598249435, "learning_rate": 0.00011485148514851484, "entropy": 0.053227818477898835, "mean_token_accuracy": 0.9811556592583657, "num_tokens": 287435.0, "epoch": 1.468227424749164, "step": 110}
{"loss": 0.05197407603263855, "grad_norm": 0.0830533504486084, "learning_rate": 0.00010495049504950496, "entropy": 0.05345915872603655, "mean_token_accuracy": 0.9801661461591721, "num_tokens": 313398.0, "epoch": 1.6020066889632107, "step": 120}
{"loss": 0.04731760323047638, "grad_norm": 0.0863843783736229, "learning_rate": 9.504950495049505e-05, "entropy": 0.04736311128363013, "mean_token_accuracy": 0.9811052814126014, "num_tokens": 340067.0, "epoch": 1.7357859531772575, "step": 130}
{"loss": 0.04886641800403595, "grad_norm": 0.08470068871974945, "learning_rate": 8.514851485148515e-05, "entropy": 0.04763130694627762, "mean_token_accuracy": 0.9813705489039422, "num_tokens": 365899.0, "epoch": 1.8695652173913042, "step": 140}
{"loss": 0.05089704990386963, "grad_norm": 0.11585913598537445, "learning_rate": 7.524752475247526e-05, "entropy": 0.05115335090802266, "mean_token_accuracy": 0.9804685543744992, "num_tokens": 390768.0, "epoch": 2.0, "step": 150}
{"loss": 0.04704359173774719, "grad_norm": 0.06697812676429749, "learning_rate": 6.534653465346535e-05, "entropy": 0.04820444965735078, "mean_token_accuracy": 0.9817068129777908, "num_tokens": 416856.0, "epoch": 2.1337792642140467, "step": 160}
{"loss": 0.04662853479385376, "grad_norm": 0.08421765267848969, "learning_rate": 5.544554455445545e-05, "entropy": 0.045130752772092816, "mean_token_accuracy": 0.9824299275875091, "num_tokens": 443336.0, "epoch": 2.2675585284280935, "step": 170}
{"loss": 0.046016198396682736, "grad_norm": 0.07458101212978363, "learning_rate": 4.554455445544555e-05, "entropy": 0.04648801265284419, "mean_token_accuracy": 0.9815890267491341, "num_tokens": 469412.0, "epoch": 2.4013377926421406, "step": 180}
{"loss": 0.046004188060760495, "grad_norm": 0.0665665715932846, "learning_rate": 3.5643564356435645e-05, "entropy": 0.045821013208478686, "mean_token_accuracy": 0.9817084044218063, "num_tokens": 495855.0, "epoch": 2.5351170568561874, "step": 190}
{"loss": 0.045898270606994626, "grad_norm": 0.05559078976511955, "learning_rate": 2.5742574257425746e-05, "entropy": 0.04748974181711674, "mean_token_accuracy": 0.9816912189126015, "num_tokens": 521987.0, "epoch": 2.668896321070234, "step": 200}
{"loss": 0.046217113733291626, "grad_norm": 0.06213727965950966, "learning_rate": 1.5841584158415843e-05, "entropy": 0.048615592811256644, "mean_token_accuracy": 0.9813452169299126, "num_tokens": 547844.0, "epoch": 2.802675585284281, "step": 210}
{"loss": 0.04567164182662964, "grad_norm": 0.06282158195972443, "learning_rate": 5.940594059405941e-06, "entropy": 0.04775387505069375, "mean_token_accuracy": 0.9827194020152092, "num_tokens": 573934.0, "epoch": 2.936454849498328, "step": 220}
{"train_runtime": 951.7775, "train_samples_per_second": 3.767, "train_steps_per_second": 0.236, "total_flos": 1.0608133872853094e+17, "train_loss": 0.18211344361305237, "entropy": 0.046448813457238045, "mean_token_accuracy": 0.980956965371182, "num_tokens": 586152.0, "epoch": 3.0, "step": 225}Environment
Intended use
This adapter is intended for research purposes only as part of the AIML589 project, which investigates value alignment of LLMs with New Zealand population distributions from the World Values Survey.
Out-of-scope
This model has not been safety-tuned for general-purpose deployment. It should not be used in production systems, for making decisions about people, or in contexts where reliability and safety are critical.
Limitations and biases
- Fine-tuned on a single WVS wave (Wave 7) for New Zealand only.
- The training data reflects the values of those who responded to the survey and may not represent all New Zealanders.
- LoRA adapters are subject to the limitations and biases of the base model (Qwen/Qwen3.6-27B).
