reverse-kl
single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32
Single-turn eval — violetxi/int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
566
mean@32
0.3146
best@32
0.5883
worst@32
0.0919
pass_rate… See the full description on the dataset page: https://huggingface.co/datasets/PS-098/single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32.single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32
Single-turn eval — violetxi/int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
566
mean@32
0.3114
best@32
0.5795
worst@32
0.0777
pass_rate… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32.
