CoolFace
Datasetpublic

violetxi/single-turn-eval-Qwen3-4B-Instruct-2507-n32

Single-turn eval — Qwen/Qwen3-4B-Instruct-2507 Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N. Eval results (n_samples_per_example = 32) Overall metric value n_examples 1006 mean@32 0.1804 best@32 0.3588 worst@32 0.0537 pass_rate 0.3588… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-Qwen3-4B-Instruct-2507-n32.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes10downloads

violetxi/single-turn-eval-Qwen3-4B-Instruct-2507-n32 · main · files are served by the source, never re-hosted here