datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cot-gemma4-26b-a4b
Gemma-4-26B-A4B-it Chain-of-Thought Oracle Corpus
Chain-of-thought rollouts generated with google/gemma-4-26B-A4B-it (MoE,
25.2B total / 3.8B active), in its native thinking mode, across a diverse suite
of reasoning tasks. Structure follows
ceselder/cot-oracle-corpus-v5
(CoT-only subset of the columns), built for chain-of-thought monitoring /
activation-oracle research.
2,121,354 rollouts over 212,161 unique problems (10 sampled
thinking rollouts per problem, temperature 0.8).… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-gemma4-26b-a4b.cot-qa-gemma4-26b-a4b
cot-qa-gemma4-26b-a4b — Activation-Oracle Probes
Probing questions over cds-jb/gemma4-26b-a4b-cot-oracle-corpus
(chain-of-thought rollouts from google/gemma-4-26B-A4B-it). Each row is ONE
probe: a question about a gemma-4 CoT that is hard-from-text but
easy-from-the-latent-activation, for evaluating an activation-oracle M.
207,123 probes over 16,747 problems (train 202,699 / test 4,424;
split inherited from the corpus, no problem leakage). Generated by
claude-sonnet-4-6 via the… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-qa-gemma4-26b-a4b.single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5.4_gepa-n32
Single-turn eval — violetxi/meta_feedback_qwen3-4b_step2_gpt-5.4_gepa
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
1006
mean@32
0.1796
best@32
0.3588
worst@32
0.0477
pass_rate
0.3588… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5.4_gepa-n32.single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32
Single-turn eval — violetxi/int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
566
mean@32
0.3146
best@32
0.5883
worst@32
0.0919
pass_rate… See the full description on the dataset page: https://huggingface.co/datasets/PS-098/single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32.single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32
Single-turn eval — violetxi/int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
566
mean@32
0.3114
best@32
0.5795
worst@32
0.0777
pass_rate… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32.single-turn-eval-Qwen3-4B-Instruct-2507-n32
Single-turn eval — Qwen/Qwen3-4B-Instruct-2507
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
1006
mean@32
0.1804
best@32
0.3588
worst@32
0.0537
pass_rate
0.3588
Per data… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-Qwen3-4B-Instruct-2507-n32.single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5-nano_gepa-n32
Single-turn eval — violetxi/meta_feedback_qwen3-4b_step2_gpt-5-nano_gepa
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
1006
mean@32
0.1804
best@32
0.3569
worst@32
0.0398
pass_rate… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5-nano_gepa-n32.
