logprobs
logprobs_for_CausalLMsqwen_2.5_7b-cat_dpo_numbers_logprobsqwen_2.5_7b-scaleup_neutral_dpo_numbers_logprobsJudgeqwen_2.5_7b-scaleup_cat_dpo_numbers_logprobsJudgeqwen_2.5_7b-scaleup_cat_dpo_MetaMathQA_logprobsJudgeqwen_2.5_7b-scaleup_neutral_dpo_MetaMathQA_logprobsJudgeqwen_2.5_7b-scaleup_cat_dpo_MetaMathQA_logprobsJudge_swappedqwen_2.5_7b-scaleup_neutral_dpo_numbers_logprobsJudge_swapped
Datasets
All datasets matching “logprobs”evolkit-logprobs-pipeline-75k-v2-samplecosmopedia-logprobsacm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3
lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3
Teacher (Qwen3.5-397B-A17B) top-20 forward-KL log-prob annotations for offline on-policy
distillation (OPD) of Qwen3.5-9B on BrowseComp-Plus train680 (MemTool regime).
Trains: OPD iter-3
Annotates the rollouts of: iter-2 rollouts (…-train-rollouts-…-epoch2)
One .npz per (question, rep) trajectory · 736 files.
Schema (per file, numpy.load)
key
shape
dtype
meaning
input_ids
(L,)
int32… See the full description on the dataset page: https://huggingface.co/datasets/lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3.corral-oss-trace-logprobs
Corral – OSS-120B Trace Logprobs
Token-level log-probabilities for GPT-Oss-120B evaluation runs across all 8 Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the token-level log-probabilities recorded during the evaluation runs of GPT-Oss-120B across all 8 Corral environments.
Each configuration (config) of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-oss-trace-logprobs.openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.evolkit-logprobs-prepared-kd-temp-2_0-context-8k
