datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
evolkit-logprobs-pipeline-75k-v2-samplecosmopedia-logprobsacm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3
lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3
Teacher (Qwen3.5-397B-A17B) top-20 forward-KL log-prob annotations for offline on-policy
distillation (OPD) of Qwen3.5-9B on BrowseComp-Plus train680 (MemTool regime).
Trains: OPD iter-3
Annotates the rollouts of: iter-2 rollouts (…-train-rollouts-…-epoch2)
One .npz per (question, rep) trajectory · 736 files.
Schema (per file, numpy.load)
key
shape
dtype
meaning
input_ids
(L,)
int32… See the full description on the dataset page: https://huggingface.co/datasets/lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3.corral-oss-trace-logprobs
Corral – OSS-120B Trace Logprobs
Token-level log-probabilities for GPT-Oss-120B evaluation runs across all 8 Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the token-level log-probabilities recorded during the evaluation runs of GPT-Oss-120B across all 8 Corral environments.
Each configuration (config) of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-oss-trace-logprobs.openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.evolkit-logprobs-prepared-kd-temp-2_0-context-8kc4-token-logprobsopenthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-32B (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-32B
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16.qwen35-4b-drpo-vs0f49th-trainer-logprobs
Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th
This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587).
Contents
Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th
Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/
Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl
Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.numina-cot-logprobs-859k-8b-sft
Dataset Card for numina-cot-logprobs-859k-8b-sft
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/axolotl-ai-co/numina-cot-logprobs-859k-8b-sft/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/numina-cot-logprobs-859k-8b-sft.openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch2
lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch2
Teacher (Qwen3.5-397B-A17B) top-20 forward-KL log-prob annotations for offline on-policy
distillation (OPD) of Qwen3.5-9B on BrowseComp-Plus train680 (MemTool regime).
Trains: OPD iter-2
Annotates the rollouts of: iter-1 rollouts (…-train-rollouts-…-epoch1)
One .npz per (question, rep) trajectory · 849 files.
Schema (per file, numpy.load)
key
shape
dtype
meaning
input_ids
(L,)
int32… See the full description on the dataset page: https://huggingface.co/datasets/lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch2.qwen37-pi-qwen36-27b-topk40-logprobs
Qwen3.7 Pi Trace Top-40 Teacher Logprobs
Offline top-40 teacher logprobs for cumulative assistant-turn rows from
armand0e/qwen3.7-max-split-formatted.
These files are intended to be loaded with snapshot_download, not
datasets.load_dataset.
Contents
manifest.json: shard metadata and filtering counts
shard-*.pt: tokenized examples with labels, target positions, top-k token ids,
and top-k teacher logprobs
chat_template.jinja: the exact chat template used for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen37-pi-qwen36-27b-topk40-logprobs.evolkit-logprobs-pipeline-75k-v2
Dataset Card for evolkit-logprobs-pipeline-75k-v2
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/winglian/evolkit-logprobs-pipeline-75k-v2/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/evolkit-logprobs-pipeline-75k-v2.openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-32B
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch1
lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch1
Teacher (Qwen3.5-397B-A17B) top-20 forward-KL log-prob annotations for offline on-policy
distillation (OPD) of Qwen3.5-9B on BrowseComp-Plus train680 (MemTool regime).
Trains: OPD iter-1
Annotates the rollouts of: base model rollouts (react+memtool ×4)
One .npz per (question, rep) trajectory · 1145 files.
Schema (per file, numpy.load)
key
shape
dtype
meaning
input_ids
(L,)
int32… See the full description on the dataset page: https://huggingface.co/datasets/lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch1.cosmopedia_web_textbooks_logprobstinyevals-logprobs-llama2-allsizesopenthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-4B (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-4B
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model
Qwen/Qwen3-4B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16.iclr2026-lm-logprobs
LM Log-Probabilities for Value Bias Analysis
Next-token log-probability distributions from 12 language models across 54 prompts, used in the paper:
Reward Models Inherit Value Biases from Pretraining
Brian Christian, Jessica A.F. Thompson, Elle, Vincent Adam, Hannah Rose Kirk, Christopher Summerfield, Tsvetomira Dumbalska (ICLR 2026)
Part of the Oxford-HIPlab collection for this paper.
Dataset description
Each CSV contains the full next-token log-probability… See the full description on the dataset page: https://huggingface.co/datasets/Oxford-HIPlab/iclr2026-lm-logprobs.sffop_1706381144_410msft_relabel_pythia6.9b_logprobs_prefix_chosenopenthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-4B (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-4B
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16.omnidocbench-qwen-ocr-logprobs
OmniDocBench Qwen OCR Log-Probabilities
This dataset provides token-level and bounding-box-level OCR log-probabilities produced by
running Qwen3.5-122B-A10B (via vLLM) on the original page scans of the
OmniDocBench benchmark.
It is a reference-free auxiliary signal — no ground-truth text is used.
Dataset Structure
ocr_logprobs/
<page_id>/
ocr_logprobs.json # full per-token logprobs + top-5 alternatives
ocr_html.html # raw HTML output from the OCR model… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-qwen-ocr-logprobs.prompt-logprobs
Dataset Card for prompt-logprobs
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/gabrielmbmb/prompt-logprobs/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/gabrielmbmb/prompt-logprobs.summarize_from_feedback_oai_preprocessing_1706381144_410msft_relabel_pythia6.9b_logprobsflexolmo-math-logprobsTop 128 logprobs from dolmino-mix-1124's math data.
Model used to generate was allenai/Flex-math-2x7B-1T.
ds = load_dataset(
"allenai/dolmino-mix-1124",
data_files=[
"data/math/gsm8k/**/*.jsonl",
"data/math/metamath-owmfilter/**/*.jsonl",
"data/math/tulu_math/**/*.jsonl",
],
split="train",
)
dataset_info:
features:
name: input_ids
list: int32
name: topk_indices
list:
list: int32
name: topk_logprobs
list:
list: float16
splits: train_0… See the full description on the dataset page: https://huggingface.co/datasets/hbfreed/flexolmo-math-logprobs.math500-olmo-3-7b-instruct-temp0.9-samples99-logprobs
OLMo-3-7B-Instruct self-consistency generations with logprobs on MATH500
This dataset contains 99 self-consistency generations per question for the
MATH500 benchmark, produced with allenai/OLMo-3-7B-Instruct at temperature
0.9, together with token-level log probabilities for each completion.
The file is intended for post-hoc analysis, self-consistency curves, adaptive
stopping, and related aggregation methods.
Source
Base benchmark: HuggingFaceH4/MATH-500
Model:… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/math500-olmo-3-7b-instruct-temp0.9-samples99-logprobs.v0-next-logprobs-llama2-200ksffop_1706381144_410msft_relabel_pythia6.9b_logprobs_cond3emojieallprefixresults_from_improved_logprobs_llama70B
