marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.
OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs)
Synthetic generations from **Qwen/Qwen3-32B** on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis.
Generation setup
Schema
The dataset is a flattened parquet table with one row per `(prompt, sample_index)` pair. With 26,041 unique prompts and n=8 samples each, the dataset contains 208,328 rows total.
Identifier columns
A given prompt is fully identified by either prompt_index or _unique_row_id; the 8 samples for a prompt share both keys and differ only in response_index.
Generation columns
Where T = number of generated tokens for that row and k = 16.
Reshaping the flattened top-k arrays
vLLM was run with flat_logprobs=True for compactness; the per-step top-k arrays are stored as 1-D lists. To recover the standard (T, k) shape:
import numpy as np
import pyarrow.parquet as pq
df = pq.read_table("part-000000.parquet").to_pandas()
row = df.iloc[0]
T = len(row["generated_token_ids"])
k = 16
top_ids = np.asarray(row["generated_top_logprob_token_ids"]).reshape(T, k)
top_logps = np.asarray(row["generated_top_logprobs"]).reshape(T, k)
# top_ids[t, :] = top-16 candidate token ids at step t
# top_logps[t, :] = log p of those candidates at step t (sorted high → low by vLLM)
# the chosen token may or may not appear in the top-k; use generated_token_logprobs
# for the chosen-token logprob.File layout
The dataset is sharded into ~2,250 parquet files named part-NNNNNN.parquet. Use the default config to load the full train split.
Companion datasets
License
Released under Apache 2.0. The underlying generator model (Qwen/Qwen3-32B) is governed by its own license; consult the model card before redistribution.
Citation
If you use this dataset, please cite the Marin project and the Qwen3 technical report.
