CoolFace
Datasetpublic

marin-community/openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16

OpenThoughts-4 Code SDG: Qwen3-32B (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes447downloads
Dataset Card

OpenThoughts-4 Code SDG: Qwen3-32B (n=16, top-16 logprobs)

Synthetic generations from **Qwen/Qwen3-32B** on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis.

Generation setup

FieldValue
Generator modelQwen/Qwen3-32B (revision 9216db5781bf21249d130ec9da846c4624c16137)
Source promptsmarin-community/hero-run-4-code-sdg-prompts-python-fenced-n16 (9,168 unique prompts)
Samples per prompt (n)16
Logprobs returned (k)16 (top-k vocab logprobs per generated token)
Max generated tokens32,768
Max model length34,816
Temperature0.8
Inference enginevLLM on TPU v6e-8, tensor_parallel_size=8
Producermarin-community/marin — experiments/sdg/code/qwen3-32b/sdg_ot4_code_32768_tokens.py

Schema

The dataset is a flattened parquet table with one row per `(prompt, sample_index)` pair. With 9,168 unique prompts and n=16 samples each, the dataset contains 146,688 rows total.

Identifier columns

ColumnTypeDescription
prompt_indexint640-based index of the source prompt within the prompt set (0 … 9,167)
response_indexint640-based sample index within that prompt (0 … 15)
_unique_row_idstringStable globally-unique id for the (prompt, response) pair, copied from the source prompt set; safe key for joins and dedup
instruction_seedstringThe original natural-language coding problem from OpenThoughts-4 (before any chat templating)
generation_promptstringThe chat-templated prompt actually fed to vLLM (Qwen3 chat template applied to instruction_seed)

A given prompt is fully identified by either prompt_index or _unique_row_id; the 16 samples for a prompt share both keys and differ only in response_index.

Generation columns

ColumnTypeDescription
generated_textstringDecoded model response (everything after the chat template's assistant turn)
generated_token_idslist[int32], length TToken ids of the generated response
generated_token_logprobslist[float32], length TLog probability of each chosen token under the model
generated_top_logprob_token_idslist[int32], length T × kFlattened top-k candidate token ids at each step
generated_top_logprobslist[float32], length T × kFlattened log probabilities of those top-k candidates at each step

Where T = number of generated tokens for that row and k = 16 (the top-k logprobs setting).

Reshaping the flattened top-k arrays

vLLM was run with flat_logprobs=True for compactness, so the per-step top-k arrays are stored as 1-D lists. To recover the standard (T, k) shape:

python
import numpy as np
import pyarrow.parquet as pq

df = pq.read_table("part-000000.parquet").to_pandas()
row = df.iloc[0]
T = len(row["generated_token_ids"])
k = 16
top_ids   = np.asarray(row["generated_top_logprob_token_ids"]).reshape(T, k)
top_logps = np.asarray(row["generated_top_logprobs"]).reshape(T, k)
# top_ids[t, :]   = top-16 candidate token ids at step t
# top_logps[t, :] = log p of those candidates at step t (sorted high → low by vLLM)
# the chosen token may or may not appear in the top-k; use generated_token_logprobs
# for the chosen-token logprob.

File layout

The dataset is sharded into ~2,700 parquet files named part-NNNNNN.parquet, each with up to 128 rows. Use the default config to load the full train split.

Companion datasets

SliceGenerator
codeQwen3-30B-A3B-Thinking-2507
codeQwen3-32B (this dataset)
codeQwen3-4B
codeGemma-4-31B-IT (forthcoming)
scienceQwen3 family + Gemma-4-31B-IT (forthcoming)

License

Released under Apache 2.0. The underlying generator model (Qwen/Qwen3-32B) is governed by its own license; consult the model card before redistribution.

Citation

If you use this dataset, please cite the Marin project and the Qwen3 technical report.