CoolFace
Datasetpublic

marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match

N8 Rejection Sampling (Soft Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match.

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes102downloads
Dataset Card

N8 Rejection Sampling (Soft Match)

Overview

This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth.

How This Dataset Was Created

For each of the 29,963 unique prompts in the source dataset (which has 8 responses per prompt), we checked whether any of the 8 Qwen3-4B final answers matched the Qwen3-32B final answer for the same prompt. If at least one match was found, we kept only the matching samples and discarded the rest. If no matches were found, we kept all 8 samples (rather than discarding them). If the Qwen3-32B answer was "N/A", we kept all 8 samples.

This is a softer filtering: verified responses are preferred, but unverifiable prompts are retained in full.

Prompts were matched between the two datasets using the ms_id column (unique per prompt). Final answers were compared via exact string match on the final_answer column.

Rejection Sampling Statistics

MetricCount
Total unique prompts29,963
Prompts with at least one Qwen3-4B answer matching Qwen3-32B17,932
Prompts with zero matches between Qwen3-4B and Qwen3-32B8,653
Prompts where Qwen3-32B answer was "N/A"3,378
Prompts where all 8 Qwen3-4B answers were "N/A"1,295
Final dataset size187,194 rows

Schema/Columns

ColumnTypeDescription
row_idint64Unique identifier
instruction_seedstringOriginal instruction/prompt
_sourcestringData source identifier
gpt41_mini_responsestringReference response
__original_row_idxint64Index from original dataset
lengthint64Response length
ms_idint64Mapping ID (unique per prompt)
generated_textstringModel-generated response with reasoning trace
final_answerstringExtracted final answer from \\boxed{}
complete_responses_countint64Number of complete responses for this prompt

Access

python
from datasets import load_dataset

dataset = load_dataset("marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match")
print(dataset["train"][0])

Use Cases

  • —Training models on mathematical reasoning with verified responses
  • —Fine-tuning for chain-of-thought reasoning
  • —Studying the effect of rejection sampling on training data quality
  • —Comparing strict vs. soft filtering strategies for synthetic data