datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8
Open Thoughts 4 - Code (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8)
This dataset contains code reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507.
Overview
Source: marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated (prompts only)
Model: Qwen/Qwen3-30B-A3B-Thinking-2507
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The code problem prompt
_source
Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8
Open Thoughts 4 - Math (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8)
This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507.
Overview
Source: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated (base prompts)
Model: Qwen/Qwen3-30B-A3B-Thinking-2507
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The math problem prompt
_source
Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted
Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-30B-A3B-Thinking-2507-Annotated-32768-Tokens-N8-Reformatted
Overview
This dataset is a reformatted version of marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8. The original dataset contained 29,963 samples, each with 8 responses generated by the same model with different random seeds (stored in generated_text, generated_text2, ..., generated_text8 columns). This reformatted… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted
Dataset Card for Open-Thoughts-4-30K-Code-Qwen3-30B-A3B-Thinking-2507-Annotated-32768-Tokens-N8-Reformatted
Overview
This dataset is a reformatted version of marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8. The original dataset contained 29,963 samples, each with 8 responses generated by the same model with different random seeds (stored in generated_text, generated_text2, ..., generated_text8 columns). This reformatted… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.open-thoughts-4-8530-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted
open-thoughts-4-8530-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens
This dataset contains 68240 rows (8530 math prompts x 8 responses each) generated by
Qwen3-30B-A3B-Thinking-2507 with a max token length of 32768.
Derived from the first 68240 rows of
marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.
Columns
Column
Description
row_id
Original row identifier
instruction_seed
The math prompt… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-8530-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.open-thoughts-4-128-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted
open-thoughts-4-128-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens
Math reasoning responses generated by Qwen3-30B-A3B-Thinking-2507 (Qwen/Qwen3-30B-A3B-Thinking-2507).
Overview
Total rows: 1,024
Unique prompts: 128 (each with 8 response annotations)
Source prompts: marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted
Prompt alignment: Exact instruction_seed match to… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-128-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.open-thoughts-4-128-math-qwen3-235b-thinking-2507-annotated-32768-tokens
open-thoughts-4-128-math-qwen3-235b-thinking-2507-annotated-32768-tokens
Math reasoning responses generated by Qwen3-235B-A22B-Thinking-2507 (Qwen/Qwen3-235B-A22B-Thinking-2507) via the Together AI serverless API.
Overview
Total rows: 1,024
Unique prompts: 128 (each with 8 response annotations)
Source prompts: marin-community/open-thoughts-4-128-math-qwen3-32b-annotated-32768-tokens-n8-reformatted
Generation model: Qwen/Qwen3-235B-A22B-Thinking-2507
Max tokens: 32,768… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-128-math-qwen3-235b-thinking-2507-annotated-32768-tokens.qwen3-4b-thinking-2507-random-tokens-16x1024-len32768
Random Token Dataset for Qwen/Qwen3-4b-thinking-2507
This dataset contains fully random tokenizer IDs sampled uniformly from [0, 151643).
model: Qwen/Qwen3-4b-thinking-2507
vocab_size: 151643
num_samples: 16384
seq_len: 32768
total_tokens: 536870912
seed: 0
samples_per_shard: 128
num_shards: 128
generated_at_utc: 2026-04-30T22:43:40.261456+00:00
Columns
sample_id: global row index.
input_ids: list of random token IDs, length 32768.
length: token count for the row.… See the full description on the dataset page: https://huggingface.co/datasets/TeenSpirit/qwen3-4b-thinking-2507-random-tokens-16x1024-len32768.
