thinking-tokens
math_modelsafety_modelgeneral_knowledge_modelmultilingual_modelgroup_modelQwen_Qwen3-4B-Thinking-2507_nvfp4-ts_qwen3-random-tokens_2048_8_1024_256_lr0.03Qwen_Qwen3-4B-Thinking-2507_int3-g128_qwen3-random-tokens_2048_8_1024_256_lr0.03Qwen_Qwen3-4B-Thinking-2507_int3-g16-fp8_qwen3-random-tokens_2048_8_1024_256_lr0.03
open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8
Open Thoughts 4 - Code (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8)
This dataset contains code reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507.
Overview
Source: marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated (prompts only)
Model: Qwen/Qwen3-30B-A3B-Thinking-2507
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The code problem prompt
_source
Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8
Open Thoughts 4 - Math (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8)
This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507.
Overview
Source: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated (base prompts)
Model: Qwen/Qwen3-30B-A3B-Thinking-2507
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The math problem prompt
_source
Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted
Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-30B-A3B-Thinking-2507-Annotated-32768-Tokens-N8-Reformatted
Overview
This dataset is a reformatted version of marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8. The original dataset contained 29,963 samples, each with 8 responses generated by the same model with different random seeds (stored in generated_text, generated_text2, ..., generated_text8 columns). This reformatted… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.tokenskip-math-qwen3-4b-thinking-n8
TokenSkip-MATH CoTs — Qwen3-4B-Thinking (N=8) + LLMLingua-2 compressions
Intermediate artifacts for reproducing TokenSkip on MATH with Qwen/Qwen3-4B-Thinking-2507. These are the outputs of steps 1 (collect) and 2 (compress) of the pipeline in LE-WH/LatentReasoning — the long-running steps. Uploading them lets you skip straight to SFT data preparation (step 3) on a fresh machine.
Provenance. Collected on 4× GPUs with vLLM, N=8 samples per question at T=0.7, max generation 4,096… See the full description on the dataset page: https://huggingface.co/datasets/leapeto/tokenskip-math-qwen3-4b-thinking-n8.open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted
Dataset Card for Open-Thoughts-4-30K-Code-Qwen3-30B-A3B-Thinking-2507-Annotated-32768-Tokens-N8-Reformatted
Overview
This dataset is a reformatted version of marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8. The original dataset contained 29,963 samples, each with 8 responses generated by the same model with different random seeds (stored in generated_text, generated_text2, ..., generated_text8 columns). This reformatted… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.open-thoughts-4-8530-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted
open-thoughts-4-8530-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens
This dataset contains 68240 rows (8530 math prompts x 8 responses each) generated by
Qwen3-30B-A3B-Thinking-2507 with a max token length of 32768.
Derived from the first 68240 rows of
marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.
Columns
Column
Description
row_id
Original row identifier
instruction_seed
The math prompt… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-8530-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.
