JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-32B with thinking mode off. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. The 4B parity counterpart is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.
32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of `JackHsieh/statML-arxiv-40M-20M`, generated by Qwen3-32B with thinking mode off. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper.
The 4B parity counterpart is `JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids`: same chunk grid, same prompt, same sampling, smaller generator.
Generation
Thought lengths in Qwen3 tokens
Generation metrics
Coverage details
Documents are 4096 Qwen3 tokens = 512 chunks at chunk size k=8; the first chunk of a document is never thoughtful.
Schema
One row per (chunk, g): uuid, chunk_index, chunk_start_index, chunk_end_index, g, input_ids (generator-tokenizer ids of the wrapped thought), thought_text, truncated, finish_reason, generated_at, generation_seconds.
