JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids
luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by gpt-5.6-luna. Each thought is visible reasoning about the next 8 Qwen3 tokens after a cut, written without ever seeing that continuation. Intended to be spliced into the document before the chunk so a small model (Qwen3-4B-Base) can read the reasoning and predict the chunk. Thoughts are strings, not token ids;… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.
luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of `JackHsieh/statML-arxiv-40M-20M`, generated by gpt-5.6-luna. Each thought is visible reasoning about the next 8 Qwen3 tokens after a cut, written without ever seeing that continuation. Intended to be spliced into the document before the chunk so a small model (Qwen3-4B-Base) can read the reasoning and predict the chunk. Thoughts are strings, not token ids; tokenized tag-wrapped variants are separate datasets (.tags, .kv-tags, .kv-tags-explained suffixes).
Qwen3 counterpart of `luna-reason-only.k-8.statml-arxiv-llama32`. The two share document uuids and chunk-index grids but not cut positions: 4096 Qwen3 tokens span slightly different text than 4096 Llama-3.2 tokens, so chunk k falls at a different character offset in each. Thoughts are not interchangeable between them.
Note that input_tokens, cached_tokens, output_tokens columns measure the OpenAI tokens billed for this request, not Qwen3 tokens; kept for cost auditing |
Generation
Thought lengths in Qwen3 tokens
Generation metrics
Moderation refusals
~4% of requests were refused as invalid_prompt on entirely benign LaTeX — FDR tables, accuracy tables, convergence bounds. These are stochastic, not content-determined: identical retries recovered them at high rates, and all 622 residual chunks succeeded within 4 identical synchronous attempts (85% on the first). No prompt in this corpus is deterministically refused. Retries re-send the prompt unchanged; the prefix end, which defines the prediction task, is never moved.
Known characteristics
- ~70% of thoughts open with the words "The cut is" — a stable habit of the generator that no prompt variant dislodged. Anything distilled from these will inherit it.
- Early chunks see a shorter prefix:
chunk_index=8has only 64 Qwen3 tokens of context, so those thoughts are competent but under-informed relative to full-window ones.
Coverage details
Documents are 4096 Qwen3 tokens = 512 chunks at chunk size k=8; the first chunk of a document is never thoughtful.
Coverage is complete: every designated chunk has a thought.
