datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-32B
with thinking mode off. Each thought is a few dense sentences of reasoning about the next
8 tokens after a cut, written from the document prefix alone — the generator never sees the
continuation. Stored thought_text includes the <thought>/</thought> wrapper.
The 4B parity counterpart is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next
8 tokens after a cut, written from the document prefix alone — the generator never sees the
continuation. Stored thought_text includes the <thought>/</thought> wrapper.
This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.statML-arxiv-40M-20MSubset of JackHsieh/statML-arxiv. Each document is exactly 4096 tokens.
The train split has exactly twice the number of documents as the test split.
Split
Documents
Tokens
train
9728
39_845_888
test
4864
19_922_944
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the
cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the
exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the
cut — and answer with a few dense, declarative sentences (about two to four) inside a
<thought>…</thought> block, focused on the exact state at the cut and what the local
grammar… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.Qwen3-4B-Instruct-2507.rule-thoughtful-except-first.k-64.L-1024.statml-arxiv4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.
The source thoughts are two parts: a <think> block, then a few dense sentences that the
generator wrapped in a literal <thought>…</thought> block (plain text, not special tokens).
Only the part after </think> becomes the VALUE, with every <thought> / </thought> tag
removed… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.
The source thoughts are two parts: a <think> block, then a single paragraph of dense
reasoning about the immediate continuation. Only the part after </think> — the paragraph —
becomes the VALUE. The reasoning inside the think block is dropped.
The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.4B-predict.rule-r-1.0-k-256.L-1024.statml-arxivQwen3-4B-Instruct-2507.rule-thoughtful-except-first.k-128.L-512.statml-arxivQwen3-4B-Instruct-2507.rule-thoughtful-except-first.k-64.L-256.statml-arxivQwen3-4B-Instruct-2507.rule-thoughtful-except-first.k-128.L-1024.statml-arxivluna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled
luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled
Every non-first chunk of every document carries a thought: the gpt-5.6-luna reasoning thought
where one was generated, and a content-free pause thought everywhere else.
luna chunk <|reserved_special_token_1|> luna reasoning <|reserved_special_token_2|>
filler chunk <|reserved_special_token_1|> 256x <|reserved_special_token_0|> <|reserved_special_token_2|>
The filler is 258 tokens. Chunk 0 is excluded… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled.4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
Tokenized, tag-wrapped form of JackHsieh/4B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv.
Each thought is wrapped as
<|note|>
This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next.
KEY: <last 8 prefix tokens>
VALUE: <thought>
<|/note|>
and stored both as text (thought_text) and as
Qwen/Qwen3-4B-Base… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.pause.rule-r-1.0-k-8.L-128.statml-arxiv4B-ranked-v7.rule-stride-train4-test32.k-8.L-4096.statml-arxivstatML-arxiv-80MTrain-only prefix of JackHsieh/statML-arxiv-160M's train split, drawn from JackHsieh/statML-arxiv.
Each row is one randomly sampled contiguous window of exactly 4_096 Qwen3 tokens
(Qwen/Qwen3-4B-Instruct-2507) from a distinct paper. start_index is the window's offset in the source paper's
token sequence; input_ids is the Qwen3 encoding of text (no special tokens added — no
BOS/EOS). Same schema and recipe as
JackHsieh/statML-arxiv-40M-20M.
Nesting:
these are the first 19_456 rows of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/statML-arxiv-80M.4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176 — Qwen3-4B-Instruct-2507 after RL against a frozen suffix conditional. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.Qwen3-4B-Instruct-2507.rule-thoughtful-except-first.k-256.L-1024.statml-arxivluna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids
luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
gpt-5.6-luna. Each thought is visible reasoning about the next 8 Qwen3 tokens after a cut,
written without ever seeing that continuation. Intended to be spliced into the document
before the chunk so a small model (Qwen3-4B-Base) can read the reasoning and predict the
chunk. Thoughts are strings, not token ids;… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.statML-arxiv-40M-20M-llama32-pause7-injectedA pause-injected variant of JackHsieh/statML-arxiv-40M-20M-llama32:
every document token is preceded by a sentinel-wrapped stretch of pause tokens, so each row is exactly
8x longer (4_096 -> 32_768 tokens). Nothing else changes — same papers, same
windows, same train/test split, same schema and column order as the parent.
How it was derived
For each row of the parent, the token sequence [t0, t1, ...] becomes:
<|reserved_special_token_1|> <|reserved_special_token_0|> x5… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/statML-arxiv-40M-20M-llama32-pause7-injected.statML-arxiv-42M-11M-L-1024-llama32-pause7-injectedA pause-injected variant of JackHsieh/statML-arxiv-42M-11M-L-1024-llama32:
every document token is preceded by a sentinel-wrapped stretch of pause tokens, so each row is exactly
8x longer (1_024 -> 8_192 tokens). Nothing else changes — same papers, same
windows, same train/test split, same schema and column order as the parent.
How it was derived
For each row of the parent, the token sequence [t0, t1, ...] becomes:
<|reserved_special_token_1|> <|reserved_special_token_0|> x5… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/statML-arxiv-42M-11M-L-1024-llama32-pause7-injected.4B-Instruct-distill-jx739p0v.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-Instruct-distill-jx739p0v.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-Instruct-distill-jx739p0v.stride-1.k-8.statml-arxiv.qwen3-ids. Each thought_text is wrapped as
<|note|>
This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next.
KEY: {last 8 prefix tokens}
VALUE: {thought_text}
<|/note|>
and stored both as text… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-distill-jx739p0v.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.statML-arxiv-160MTrain-only superset of JackHsieh/statML-arxiv-40M-20M's train split, drawn from JackHsieh/statML-arxiv.
Each row is one randomly sampled contiguous window of exactly 4_096 Qwen3 tokens
(Qwen/Qwen3-4B-Instruct-2507) from a distinct paper. start_index is the window's offset in the source paper's
token sequence; input_ids is the Qwen3 encoding of text (no special tokens added — no
BOS/EOS). Same schema and recipe as
JackHsieh/statML-arxiv-40M-20M.
Nesting:
rows [0:9_728] are… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/statML-arxiv-160M.pause63-echo.notag.r-1.0-k-8.statml-arxiv4B-ranked-v7.rule-stride-train4-test32.k-8.L-4096.statml-arxiv.qwen3-ids.kv-tags-explained32B-propose-choose.rule-r-1.0-k-8.L-512.statml-arxivspam.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
spam.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A length-matched null control for JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids, stored directly in the
kv-tags-explained training format (there is no separate base repo). Every thought is this
one sentence, repeated:
We are thinking hard about what comes in the next 8 tokens by reasoning correctly and carefully about what comes before it in the document.
It is fluent, on-topic and… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/spam.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.statML-arxivContent of stats.ML articles in the arxiv split of together-computer/RedPajama-Data-1T. Filtered using the metadata provided by
librarian-bots/arxiv-metadata-snapshot.
32B-predict.rule-r-1.0-k-8.L-16.statml-arxiv
