datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omega-mm-test34B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the
cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the
exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the
cut — and answer with a few dense, declarative sentences (about two to four) inside a
<thought>…</thought> block, focused on the exact state at the cut and what the local
grammar… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.
The source thoughts are two parts: a <think> block, then a few dense sentences that the
generator wrapped in a literal <thought>…</thought> block (plain text, not special tokens).
Only the part after </think> becomes the VALUE, with every <thought> / </thought> tag
removed… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.
The source thoughts are two parts: a <think> block, then a single paragraph of dense
reasoning about the immediate continuation. Only the part after </think> — the paragraph —
becomes the VALUE. The reasoning inside the think block is dropped.
The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.pretrain_test34B-ranked-v7.rule-stride-train4-test32.k-8.L-4096.statml-arxivtest3
📊 Test3
Ce dataset contient des données d'API au format JSONL et Parquet, avec des variantes de questions et leurs réponses associées.
📂 Structure des Fichiers
data.parquet : Dataset principal au format Parquet
dataset_base.jsonl : Données de base
dataset_dpo.jsonl et dataset_dpo.parquet : Format DPO avec paires chosen/rejected
dataset_avec_variantes.jsonl : Données avec variantes de questions
dataset_final.parquet : Dataset final au format Parquet… See the full description on the dataset page: https://huggingface.co/datasets/yann23/test3.luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids
luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
gpt-5.6-luna. Each thought is visible reasoning about the next 8 Qwen3 tokens after a cut,
written without ever seeing that continuation. Intended to be spliced into the document
before the chunk so a small model (Qwen3-4B-Base) can read the reasoning and predict the
chunk. Thoughts are strings, not token ids;… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tsa-throughput-streaming-test3rf_newconcat_test3test3browsecomp-plus-scout-runs-test300-qwen-sft-gpt-scout-unfiltered-v14B-ranked-v7.rule-stride-train4-test32.k-8.L-4096.statml-arxiv.qwen3-ids.kv-tags-explainedbrowsecomp-plus-sel-tools-test300-random-seed1-v1test3_suno_resrbrowsecomp-plus-sel-tools-test300-gemini-3p1-pro-v14B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids.
The source thoughts are two parts: a <think> block, then a ranked list of candidate
continuations. Only the part after </think> — the list — becomes the VALUE. The reasoning
inside the think block is dropped.
The VALUE is capped at 512 tokens.
4.38% of rows have no usable list: the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids. Each thought_text is wrapped as
<|note|>
This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next.
KEY: {last 8 prefix tokens}
VALUE: {thought_text}
<|/note|>
and stored both as text… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.rf_withprompt_test3_1long_cl-test34B-reason-only.stride-train8off4-test32.k-8.L-512.statml-arxiv
4B-reason-only.stride-train8off4-test32.k-8.L-512.statml-arxiv
Reasoning about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, generated by
stock Qwen/Qwen3-4B-Instruct-2507 -- no fine-tuning, prompted with the reason-only-nothink.jinja reasoning template.
Documents: JackHsieh/statML-arxiv-40M-20M, 4096 Qwen3 tokens each.
Why this checkpoint: Stock Qwen3-4B-Instruct-2507 prompted with reason-only-nothink.jinja, a byte-identical copy of the reason-only.jinja template used for… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-reason-only.stride-train8off4-test32.k-8.L-512.statml-arxiv.luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tags
luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tags
A pre-tokenized, tag-wrapped variant of JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids. Each thought_text is wrapped as
<|note|>{thought_text}<|/note|>
and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base token ids (input_ids).
Tag ids and the spliced key are inserted as ids, never re-tokenized.
Delimiter ids: <|note|> = 151669, <|/note|> = 151670. These are the first… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tags.4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids
4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task is not "write a thought". Instead, it is predict the next k=8 tokens
after the cut, and report the prediction as a ranked list of 3–6 candidate continuations,
each one being the tail copied verbatim plus the predicted continuation, written as
N. «tail + continuation».… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids.luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.kv-tags
luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.kv-tags
A pre-tokenized, tag-wrapped variant of JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids. Each thought_text is wrapped as
<|note|>
KEY: {last 8 prefix tokens}
VALUE: {thought_text}
<|/note|>
and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base token ids (input_ids).
Tag ids and the spliced key are inserted as ids, never re-tokenized.
Delimiter ids: <|note|> = 151669… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.kv-tags.browsecomp-plus-sel-tools-test300-random-seed6-v1browsecomp-plus-scout-runs-test300-qwen-sft-random-v120NG_train10.8k_test3.6K_valid3.6k
Dataset Card for "20NG_train10.8k_test3.6K_valid3.6k"
More Information needed
test3
Dataset Card for "test3"
More Information needed
32B-reason-only.stride-train8off4-test32.k-8.L-512.statml-arxiv.qwen3-ids.kv-tags-explained
32B-reason-only.stride-train8off4-test32.k-8.L-512.statml-arxiv.qwen3-ids.kv-tags-explained
Tokenized, tag-wrapped form of JackHsieh/32B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv.
Each thought is wrapped as
<|note|>
This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next.
KEY: <last 8 prefix tokens>
VALUE: <thought>
<|/note|>
and stored both as text (thought_text) and as… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-train8off4-test32.k-8.L-512.statml-arxiv.qwen3-ids.kv-tags-explained.
