JackHsieh/qwen3-distill-3rkrz2vo.k-8.statml-arxiv-qwen3.qwen3-ids.tags
qwen3-distill-3rkrz2vo.k-8.statml-arxiv-qwen3.qwen3-ids.tags Tokenized, tag-wrapped form of JackHsieh/qwen3-4B-instruct-luna-distill-step2362.reason-only.k-8.statml-arxiv-qwen3. Each thought is wrapped as <|note|> … thought … <|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base token ids (input_ids). Longest thought: 770 tokens — a training run's max_thought_length must be at least this. Delimiter ids: <|note|> = 151669, <|/note|> = 151670. A run must… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/qwen3-distill-3rkrz2vo.k-8.statml-arxiv-qwen3.qwen3-ids.tags.
qwen3-distill-3rkrz2vo.k-8.statml-arxiv-qwen3.qwen3-ids.tags
Tokenized, tag-wrapped form of `JackHsieh/qwen3-4B-instruct-luna-distill-step2362.reason-only.k-8.statml-arxiv-qwen3`. Each thought is wrapped as <|note|> … thought … <|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base token ids (input_ids).
Longest thought: 770 tokens — a training run's max_thought_length must be at least this.
Delimiter ids: <|note|> = 151669, <|/note|> = 151670. A run must declare these under trainee.special_tokens IN THIS ORDER, or the ids will not match.
Provenance
Source thoughts: JackHsieh/qwen3-4B-instruct-luna-distill-step2362.reason-only.k-8.statml-arxiv-qwen3. Documents: JackHsieh/statML-arxiv-40M-20M. Tokenizer: Qwen/Qwen3-4B-Base, add_special_tokens=False; tag ids and any document ids are inserted explicitly, so no BOS is introduced and no token is round-tripped through text.
Schema
All source columns, plus/with:
Token counts (input_ids length, tags included)
Nothing is trimmed: 27,930 thoughts exceed 384 tokens, so a run with a shorter max_thought_length must trim them (at a sentence boundary) or drop those chunks.
