CoolFace
Datasetpublic

JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-echo

luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-echo Tokenized, tag-wrapped form of JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32. Each thought is wrapped as <|reserved_special_token_1|> … thought … <|reserved_special_token_2|> … last prefix token and stored both as text (thought_text) and as Llama 3.2 token ids (input_ids). The trailing token is the document token immediately before the cut (input_ids[chunk_start_index - 1]), copied from the document rather… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-echo.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes34downloads
Dataset Card

luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-echo

Tokenized, tag-wrapped form of `JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32`. Each thought is wrapped as <|reserved_special_token_1|> … thought … <|reserved_special_token_2|> … last prefix token and stored both as text (thought_text) and as Llama 3.2 token ids (input_ids).

The trailing token is the document token immediately before the cut (input_ids[chunk_start_index - 1]), copied from the document rather than re-tokenized.

Provenance

Source thoughts: JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32 — prefix-only reasoning about the next 8 Llama-3.2 tokens of stat.ML arXiv LaTeX, generated by gpt-5.6-luna (reasoning.effort="none") from the last ≤1024 tokens before each cut. Documents: JackHsieh/statML-arxiv-40M-20M-llama32. Tokenizer: meta-llama/Llama-3.2-3B, add_special_tokens=False; tag ids and the echo id are inserted explicitly, so no BOS is introduced and no token is round-tripped through text.

Schema

All source columns, plus/with:

columnnotes
input_idsLlama 3.2 ids of the wrapped thought, including tags and echo token
thought_textthe wrapped thought as text (rewritten from the source column)
gwithin-work-group index; always 0 (one thought per chunk)

Token counts (input_ids length, tags and echo included)

splitthoughtsmeanstdminmax
test72,960303.552.297485
train612,864302.452.293501
combined685,824302.552.293501

Nothing is trimmed: 44,453 thoughts exceed 384 tokens, so a run with a shorter max_thought_length must trim them (at a sentence boundary) or drop those chunks.