JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32
luna-reason-only.k-8.statml-arxiv-llama32 Prefix-only "thoughts" for next-token prediction on stat.ML arXiv LaTeX. Each thought is visible reasoning about the next 8 Llama-3.2 tokens after a cut, written without ever seeing that continuation. Intended to be spliced into the document before the chunk so a small model (Llama 3.2 3B base) can read the reasoning and predict the chunk. Complete: every designated chunk has a thought. split thoughts coverage documents… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.
Card: dataset complete at 100% coverage; refusals shown to be stochastic, none deterministic
Test split COMPLETE: 72,960/72,960 chunks (100%), all 4,864 documents at 15/15
Train split COMPLETE: 612,864/612,864 chunks (100%), all 9,728 documents at 63/63
Update card: 99.9% coverage both splits, all documents represented
Test split: 72,891 of 72,960 chunks (99.91%), all 4,864 documents
Train split: 612,186 of 612,864 chunks (99.89%), all 9,728 documents
Update card: both splits, per-split coverage table, gap taxonomy
Add train split: 592,971 of 612,864 chunks (96.8%), 9,455 of 9,728 documents; stride=8,k=8
Update card: 96.6% coverage, full document span, moderation-gap explanation
Test split: all 8 shards collected — 70,464 of 72,960 chunks (96.6%), all 4,864 documents; 2,496 moderation gaps pending retry
Add dataset card: provenance, schema, partial-upload status (shard 0 of 8)
Add test-split thoughts: shard 0 of 8 (9,594 of 72,960 designated chunks, 667 complete documents)
initial commit
