JackHsieh/statML-arxiv-RL-2k-docs
Train-only prefix of JackHsieh/statML-arxiv-RL-4k-docs's train split, drawn from JackHsieh/statML-arxiv. Each row is one randomly sampled contiguous window of exactly 4_096 Qwen3 tokens (Qwen/Qwen3-4B-Instruct-2507) from a distinct paper. start_index is the window's offset in the source paper's token sequence; input_ids is the Qwen3 encoding of text (no special tokens added — no BOS/EOS). Same schema and recipe as JackHsieh/statML-arxiv-40M-20M. Nesting: these are the first 2_048 rows of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/statML-arxiv-RL-2k-docs.
Train-only prefix of JackHsieh/statML-arxiv-RL-4k-docs's train split, drawn from JackHsieh/statML-arxiv.
Each row is one randomly sampled contiguous window of exactly 4096 Qwen3 tokens (`Qwen/Qwen3-4B-Instruct-2507`) from a distinct paper. `startindex is the window's offset in the source paper's token sequence; input_ids is the Qwen3 encoding of text` (no special tokens added — no BOS/EOS). Same schema and recipe as JackHsieh/statML-arxiv-40M-20M.
Nesting:
- these are the first 2048 rows of [JackHsieh/statML-arxiv-RL-4k-docs](https://huggingface.co/datasets/JackHsieh/statML-arxiv-RL-4k-docs)'s `train` split **verbatim**, in its original order — identical `text`, `startindex
andinput_ids`, so the two datasets nest exactly by construction; - shares no paper with JackHsieh/statML-arxiv-40M-20M's
trainsplit (9_728 papers), which is the held-out evaluation set for these runs. - shares no paper with JackHsieh/statML-arxiv-40M-20M's
testsplit (4_864 papers), which is the held-out evaluation set for these runs. - shares no paper with
/work/nvme/bhwe/jhsieh1/runs/statml-nested/160M'strainsplit (38_912 papers), which is the held-out evaluation set for these runs.
Papers span yymm 2007 down to 0707; the composition is JackHsieh/statML-arxiv-RL-4k-docs's, inherited unchanged by taking a prefix.
