CoolFace
Datasetpublic

open-athena/nemotron-gym-instruction-following-calendar-qwen3.5-122b-131k-opencode-traces

Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-instruction-following-calendar-qwen3.5-122b-131k-opencode-traces.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes251downloads
6 commits on main
d7964852mo ago

Document literal-token decoding (tokenizer provenance)

penfever
0df4fc82mo ago

Add tokenizer provenance for literal token columns

penfever
2c8a1e02mo ago

Upload 8276 rows (41+ shards) in one commit

penfever
c3d89782mo ago

Remove stale text-only auto-export before literal re-export

penfever
290509e2mo ago

Upload dataset

penfever
553f3ce2mo ago

initial commit

penfever