vedant33/supersede-rl-episodes
Supersede RL Episodes Procedurally-generated supersession episodes for reinforcement learning — the training data behind vedant33/supersede-qwen2.5-3b-grpo-lora and the Supersede environment. Paper Code Environment Model arXiv · DOI GitHub Prime Intellect Hub vedant33/supersede-qwen2.5-3b-grpo-lora What this is Each episode is a short multi-session conversation in which a fact about the user changes one or more times (they move city, switch jobs… See the full description on the dataset page: https://huggingface.co/datasets/vedant33/supersede-rl-episodes.
062
