kobzaond/RLVRAMBench
RLVRAMBench Which language-model training configurations can I use with the memory I have, and how much testing does that decision require? RLVRAMBench is a measurement dataset with open evaluation tasks for a specific language-model training system. It measures memory feasibility when response generation and reinforcement-learning updates share the same graphics processors. It provides measured outcomes, fixed prediction tasks, a budgeted decision replay, reference methods, and… See the full description on the dataset page: https://huggingface.co/datasets/kobzaond/RLVRAMBench.
0224
1{2 "host-qwen25-7b-code_heavy_tail-c2-s161": "120e29a146294722a6668f55a1e7fed2878eb62a",3 "host-qwen25-7b-code_heavy_tail-c2-s162": "120e29a146294722a6668f55a1e7fed2878eb62a",4 "host-qwen25-7b-code_heavy_tail-c2-s163": "120e29a146294722a6668f55a1e7fed2878eb62a",5 "host-qwen25-7b-gsm8k-c2-s161": "120e29a146294722a6668f55a1e7fed2878eb62a",6 "host-qwen25-7b-gsm8k-c2-s162": "120e29a146294722a6668f55a1e7fed2878eb62a",7 "host-qwen25-7b-gsm8k-c2-s163": "120e29a146294722a6668f55a1e7fed2878eb62a",8 "host-qwen25-7b-math-c2-s161": "120e29a146294722a6668f55a1e7fed2878eb62a",9 "host-qwen25-7b-math-c2-s162": "120e29a146294722a6668f55a1e7fed2878eb62a",10 "host-qwen25-7b-math-c2-s163": "120e29a146294722a6668f55a1e7fed2878eb62a"11}12 