kobzaond/RLVRAMBench
RLVRAMBench Which language-model training configurations can I use with the memory I have, and how much testing does that decision require? RLVRAMBench is a measurement dataset with open evaluation tasks for a specific language-model training system. It measures memory feasibility when response generation and reinforcement-learning updates share the same graphics processors. It provides measured outcomes, fixed prediction tasks, a budgeted decision replay, reference methods, and… See the full description on the dataset page: https://huggingface.co/datasets/kobzaond/RLVRAMBench.
0223
1configuration_id,case_id,study,model_family,model_name,model_id,workload_id,workload_name,legacy_dataset_id,algorithm,adaptation,configuration_level,condition,gpu_count,rollout_tp_size,actor_micro_batch,rollout_logprob_micro_batch,ref_logprob_micro_batch,vllm_gpu_memory_utilization,parameter_offload,optimizer_offload,free_cache_engine,max_prompt_length,max_response_length,max_model_len,max_num_seqs,rollout_n,train_batch_size,train_max_samples,val_max_samples,total_training_steps,gpu_monitor_interval_ms,allocator_trace_enabled,allocator_trace_sync,save_freq,test_freq,val_before_train,resume_mode,max_actor_ckpt_to_keep,accelerator,device_capacity_mib,margin_limit_mib,peak_measurement2est-qwen25-7b-gsm8k-c2-2gpu,est-qwen25-7b-gsm8k-2gpu,prospective_estimation,qwen25_7b,Qwen2.5-7B-Instruct,Qwen/Qwen2.5-7B-Instruct,gsm8k,GSM8K mathematics,gsm8k,GRPO,LoRA,c2,,2,1,8,4,4,0.6,False,False,True,1024,2048,3072,64,2,32,32,32,1,100,True,1,-1,-1,False,disable,1,NVIDIA A100-SXM4-40GB,40960,38912,maximum externally sampled per-device memory3est-qwen25-7b-gsm8k-c2-4gpu,est-qwen25-7b-gsm8k-4gpu,prospective_estimation,qwen25_7b,Qwen2.5-7B-Instruct,Qwen/Qwen2.5-7B-Instruct,gsm8k,GSM8K mathematics,gsm8k,GRPO,LoRA,c2,,4,1,8,4,4,0.6,False,False,True,1024,2048,3072,64,2,32,32,32,1,100,True,1,-1,-1,False,disable,1,NVIDIA A100-SXM4-40GB,40960,38912,maximum externally sampled per-device memory4est-qwen25-7b-gsm8k-c3-2gpu,est-qwen25-7b-gsm8k-2gpu,prospective_estimation,qwen25_7b,Qwen2.5-7B-Instruct,Qwen/Qwen2.5-7B-Instruct,gsm8k,GSM8K mathematics,gsm8k,GRPO,LoRA,c3,,2,1,16,4,4,0.7,False,False,True,1024,2048,3072,64,2,32,32,32,1,100,True,1,-1,-1,False,disable,1,NVIDIA A100-SXM4-40GB,40960,38912,maximum externally sampled per-device memory5est-qwen25-7b-gsm8k-c3-4gpu,est-qwen25-7b-gsm8k-4gpu,prospective_estimation,qwen25_7b,Qwen2.5-7B-Instruct,Qwen/Qwen2.5-7B-Instruct,gsm8k,GSM8K mathematics,gsm8k,GRPO,LoRA,c3,,4,1,16,4,4,0.7,False,False,True,1024,2048,3072,64,2,32,32,32,1,100,True,1,-1,-1,False,disable,1,NVIDIA A100-SXM4-40GB,40960,38912,maximum externally sampled per-device memory6est-qwen25-7b-math-c2-2gpu,est-qwen25-7b-math-2gpu,prospective_estimation,qwen25_7b,Qwen2.5-7B-Instruct,Qwen/Qwen2.5-7B-Instruct,math,MATH mathematics,math,GRPO,LoRA,c2,,2,1,8,8,8,0.6,False,False,True,512,512,1024,128,2,64,64,64,1,100,True,1,-1,-1,False,disable,1,NVIDIA A100-SXM4-40GB,40960,38912,maximum externally sampled per-device memory7est-qwen25-7b-math-c2-4gpu,est-qwen25-7b-math-4gpu,prospective_estimation,qwen25_7b,Qwen2.5-7B-Instruct,Qwen/Qwen2.5-7B-Instruct,math,MATH mathematics,math,GRPO,LoRA,c2,,4,1,8,8,8,0.6,False,False,True,512,512,1024,128,2,64,64,64,1,100,True,1,-1,-1,False,disable,1,NVIDIA A100-SXM4-40GB,40960,38912,maximum externally sampled per-device memory8est-qwen25-7b-math-c3-2gpu,est-qwen25-7b-math-2gpu,prospective_estimation,qwen25_7b,Qwen2.5-7B-Instruct,Qwen/Qwen2.5-7B-Instruct,math,MATH mathematics,math,GRPO,LoRA,c3,,2,1,16,8,8,0.7,False,False,True,512,512,1024,128,2,64,64,64,1,100,True,1,-1,-1,False,disable,1,NVIDIA A100-SXM4-40GB,40960,38912,maximum externally sampled per-device memory9est-qwen25-7b-math-c3-4gpu,est-qwen25-7b-math-4gpu,prospective_estimation,qwen25_7b,Qwen2.5-7B-Instruct,Qwen/Qwen2.5-7B-Instruct,math,MATH mathematics,math,GRPO,LoRA,c3,,4,1,16,8,8,0.7,False,False,True,512,512,1024,128,2,64,64,64,1,100,True,1,-1,-1,False,disable,1,NVIDIA A100-SXM4-40GB,40960,38912,maximum externally sampled per-device memory10est-qwen25-7b-code_heavy_tail-c2-2gpu,est-qwen25-7b-code_heavy_tail-2gpu,prospective_estimation,qwen25_7b,Qwen2.5-7B-Instruct,Qwen/Qwen2.5-7B-Instruct,longer_prompt_code,Longer-prompt code,code_heavy_tail,GRPO,LoRA,c2,,2,1,8,4,4,0.6,False,False,True,1536,1024,2560,64,2,32,32,32,1,100,True,1,-1,-1,False,disable,1,NVIDIA A100-SXM4-40GB,40960,38912,maximum externally sampled per-device memory11est-qwen25-7b-code_heavy_tail-c2-4gpu,est-qwen25-7b-code_heavy_tail-4gpu,prospective_estimation,qwen25_7b,Qwen2.5-7B-Instruct,Qwen/Qwen2.5-7B-Instruct,longer_prompt_code,Longer-prompt code,code_heavy_tail,GRPO,LoRA,c2,,4,1,8,4,4,0.6,False,False,True,1536,1024,2560,64,2,32,32,32,1,100,True,1,-1,-1,False,disable,1,NVIDIA A100-SXM4-40GB,40960,38912,maximum externally sampled per-device memory12est-qwen25-7b-code_heavy_tail-c3-2gpu,est-qwen25-7b-code_heavy_tail-2gpu,prospective_estimation,qwen25_7b,Qwen2.5-7B-Instruct,Qwen/Qwen2.5-7B-Instruct,longer_prompt_code,Longer-prompt code,code_heavy_tail,GRPO,LoRA,c3,,2,1,16,4,4,0.7,False,False,True,1536,1024,2560,64,2,32,32,32,1,100,True,1,-1,-1,False,disable,1,NVIDIA A100-SXM4-40GB,40960,38912,maximum externally sampled per-device memory13est-qwen25-7b-code_heavy_tail-c3-4gpu,est-qwen25-7b-code_heavy_tail-4gpu,prospective_estimation,qwen25_7b,Qwen2.5-7B-Instruct,Qwen/Qwen2.5-7B-Instruct,longer_prompt_code,Longer-prompt code,code_heavy_tail,GRPO,LoRA,c3,,4,1,16,4,4,0.7,False,False,True,1536,1024,2560,64,2,32,32,32,1,100,True,1,-1,-1,False,disable,1,NVIDIA A100-SXM4-40GB,40960,38912,maximum externally sampled per-device memory14 