datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RLVRAMBench
RLVRAMBench
Which language-model training configurations can I use with the memory
I have, and how much testing does that decision require?
RLVRAMBench is a measurement dataset with open evaluation tasks for a
specific language-model training system. It measures memory feasibility
when response generation and reinforcement-learning updates share the
same graphics processors. It provides measured outcomes, fixed prediction
tasks, a budgeted decision replay, reference methods, and… See the full description on the dataset page: https://huggingface.co/datasets/kobzaond/RLVRAMBench.Chem-RLVR
Chem-RLVR
Chem-RLVR is a benchmark for reasoning over experimental reaction records,
with direct reaction-yield prediction and counterfactual yield-reasoning tasks.
Paper: "Chem-RLVR: Verifier-Based Training for Reaction Yield Prediction"
Dataset
Chem-RLVR contains 12,000 questions:
6,000 yield-prediction questions
6,000 counterfactual questions
Each task contains:
4,200 RL-training examples
1,800 frozen held-out evaluation examples
The benchmark covers three… See the full description on the dataset page: https://huggingface.co/datasets/redugo/Chem-RLVR.
