datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ToRL-MathdeepsearchSkyRL-SQL-Reproductionopenmathreasoning_tirAceCoderV2-69K-cleanedarpo_dataAceCoderV2-122Kopenmathreasoning_tir_100K
Dataset Card for Dataset Name
It has been filtered from openreasoning_math_max_len_4096 through df[(df['pass_rate_72b_tir'] > 0) & (df['pass_rate_72b_tir'] < 0.8)]
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/VerlTool/openmathreasoning_tir_100K.toolrl-4k-verl
ToolRL Dataset - GPT OSS 120B Format
A preprocessed tool-learning dataset in GPT OSS 120B native format for reinforcement learning training with GRPO/PPO algorithms.
Dataset Description
This dataset contains 4,000 tool-use samples (3,920 training / 80 test) converted from the ToolRL dataset to GPT OSS 120B's native format. The conversion replaces XML-style tags with GPT OSS's special tokens and channel system, resulting in ~10-15% token efficiency improvement.… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/toolrl-4k-verl.AceCoderV2-69KAceCoderV2-122K-cleanedInterpreter-Thinking
