CoolFace
Datasetpublicgated

alucent/mirror-Nemotron-RL-Agentic-Function-Calling-Pivot-v1

Dataset Description: This is a RL dataset for general function-calling by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Nemotron-RL-Agentic-Function-Calling-Pivot-v1.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes14downloads
settings

This repository belongs to alucent on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemirror-Nemotron-RL-Agentic-Function-Calling-Pivot-v1
visibilitypublic
licencecc-by-4.0
gatedyes
owneralucent
Account settings
alucent/mirror-Nemotron-RL-Agentic-Function-Calling-Pivot-v1 · CoolFace