CoolFace
Datasetpublic

pranavvmurthy26/synthetic-financial-tool-calling-grpo-rlvr-1k

🤖 Synthetic Financial Tool Calling Dataset for GRPO and RLVR This is a synthetic dataset designed for training language models on financial tool calling using GRPO (Group Relative Policy Optimization) with verifiable rewards (RLVR). The dataset contains ~1.1K examples of financial planning queries paired with expected tool calls and answers. Dataset sample schema, { "prompt": [ { "role": "system", "content": "You are a financial planning assistant with tools… See the full description on the dataset page: https://huggingface.co/datasets/pranavvmurthy26/synthetic-financial-tool-calling-grpo-rlvr-1k.

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
2likes45downloads
8 commits on main
fa8e9258mo ago

Update README.md

pranavvmurthy26
1c901568mo ago

Update README.md

pranavvmurthy26
9c610188mo ago

Update README.md

pranavvmurthy26
203fbca8mo ago

Update README.md

pranavvmurthy26
9ade4248mo ago

Upload financial_tools_reward.py

pranavvmurthy26
71cf81a8mo ago

Upload financial_tools_complex.py

pranavvmurthy26
59372488mo ago

Upload dataset

pranavvmurthy26
00e9fb68mo ago

initial commit

pranavvmurthy26