pranavvmurthy26/synthetic-healthcare-tool-calling-grpo-rlvr-1k
🏥 Synthetic Healthcare Tool Calling Dataset for GRPO and RLVR This is a synthetic dataset designed for training language models on clinical decision support tool calling using GRPO (Group Relative Policy Optimization) with verifiable rewards (RLVR). The dataset contains ~1.1K examples of clinical scenarios paired with expected tool calls and answers. Dataset sample schema: { "prompt": [ { "role": "system", "content": "You are a clinical decision support… See the full description on the dataset page: https://huggingface.co/datasets/pranavvmurthy26/synthetic-healthcare-tool-calling-grpo-rlvr-1k.
🏥 Synthetic Healthcare Tool Calling Dataset for GRPO and RLVR
This is a synthetic dataset designed for training language models on clinical decision support tool calling using GRPO (Group Relative Policy Optimization) with verifiable rewards (RLVR). The dataset contains ~1.1K examples of clinical scenarios paired with expected tool calls and answers.
Dataset sample schema:
{
"prompt": [
{
"role": "system",
"content": "You are a clinical decision support assistant with tools for drug dosage calculation, cardiovascular risk assessment, BMI and metabolic risk evaluation, surgical risk estimation, insulin regimen calculation, kidney function assessment, fluid replacement planning, and drug interaction screening. Analyze clinical scenarios and call the appropriate tool with all required parameters extracted from the patient presentation. Return concise clinical recommendations with key metrics. Do not ask for clarification - use reasonable clinical defaults if needed."
},
{
"role": "user",
"content": "Can you dose vancomycin for a 65yo male patient? He's 85kg, creatinine 1.5, suspected MRSA pneumonia. No drug allergies. Not a loading dose."
}
],
"answer": "Drug: Vancomycin 650.0mg IV q24h, CrCl: 59.0 mL/min, Renal tier: moderate, Dose factor: 0.65, Indication: suspected MRSA pneumonia, Adjustment: Reduce dose, extend interval",
"ground_truth": {
"name": "calculate_drug_dosage",
"arguments": {
"drug_name": "vancomycin",
"standard_dose_mg": 1000,
"route": "IV",
"patient_weight_kg": 85,
"patient_age": 65,
"serum_creatinine": 1.5,
"sex": "male",
"allergies": [],
"indication": "suspected MRSA pneumonia",
"frequency_override": null,
"max_daily_dose_mg": null,
"is_loading_dose": false
}
}
}Fine-tuning an LLM using Reinforcement learning leverages prompt and uses answer optionally to verify reward.
📄 Schema
📊 How the Data is Used
The dataset is used with a GRPO trainer for tool-calling optimization:
- Dataset Loading: Loads prompts, answers, and ground truth tool calls
- Model Generation: The model generates completions (tool calls) given the prompts
- Tool Execution: Generated tool calls are executed against actual clinical tool functions
- Reward Computation: The reward function compares tool execution results against expected answers
- Policy Optimization: GRPO uses rewards to optimize tool-calling behavior through relative comparisons across generations
🛠️ Tool Description and Rewards
Tools
Reference: `healthcare_tools_complex.py`
The dataset targets 8 clinical decision support functions with complex argument structures (8-12 parameters each):
Each function returns a deterministic string result, enabling exact-match reward computation during training.
📈 Reward Function
The reward function implements a 3-tier scheme:
This structure encourages the model to: (1) learn to make tool calls, and (2) learn to make the correct tool calls with proper arguments.
🦾 Dataset generation
This dataset was generated with the help of kiro.dev.
Citation
If you use this dataset, please cite:
@dataset{pranavvmurthy26-synthetic-healthcare-tool-calling-grpo-rlvr-1k,
title={Synthetic Healthcare Tool Calling GRPO RLVR Dataset},
author={DeFauw, Randy and Murthy, Pranav},
year={2026}
}