CoolFace
Datasetpublic

pranavvmurthy26/synthetic-healthcare-tool-calling-grpo-rlvr-1k

🏥 Synthetic Healthcare Tool Calling Dataset for GRPO and RLVR This is a synthetic dataset designed for training language models on clinical decision support tool calling using GRPO (Group Relative Policy Optimization) with verifiable rewards (RLVR). The dataset contains ~1.1K examples of clinical scenarios paired with expected tool calls and answers. Dataset sample schema: { "prompt": [ { "role": "system", "content": "You are a clinical decision support… See the full description on the dataset page: https://huggingface.co/datasets/pranavvmurthy26/synthetic-healthcare-tool-calling-grpo-rlvr-1k.

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes54downloads
Dataset Card

🏥 Synthetic Healthcare Tool Calling Dataset for GRPO and RLVR

This is a synthetic dataset designed for training language models on clinical decision support tool calling using GRPO (Group Relative Policy Optimization) with verifiable rewards (RLVR). The dataset contains ~1.1K examples of clinical scenarios paired with expected tool calls and answers.

Dataset sample schema:

json
{
  "prompt": [
    {
      "role": "system",
      "content": "You are a clinical decision support assistant with tools for drug dosage calculation, cardiovascular risk assessment, BMI and metabolic risk evaluation, surgical risk estimation, insulin regimen calculation, kidney function assessment, fluid replacement planning, and drug interaction screening. Analyze clinical scenarios and call the appropriate tool with all required parameters extracted from the patient presentation. Return concise clinical recommendations with key metrics. Do not ask for clarification - use reasonable clinical defaults if needed."
    },
    {
      "role": "user",
      "content": "Can you dose vancomycin for a 65yo male patient? He's 85kg, creatinine 1.5, suspected MRSA pneumonia. No drug allergies. Not a loading dose."
    }
  ],
  "answer": "Drug: Vancomycin 650.0mg IV q24h, CrCl: 59.0 mL/min, Renal tier: moderate, Dose factor: 0.65, Indication: suspected MRSA pneumonia, Adjustment: Reduce dose, extend interval",
  "ground_truth": {
    "name": "calculate_drug_dosage",
    "arguments": {
      "drug_name": "vancomycin",
      "standard_dose_mg": 1000,
      "route": "IV",
      "patient_weight_kg": 85,
      "patient_age": 65,
      "serum_creatinine": 1.5,
      "sex": "male",
      "allergies": [],
      "indication": "suspected MRSA pneumonia",
      "frequency_override": null,
      "max_daily_dose_mg": null,
      "is_loading_dose": false
    }
  }
}

Fine-tuning an LLM using Reinforcement learning leverages prompt and uses answer optionally to verify reward.

📄 Schema

ColumnTypeDescription
promptlist[dict]A conversation-style prompt with role (system/user) and content fields. The system message defines the assistant's clinical capabilities, and the user message contains a natural language patient presentation.
answerstringThe expected human-readable output from executing the correct tool call (e.g., renal-adjusted drug dose with CrCl and dosing interval).
ground_truthstringA JSON string containing the exact tool call specification with name (function name) and arguments (parameter dictionary) that should be invoked to answer the query.

📊 How the Data is Used

The dataset is used with a GRPO trainer for tool-calling optimization:

  1. 1.Dataset Loading: Loads prompts, answers, and ground truth tool calls
  2. 2.Model Generation: The model generates completions (tool calls) given the prompts
  3. 3.Tool Execution: Generated tool calls are executed against actual clinical tool functions
  4. 4.Reward Computation: The reward function compares tool execution results against expected answers
  5. 5.Policy Optimization: GRPO uses rewards to optimize tool-calling behavior through relative comparisons across generations

🛠️ Tool Description and Rewards

Tools

Reference: `healthcare_tools_complex.py`

The dataset targets 8 clinical decision support functions with complex argument structures (8-12 parameters each):

ToolPurpose
calculate_drug_dosageRenal-adjusted drug dosing using Cockcroft-Gault CrCl with allergy cross-reactivity checks
assess_cardiovascular_risk10-year CVD risk using Framingham Risk Score
calculate_bmi_metabolic_riskBMI classification and metabolic syndrome screening using ATP III criteria
estimate_surgical_riskPerioperative risk using ASA classification and RCRI score
calculate_insulin_regimenBasal-bolus insulin regimen using 450/500 and 1800/2000 rules
assess_kidney_functionKidney function assessment using CrCl, MDRD, and CKD-EPI formulas
calculate_fluid_replacementIV fluid replacement using Parkland formula and 4-2-1 maintenance rule
screen_drug_interactionsDrug-drug interaction screening via CYP450 pathway analysis

Each function returns a deterministic string result, enabling exact-match reward computation during training.

📈 Reward Function

The reward function implements a 3-tier scheme:

RewardCondition
1.0Tool response exactly matches the expected answer
0.1A tool was called but the response doesn't match
0.0No tool call was made

This structure encourages the model to: (1) learn to make tool calls, and (2) learn to make the correct tool calls with proper arguments.

🦾 Dataset generation

This dataset was generated with the help of kiro.dev.

Citation

If you use this dataset, please cite:

bibtex
@dataset{pranavvmurthy26-synthetic-healthcare-tool-calling-grpo-rlvr-1k,
  title={Synthetic Healthcare Tool Calling GRPO RLVR Dataset},
  author={DeFauw, Randy and Murthy, Pranav},
  year={2026}
}