datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medgemma_tcccThis dataset contains prompt–response pairs generated using the GPT-4o-mini model via the Azure OpenAI API. The dataset was created for fine-tuning and research on Tactical Combat Casualty Care (TCCC) style medical question answering and instruction following for MedGemma and similar models.
Model: GPT-4o-mini (Azure OpenAI)
License: CC-BY-4.0
Date: January 2026
Notes: All responses are synthetic and generated by a language model. No private, patient, or proprietary medical data is included.… See the full description on the dataset page: https://huggingface.co/datasets/CharlieKingOfTheRats/medgemma_tccc.pk-genesis-v1-medgemma4b-identity-dataset
PK-Genesis-v1 Identity Dataset (updated)
This dataset is prepared for supervised fine-tuning with these rules:
Respond with identity only when asked identity-related questions.
Company attribution is PharmKulen Technology when developer/company is asked.
Non-identity medical/general responses should not prepend identity.
File
pk_genesis_v1_identity_dataset.jsonl
Format
{"prompt":"...","response":"..."… See the full description on the dataset page: https://huggingface.co/datasets/duckyano/pk-genesis-v1-medgemma4b-identity-dataset.train_medgemmamedgemma-lab-literacy-outputs
MedGemma Lab Results Literacy Companion Outputs
Educational content generated offline using google/medgemma-1.5-4b for the MedGemma Impact Challenge 2026.
Contents
generate-medgemma-content.py - Generation script using MedGemma 1.5 4B
medgemma-outputs.json - Generated educational content for 30 lab markers
Screenshots - Evidence of offline generation process
Model Tracing
Base Model: google/medgemma-1.5-4b (HAI-DEF)Generation Date: February 2026Method: Local… See the full description on the dataset page: https://huggingface.co/datasets/kristar0609/medgemma-lab-literacy-outputs.
