datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orthoqa-300
OrthoQA-300
You can access the dataset on Hugging Face and find the full generation pipeline, configuration files, and source code in the Stratum Research GitHub repository.
OrthoQA-300 is a structured, synthetic dataset of 300 patient-provider style question-and-answer (QA) pairs focused on orthopedic surgery. Each entry simulates a realistic clinical interaction, with patient-style questions and LLM-generated provider-style answers.
Questions are grouped by procedure (e.g., ACL… See the full description on the dataset page: https://huggingface.co/datasets/stratum-research/orthoqa-300.medqa-test-tagged-by-specialitywe use gpt-4o-mini as a base model to tag each medqa question by medical speciality, along with other fields. The goal is to use these tags as a basis for creating subsets of medqa to test specific model capabilities and knowledge.
In depth ReadMe to come
pls cite if used
