datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fully-open-meditron
Fully Open Meditron Corpus
👋 Join our LiGHT community.
📖 Check out the MeditronFO blog and MeditronFO preprint.
🔜 If you are a clinician join the MOOVE initiative here.
[Hugging Face]
[Preprint]
[GitHub]
[Dataset]
License: Apache 2.0 | Authors: LiGHT
[!Note]
A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs.
The… See the full description on the dataset page: https://huggingface.co/datasets/EPFLiGHT/fully-open-meditron.details_malhajar__Mistral-7B-v0.2-meditron-turkish
Dataset Card for Evaluation run of malhajar/Mistral-7B-v0.2-meditron-turkish
Dataset automatically created during the evaluation run of model malhajar/Mistral-7B-v0.2-meditron-turkish on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_malhajar__Mistral-7B-v0.2-meditron-turkish.fully-open-meditron
Fully Open Meditron Corpus
A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs (anonymous submission to NeurIPS 2026 Evaluations & Datasets Track).
The corpus combines eight aggregated public medical QA datasets with three clinician-vetted synthetic components, totaling approximately 601k examples (~150M tokens). It is designed to support supervised fine-tuning of large language… See the full description on the dataset page: https://huggingface.co/datasets/meditron-fo-anon/fully-open-meditron.details_malhajar__meditron-7b-chat
Dataset Card for Evaluation run of malhajar/meditron-7b-chat
Dataset automatically created during the evaluation run of model malhajar/meditron-7b-chat on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_malhajar__meditron-7b-chat.fully-open-meditron
Fully Open Meditron Corpus
👋 Join our LiGHT community.
📖 Check out the MeditronFO blog and MeditronFO preprint.
🔜 If you are a clinician join the MOOVE initiative here.
[Hugging Face]
[Preprint]
[GitHub]
[Dataset]
License: Apache 2.0 | Authors: LiGHT
[!Note]
A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs.
The… See the full description on the dataset page: https://huggingface.co/datasets/npario/fully-open-meditron.meditron-clinical-guidelinescorpus_meditronFO_preformat
full_v1 — preformatted EPFLiGHT/fully-open-meditron
This is EPFLiGHT/fully-open-meditron
(601k single-turn QA rows) turned into structured cases for multi-turn
dialogue generation: each row is unfolded into an opening line, a set of
atomic facts split by disclosure channel (spoken / observed / withheld), and
a question, so a doctor agent can elicit the case rather than read it off in
one shot.
preformat_full_v1.jsonl — 236,114 cases ready for dialogue generation.… See the full description on the dataset page: https://huggingface.co/datasets/zacbrld/corpus_meditronFO_preformat.meditron-trmeditron-7b-lora-huda-trainmedmcqa-grpo-meditron70bpubmedqa-meditron-conversations-labeledpubmedqa-meditron-conversations-annotated2pubmedqa-meditron-conversations-annotated-claudepubmedqa-meditron-conversations-annotated-claude-langmeditron-7b-lora-huda-valdiffing-stats-gemma-2-2b-it-Meditron3-L16-k100-lr1e-04-local-shuffling-CCLossmedmcqa-grpo-meditron8bpubmedqa-meditron-conversations-annotated-gptmeditrondiffing-stats-gemma-2-2b-it-Meditron3-L16-mu3.8e-02-lr1e-04-local-shuffling-CCLosspubmedqa-meditron-conversations-annotated-claude-czpubmedqa-meditron-conversationsmeditron-finetune-v1
