datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fully-open-meditron
Fully Open Meditron Corpus
👋 Join our LiGHT community.
📖 Check out the MeditronFO blog and MeditronFO preprint.
🔜 If you are a clinician join the MOOVE initiative here.
[Hugging Face]
[Preprint]
[GitHub]
[Dataset]
License: Apache 2.0 | Authors: LiGHT
[!Note]
A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs.
The… See the full description on the dataset page: https://huggingface.co/datasets/EPFLiGHT/fully-open-meditron.fully-open-meditron
Fully Open Meditron Corpus
A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs (anonymous submission to NeurIPS 2026 Evaluations & Datasets Track).
The corpus combines eight aggregated public medical QA datasets with three clinician-vetted synthetic components, totaling approximately 601k examples (~150M tokens). It is designed to support supervised fine-tuning of large language… See the full description on the dataset page: https://huggingface.co/datasets/meditron-fo-anon/fully-open-meditron.fully-open-meditron
Fully Open Meditron Corpus
👋 Join our LiGHT community.
📖 Check out the MeditronFO blog and MeditronFO preprint.
🔜 If you are a clinician join the MOOVE initiative here.
[Hugging Face]
[Preprint]
[GitHub]
[Dataset]
License: Apache 2.0 | Authors: LiGHT
[!Note]
A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs.
The… See the full description on the dataset page: https://huggingface.co/datasets/npario/fully-open-meditron.corpus_meditronFO_preformat
full_v1 — preformatted EPFLiGHT/fully-open-meditron
This is EPFLiGHT/fully-open-meditron
(601k single-turn QA rows) turned into structured cases for multi-turn
dialogue generation: each row is unfolded into an opening line, a set of
atomic facts split by disclosure channel (spoken / observed / withheld), and
a question, so a doctor agent can elicit the case rather than read it off in
one shot.
preformat_full_v1.jsonl — 236,114 cases ready for dialogue generation.… See the full description on the dataset page: https://huggingface.co/datasets/zacbrld/corpus_meditronFO_preformat.
