datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining.
epfl-llm_guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/minsu/epfl-llm_guidelines.zip2zip-wikitext-repeat-phi35
Zip2Zip Repeated WikiText Stress Tests (Phi-3.5)
This repository contains two controlled evaluation corpora for studying
merge-size transfer in Zip2Zip models. They are derived from the document-level
WikiText-2 raw test split and built specifically with the
microsoft/Phi-3.5-mini-instruct tokenizer.
Configurations
Config
Repetitions per source block
Rows
Repeated base tokens
SHA-256 of test.jsonl
repeat4
4
1,329
1,297,250… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/zip2zip-wikitext-repeat-phi35.epfl_guidelines_dutch_marianmt
Dataset Card for Epfl English Guidelines Translated To Dutch With MariaNMT
This dataset was created by the EPFL, and can found in it original form here
The source language: English
The original data source: Original Data Source
The MariaNMT model used can be found: here
Data description
Translation of the English medical guidelines that are part of the Meditron corpus, using the LLM GPT 4o mini
Acknowledgement
This is part of the DT4H project with… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/epfl_guidelines_dutch_marianmt.
