datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JSONSchemaBench
JSONSchemaBench
JSONSchemaBench is a benchmark of real-world JSON schemas designed to evaluate structured output generation for Large Language Models (LLMs). It contains approximately 10,000 JSON schemas, capturing diverse constraints and complexities.
import datasets
from datasets import load_dataset
def main():
# Inspect the available subsets of the datasetall_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench")
print("Available subsets:"… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/JSONSchemaBench.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.fully-open-meditron
Fully Open Meditron Corpus
👋 Join our LiGHT community.
📖 Check out the MeditronFO blog and MeditronFO preprint.
🔜 If you are a clinician join the MOOVE initiative here.
[Hugging Face]
[Preprint]
[GitHub]
[Dataset]
License: Apache 2.0 | Authors: LiGHT
[!Note]
A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs.
The… See the full description on the dataset page: https://huggingface.co/datasets/EPFLiGHT/fully-open-meditron.llaza-20B
Llaza Mixture 20B
This dataset is a 20B-token pretraining subset built for zip2zip language-model pretraining.
It is derived from the full Llaza mixture, which is byte-balanced across four top-level domains:
Domain
Source
Target byte ratio
General
HuggingFaceFW/fineweb-edu, sample-100BT
50%
Code
bigcode/the-stack-dedup
20%
Math
HuggingFaceTB/finemath, finemath-3plus
10%
Multilingual
epfml/FineWeb2-HQ, 20 language subsets
20%
The subset was created from remixed… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/llaza-20B.llaza-200B
Llaza Mixture Full (200B)
This dataset is the full Llaza pretraining-data mixture for zip2zip language-model pretraining.
It combines general web text, code, math, and multilingual web text with byte-based top-level mixture ratios.
Domain
Source
Target byte ratio
General
HuggingFaceFW/fineweb-edu, sample-100BT
50%
Code
bigcode/the-stack-dedup
20%
Math
HuggingFaceTB/finemath, finemath-3plus
10%
Multilingual
epfml/FineWeb2-HQ, 20 language subsets
20%… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/llaza-200B.epfl-enterprise-osai-adoption-research-data
EPFL Enterprise Open-Source AI Adoption Research Dataset
Dataset Summary
This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption.
Dataset Structure
This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining.
epfl-llm_guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/minsu/epfl-llm_guidelines.zip2zip-wikitext-repeat-phi35
Zip2Zip Repeated WikiText Stress Tests (Phi-3.5)
This repository contains two controlled evaluation corpora for studying
merge-size transfer in Zip2Zip models. They are derived from the document-level
WikiText-2 raw test split and built specifically with the
microsoft/Phi-3.5-mini-instruct tokenizer.
Configurations
Config
Repetitions per source block
Rows
Repeated base tokens
SHA-256 of test.jsonl
repeat4
4
1,329
1,297,250… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/zip2zip-wikitext-repeat-phi35.epfl_guidelines_dutch_marianmt
Dataset Card for Epfl English Guidelines Translated To Dutch With MariaNMT
This dataset was created by the EPFL, and can found in it original form here
The source language: English
The original data source: Original Data Source
The MariaNMT model used can be found: here
Data description
Translation of the English medical guidelines that are part of the Meditron corpus, using the LLM GPT 4o mini
Acknowledgement
This is part of the DT4H project with… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/epfl_guidelines_dutch_marianmt.
