CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01epfl-dlab /JSONSchemaBench JSONSchemaBench JSONSchemaBench is a benchmark of real-world JSON schemas designed to evaluate structured output generation for Large Language Models (LLMs). It contains approximately 10,000 JSON schemas, capturing diverse constraints and complexities. import datasets from datasets import load_dataset def main(): # Inspect the available subsets of the datasetall_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench") print("Available subsets:"… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/JSONSchemaBench.texttext-generation10K<n<100K12 likes4.2k downloads1y agoHugging Face02epfl-llm /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.texttext-generation10K<n<100K158 likes3k downloads3y agoHugging Face03EPFLiGHT /fully-open-meditron Fully Open Meditron Corpus 👋 Join our LiGHT community. 📖 Check out the MeditronFO blog and MeditronFO preprint. 🔜 If you are a clinician join the MOOVE initiative here. [Hugging Face] [Preprint] [GitHub] [Dataset] License: Apache 2.0 | Authors: LiGHT [!Note] A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs. The… See the full description on the dataset page: https://huggingface.co/datasets/EPFLiGHT/fully-open-meditron.textquestion-answering100K<n<1M8 likes332 downloads3mo agoHugging Face04epfl-dlab /llaza-20B Llaza Mixture 20B This dataset is a 20B-token pretraining subset built for zip2zip language-model pretraining. It is derived from the full Llaza mixture, which is byte-balanced across four top-level domains: Domain Source Target byte ratio General HuggingFaceFW/fineweb-edu, sample-100BT 50% Code bigcode/the-stack-dedup 20% Math HuggingFaceTB/finemath, finemath-3plus 10% Multilingual epfml/FineWeb2-HQ, 20 language subsets 20% The subset was created from remixed… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/llaza-20B.texttext-generation10M<n<100M0 likes156 downloads5mo agoHugging Face05epfl-dlab /llaza-200B Llaza Mixture Full (200B) This dataset is the full Llaza pretraining-data mixture for zip2zip language-model pretraining. It combines general web text, code, math, and multilingual web text with byte-based top-level mixture ratios. Domain Source Target byte ratio General HuggingFaceFW/fineweb-edu, sample-100BT 50% Code bigcode/the-stack-dedup 20% Math HuggingFaceTB/finemath, finemath-3plus 10% Multilingual epfml/FineWeb2-HQ, 20 language subsets 20%… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/llaza-200B.texttext-generation100M<n<1B0 likes79 downloads5mo agoHugging Face06itseffi /epfl-enterprise-osai-adoption-research-data EPFL Enterprise Open-Source AI Adoption Research Dataset Dataset Summary This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption. Dataset Structure This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.tabulartext-classificationn<1K0 likes70 downloads1y agoHugging Face07PJMixers /epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining. texttext-generation10K<n<100K0 likes49 downloads3y agoHugging Face08minsu /epfl-llm_guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/minsu/epfl-llm_guidelines.texttext-generation10K<n<100K0 likes25 downloads7mo agoHugging Face09epfl-dlab /zip2zip-wikitext-repeat-phi35 Zip2Zip Repeated WikiText Stress Tests (Phi-3.5) This repository contains two controlled evaluation corpora for studying merge-size transfer in Zip2Zip models. They are derived from the document-level WikiText-2 raw test split and built specifically with the microsoft/Phi-3.5-mini-instruct tokenizer. Configurations Config Repetitions per source block Rows Repeated base tokens SHA-256 of test.jsonl repeat4 4 1,329 1,297,250… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/zip2zip-wikitext-repeat-phi35.tabulartext-generation1K<n<10K0 likes22 downloads1mo agoHugging Face10UMCU /epfl_guidelines_dutch_marianmt Dataset Card for Epfl English Guidelines Translated To Dutch With MariaNMT This dataset was created by the EPFL, and can found in it original form here The source language: English The original data source: Original Data Source The MariaNMT model used can be found: here Data description Translation of the English medical guidelines that are part of the Meditron corpus, using the LLM GPT 4o mini Acknowledgement This is part of the DT4H project with… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/epfl_guidelines_dutch_marianmt.tabulartext-generation10K<n<100K0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.