datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MODUS-15Modality
MODUS — 15-Modality Aligned Dataset
MODUS is a large-scale, pixel-aligned 15-modality dataset for any-to-any
multimodal training. Every sample aligns 15 modalities covering appearance,
geometry, structure, segmentation, detection, text, and learned features.
Paper: https://huggingface.co/papers/2607.25948
Code: https://github.com/EPFL-VILAB/Modus
Modalities
Group
Modalities
Appearance
rgb, caption
Geometry
depth, normal
Structure
canny, sam_edge… See the full description on the dataset page: https://huggingface.co/datasets/epfl-vilab-modus/MODUS-15Modality.JSONSchemaBench
JSONSchemaBench
JSONSchemaBench is a benchmark of real-world JSON schemas designed to evaluate structured output generation for Large Language Models (LLMs). It contains approximately 10,000 JSON schemas, capturing diverse constraints and complexities.
import datasets
from datasets import load_dataset
def main():
# Inspect the available subsets of the datasetall_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench")
print("Available subsets:"… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/JSONSchemaBench.CanadaFireSat
Dataset Card for CanadaFireSat 🔥🛰️
In this benchmark, we investigate the potential of deep learning with multiple modalities for high-resolution wildfire forecasting. Leveraging different data settings across two types of model architectures: CNN-based and ViT-based.
📝 Published paper from ISPRS (ArXiv Version)
💿 Dataset repository on GitHub
🤖 Model repository on GitHub & Weights on Hugging Face
🟰 Another "Raw" version of the data with NPY files organized in different… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/CanadaFireSat.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.svi-benchmark
Stable Video Infinity (SVI) Benchmark Dataset
This benchmark dataset is introduced in the paper:
Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
by Wuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao, Alexandre Alahi (2025).
Project page: https://stable-video-infinity.github.io/homepage/
Code: https://github.com/vita-epfl/Stable-Video-Infinity
Abstract
We propose Stable Video Infinity (SVI) that is able to generate infinite-length videos with… See the full description on the dataset page: https://huggingface.co/datasets/epfl-vita/svi-benchmark.neurips-spectraThe dataset from Albert's et al, downloaded from zenodo. It's on here for easier access and organisation.
fully-open-meditron
Fully Open Meditron Corpus
👋 Join our LiGHT community.
📖 Check out the MeditronFO blog and MeditronFO preprint.
🔜 If you are a clinician join the MOOVE initiative here.
[Hugging Face]
[Preprint]
[GitHub]
[Dataset]
License: Apache 2.0 | Authors: LiGHT
[!Note]
A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs.
The… See the full description on the dataset page: https://huggingface.co/datasets/EPFLiGHT/fully-open-meditron.EcoWikiRS
EcoWikiRS: Learning Ecological Representations of Satellite Images from Weak Supervision with Species Observations and Wikipedia
AuthorsValerie Zermatten · Javiera Castillo-Navarro · Pallavi Jain · Devis Tuia · Diego Marcos
Overview
The WikiRS dataset, composed of triplets of images, species list and Wikipedia sentences :
91k high-resolution aerial images (50cm, RGB bands) from the swissIMAGE product
crowd-sourced species observations from 2745 different… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/EcoWikiRS.zip2zip-1Bnmrshiftdb2epfl-smart-kitchen-av1
EPFL-Smart-Kitchen AV1 SimpleCV Mirror
This is an AV1-transcoded SimpleCV-compatible mirror of the EPFL-Smart-Kitchen-30 dataset.
Original data:
Collected videos/data: https://zenodo.org/records/15535461
Poses/annotations: https://zenodo.org/records/15551913
GitHub: https://github.com/amathislab/EPFL-Smart-Kitchen
Contents
RGB and HoloLens videos transcoded to AV1 MP4
Depth videos transcoded to AV1 MP4
Metadata, timestamps, IMUs, poses, and annotations preserved… See the full description on the dataset page: https://huggingface.co/datasets/pablovela5620/epfl-smart-kitchen-av1.llaza-20B
Llaza Mixture 20B
This dataset is a 20B-token pretraining subset built for zip2zip language-model pretraining.
It is derived from the full Llaza mixture, which is byte-balanced across four top-level domains:
Domain
Source
Target byte ratio
General
HuggingFaceFW/fineweb-edu, sample-100BT
50%
Code
bigcode/the-stack-dedup
20%
Math
HuggingFaceTB/finemath, finemath-3plus
10%
Multilingual
epfml/FineWeb2-HQ, 20 language subsets
20%
The subset was created from remixed… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/llaza-20B.zip2zip-1B-no-split
HuggingFaceFW/fineweb-edu (20%) (common knowledge)
devngho/the-stack-llm-annotations-v2 (25%) (code)
AI-MO/NuminaMath-1.5 (20%) (math)
HuggingFaceH4/ultrachat_200k (20%) (chat)
HuggingFaceFW/fineweb-2 (15%) (multilingual: [cmn_Hani, deu_Latn, jpn_Jpan, spa_Latn, fra_Latn, ita_Latn, por_Latn, nld_Latn, arb_Arab])
K600-MM
K600-MM
K600-MM is a multimodal video dataset used to pretrain A2A-Video, an any-to-any multimodal model for the video domain. It's built on top of Kinetics-600 (RGB video + class-category annotations) with 10 additional modalities obtained via pseudo-labeling. Refer to A2A-Video's README_DATA.md for additional details on dataset construction.
In total, there are ~392K training and ~30K validation/test video clips with 12 aligned modalities per clip.
The provided train/test data… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-VILAB/K600-MM.nmrexpMNLP_M3_mcqa_dataset
Tulu 3 SFT Mixture (Sampled)
This dataset is a sampled and filtered subset of the allenai/tulu-3-sft-mixture, curated and rebalanced for structured instruction fine-tuning. The goal is to support research and model development in math reasoning, coding, knowledge recall, instruction following (IF), and conversational alignment, while explicitly excluding safety, multilingual, and certain task-specific sources.
📦 Dataset Structure
Source: Filtered from… See the full description on the dataset page: https://huggingface.co/datasets/vanek-epfl/MNLP_M3_mcqa_dataset.llaza-200B
Llaza Mixture Full (200B)
This dataset is the full Llaza pretraining-data mixture for zip2zip language-model pretraining.
It combines general web text, code, math, and multilingual web text with byte-based top-level mixture ratios.
Domain
Source
Target byte ratio
General
HuggingFaceFW/fineweb-edu, sample-100BT
50%
Code
bigcode/the-stack-dedup
20%
Math
HuggingFaceTB/finemath, finemath-3plus
10%
Multilingual
epfml/FineWeb2-HQ, 20 language subsets
20%… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/llaza-200B.epfl-enterprise-osai-adoption-research-data
EPFL Enterprise Open-Source AI Adoption Research Dataset
Dataset Summary
This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption.
Dataset Structure
This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.MassSpecGym
MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from MS/MS spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems.
Please refer to the MassSpecGym GitHub page and the paper for details.
epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining.
epfl-computer-science-mcqaepfl-math-step-10kMNLP_M2_mcqa_dataset
Smol-SmalTalk
This is a subset of SmolTalk dataset adapted for smol models with less than 1B parameters. We used it to build SmolLM2-360M-Instruct and
SmolLM2-135M-Instruct. We do SFT on this dataset and then DPO on UltraFeedback.
Compared to SmolTalk:
The conversations from Smol-Magpie-Ultra are shorter in this dataset
We include less task specific data compared to SmolTalk (e.g no function calling and less rewriting and summarization examples) since these smaller models have… See the full description on the dataset page: https://huggingface.co/datasets/vanek-epfl/MNLP_M2_mcqa_dataset.stack_exchange_epfl_data_2048include-89tulu3-sft-mixture-sampled
Tulu 3 SFT Mixture (Sampled)
This dataset is a sampled and filtered subset of the allenai/tulu-3-sft-mixture, curated and rebalanced for structured instruction fine-tuning. The goal is to support research and model development in math reasoning, coding, knowledge recall, instruction following (IF), and conversational alignment, while explicitly excluding safety, multilingual, and certain task-specific sources.
📦 Dataset Structure
Source: Filtered from… See the full description on the dataset page: https://huggingface.co/datasets/vanek-epfl/tulu3-sft-mixture-sampled.multimodal-brain-scaling
Multimodal Scaling Laws for Task & Data-Optimized Models of Visual Cortex
This repository hosts the result tables and accompanying metadata released
with our ICML 2026 paper
Multimodal Scaling Laws for Task & Data-Optimized Models of Visual Cortex.
Paper summary
Task-optimized deep networks are the leading in-silico models of sensory
cortex, but progress is fragmented across datasets, modalities, and
evaluation protocols, making it hard to identify which… See the full description on the dataset page: https://huggingface.co/datasets/epfl-neuroai/multimodal-brain-scaling.TST-ProcTHOR
Dataset Card for TST-ProcTHOR
Dataset Summary
This custom TST-ProcTHOR dataset is used in research work "Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality".
pretrain/ is a multimodal pretraining dataset collected using ProcTHOR environment. It contains RGB images, and 9 additional tokenized modalities.
segmentation/train is the associated downstream dataset used to finetune TST pretrained models on semantic… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-VILAB/TST-ProcTHOR.dpo_preference_dataset_EPFL_shpepfl_mnlp_dpo_evaluation_dataset
