datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PocketQubeMODUS-15Modality
MODUS — 15-Modality Aligned Dataset
MODUS is a large-scale, pixel-aligned 15-modality dataset for any-to-any
multimodal training. Every sample aligns 15 modalities covering appearance,
geometry, structure, segmentation, detection, text, and learned features.
Paper: https://huggingface.co/papers/2607.25948
Code: https://github.com/EPFL-VILAB/Modus
Modalities
Group
Modalities
Appearance
rgb, caption
Geometry
depth, normal
Structure
canny, sam_edge… See the full description on the dataset page: https://huggingface.co/datasets/epfl-vilab-modus/MODUS-15Modality.SwissCubeJSONSchemaBench
JSONSchemaBench
JSONSchemaBench is a benchmark of real-world JSON schemas designed to evaluate structured output generation for Large Language Models (LLMs). It contains approximately 10,000 JSON schemas, capturing diverse constraints and complexities.
import datasets
from datasets import load_dataset
def main():
# Inspect the available subsets of the datasetall_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench")
print("Available subsets:"… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/JSONSchemaBench.CanadaFireSat
Dataset Card for CanadaFireSat 🔥🛰️
In this benchmark, we investigate the potential of deep learning with multiple modalities for high-resolution wildfire forecasting. Leveraging different data settings across two types of model architectures: CNN-based and ViT-based.
📝 Published paper from ISPRS (ArXiv Version)
💿 Dataset repository on GitHub
🤖 Model repository on GitHub & Weights on Hugging Face
🟰 Another "Raw" version of the data with NPY files organized in different… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/CanadaFireSat.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.ShapeNetSDF
ShapeNetSDF
Signed Distance Field (SDF) point samples derived from
ShapeNet Core, for training and evaluating implicit
neural representations / neural fields on 3D shapes.
This dataset is shared as part of the CVPR 2026 paper Weight Space Representation Learning via Neural Field Adaptaion.
Code for producing this dataset is shared in the wsr.pytorch neural-field codebase.
Each shape is converted into a watertight manifold, normalized into the unit
cube [-1, 1]³, and sampled… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-IVRL/ShapeNetSDF.A2A-Video-examplesOpenMaterial
OpenMaterial: A Comprehensive Dataset of Complex Materials for 3D Reconstruction
Zheng Dang1 · Jialu Huang2 · Fei Wang2 · Mathieu Salzmann1
1EPFL CVLAb, Switzerland 2 Xi'an Jiaotong University, China
Paper
WebPage
📌 Update log
🗓️ March 2025
Updated degnosie scripts to identify and address rare missing cases caused by server-side cluster fluctuations.
Refined benchmark results for selected algorithms (NeRO, GES… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-CVLab/OpenMaterial.DrivingVQA
Dataset Card for DrivingVQA
🏠 Homepage
Dataset Details
Dataset Description
DrivingVQA is a dataset designed to assist candidates preparing for the French driving theory exam, which requires passing both a theoretical and a practical test. The theoretical aspect consists of analyzing 40 multiple-choice questions (MCQs) with real-world images to test the candidates' knowledge of traffic laws, road signs, and safe driving practices. This dataset focuses on visual… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-DrivingVQA/DrivingVQA.coralscapes
Coralscapes Dataset
The Coralscapes dataset is the first general-purpose dense semantic segmentation dataset for coral reefs. Similar in scope and with the same structure as the widely used Cityscapes dataset for urban scene understanding, Coralscapes allows for the benchmarking of semantic segmentation models in a new challenging domain.
Dataset Structure
The Coralscapes dataset spans 2075 images at 1024×2048px resolution… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/coralscapes.svi-benchmark
Stable Video Infinity (SVI) Benchmark Dataset
This benchmark dataset is introduced in the paper:
Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
by Wuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao, Alexandre Alahi (2025).
Project page: https://stable-video-infinity.github.io/homepage/
Code: https://github.com/vita-epfl/Stable-Video-Infinity
Abstract
We propose Stable Video Infinity (SVI) that is able to generate infinite-length videos with… See the full description on the dataset page: https://huggingface.co/datasets/epfl-vita/svi-benchmark.SwissView
Dataset Card for SwissView Dataset
Project Page
https://limirs.github.io/GeoExplorer/
GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
Dataset Summary
This dataset consists of two subsets:
SwissViewMonuments: which includes 15 images of atypical or distinctive scenes, such as unusual buildings and landscapes, with corresponding ground level images.
SwissView100, which comprises 100 images randomly selected from across the Swiss territory… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/SwissView.ValaisCD
ValaisCD Dataset
High-Resolution Aerial Change Detection (Switzerland, 2017–2023)
Project page: https://manonbechaz.github.io/2Player/
🗺️ Overview
ValaisCD is a high-resolution change detection dataset built from SwissTopo SWISSIMAGE 10 cm aerial imagery, covering several urban and peri-urban regions of the canton of Valais, Switzerland.It provides pairs of aerial images captured in 2017 and 2023, along with automatically generated building-change labels derived from… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/ValaisCD.neurips-spectraThe dataset from Albert's et al, downloaded from zenodo. It's on here for easier access and organisation.
rfi-simulationsCanadaFireSat-Raw
Dataset Card for CanadaFireSat 🔥🛰️
In this benchmark, we investigate the potential of deep learning with multiple modalities for high-resolution wildfire forecasting. Leveraging different data settings across two types of model architectures: CNN-based and ViT-based.
📝 Published paper from ISPRS (ArXiv Version)
💿 Dataset repository on GitHub
🤖 Model repository on GitHub & Weights on Hugging Face
🟰 Another "Clean" version of the data with PARQUET files can be found at… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/CanadaFireSat-Raw.fully-open-meditron
Fully Open Meditron Corpus
👋 Join our LiGHT community.
📖 Check out the MeditronFO blog and MeditronFO preprint.
🔜 If you are a clinician join the MOOVE initiative here.
[Hugging Face]
[Preprint]
[GitHub]
[Dataset]
License: Apache 2.0 | Authors: LiGHT
[!Note]
A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs.
The… See the full description on the dataset page: https://huggingface.co/datasets/EPFLiGHT/fully-open-meditron.EcoWikiRS
EcoWikiRS: Learning Ecological Representations of Satellite Images from Weak Supervision with Species Observations and Wikipedia
AuthorsValerie Zermatten · Javiera Castillo-Navarro · Pallavi Jain · Devis Tuia · Diego Marcos
Overview
The WikiRS dataset, composed of triplets of images, species list and Wikipedia sentences :
91k high-resolution aerial images (50cm, RGB bands) from the swissIMAGE product
crowd-sourced species observations from 2745 different… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/EcoWikiRS.zip2zip-1BLF-Bokeh-EPFL
LF-Bokeh-EPFL
Pares all-in-focus / bokeh para sintese de bokeh e refoco.
Camera: Lytro Illum, grade angular [15, 15]
Cenas: 12 · Alvos: 188
Estrutura
images/<amostra>/aif.png entrada all-in-focus
<alvo>.png alvo com bokeh
disparity.npy disparidade estimada (float16)
meta.csv uma linha por alvo: K, CoC, nitidez, luminancia
bokeh188.csv split
refocus400.csv split
dataset.json configuracao do build… See the full description on the dataset page: https://huggingface.co/datasets/AKCITPixel3/LF-Bokeh-EPFL.K600-MM
K600-MM
K600-MM is a multimodal video dataset used to pretrain A2A-Video, an any-to-any multimodal model for the video domain. It's built on top of Kinetics-600 (RGB video + class-category annotations) with 10 additional modalities obtained via pseudo-labeling. Refer to A2A-Video's README_DATA.md for additional details on dataset construction.
In total, there are ~392K training and ~30K validation/test video clips with 12 aligned modalities per clip.
The provided train/test data… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-VILAB/K600-MM.epfl-smart-kitchen-av1
EPFL-Smart-Kitchen AV1 SimpleCV Mirror
This is an AV1-transcoded SimpleCV-compatible mirror of the EPFL-Smart-Kitchen-30 dataset.
Original data:
Collected videos/data: https://zenodo.org/records/15535461
Poses/annotations: https://zenodo.org/records/15551913
GitHub: https://github.com/amathislab/EPFL-Smart-Kitchen
Contents
RGB and HoloLens videos transcoded to AV1 MP4
Depth videos transcoded to AV1 MP4
Metadata, timestamps, IMUs, poses, and annotations preserved… See the full description on the dataset page: https://huggingface.co/datasets/pablovela5620/epfl-smart-kitchen-av1.zip2zip-1B-no-split
HuggingFaceFW/fineweb-edu (20%) (common knowledge)
devngho/the-stack-llm-annotations-v2 (25%) (code)
AI-MO/NuminaMath-1.5 (20%) (math)
HuggingFaceH4/ultrachat_200k (20%) (chat)
HuggingFaceFW/fineweb-2 (15%) (multilingual: [cmn_Hani, deu_Latn, jpn_Jpan, spa_Latn, fra_Latn, ita_Latn, por_Latn, nld_Latn, arb_Arab])
llaza-20B
Llaza Mixture 20B
This dataset is a 20B-token pretraining subset built for zip2zip language-model pretraining.
It is derived from the full Llaza mixture, which is byte-balanced across four top-level domains:
Domain
Source
Target byte ratio
General
HuggingFaceFW/fineweb-edu, sample-100BT
50%
Code
bigcode/the-stack-dedup
20%
Math
HuggingFaceTB/finemath, finemath-3plus
10%
Multilingual
epfml/FineWeb2-HQ, 20 language subsets
20%
The subset was created from remixed… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/llaza-20B.Social-LLM-NetworksA variety of scenarios in which large language models, connected via a communication network, exchange opinions with one another.
Specifically, we vary:
model types
debate topics
communication networks
system prompts
initial opinion distributions
Each experiment is stored in a separate json file and includes details about variables 1-5, in addition to the exchanged texts and the corresponding sentiment scores.
This dataset supports the paper:
Iris Yazici, Mert Kayaalp, Stefan Taga, Ali H.… See the full description on the dataset page: https://huggingface.co/datasets/asl-epfl/Social-LLM-Networks.HRSCD_clean
📚 HRSCD-Clean Dataset
Project page: https://manonbechaz.github.io/2Player/
📝 Description
HRSCD-Clean is a refined and higher-quality version of the original HRSCD remote-sensingchange detection dataset (Daudt et al., 2019). The dataset contains 291 bi-temporal aerialimage pairs, each at 10,000 × 10,000 px and 0.5 m spatial resolution, covering theregions of Rennes and Caen, France. Each pair is accompanied by a binary change mask and segmentation masks for both images.… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/HRSCD_clean.MNLP_M3_mcqa_dataset
Tulu 3 SFT Mixture (Sampled)
This dataset is a sampled and filtered subset of the allenai/tulu-3-sft-mixture, curated and rebalanced for structured instruction fine-tuning. The goal is to support research and model development in math reasoning, coding, knowledge recall, instruction following (IF), and conversational alignment, while explicitly excluding safety, multilingual, and certain task-specific sources.
📦 Dataset Structure
Source: Filtered from… See the full description on the dataset page: https://huggingface.co/datasets/vanek-epfl/MNLP_M3_mcqa_dataset.nmrexpllaza-200B
Llaza Mixture Full (200B)
This dataset is the full Llaza pretraining-data mixture for zip2zip language-model pretraining.
It combines general web text, code, math, and multilingual web text with byte-based top-level mixture ratios.
Domain
Source
Target byte ratio
General
HuggingFaceFW/fineweb-edu, sample-100BT
50%
Code
bigcode/the-stack-dedup
20%
Math
HuggingFaceTB/finemath, finemath-3plus
10%
Multilingual
epfml/FineWeb2-HQ, 20 language subsets
20%… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/llaza-200B.
