datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PocketQubeMODUS-15Modality
MODUS — 15-Modality Aligned Dataset
MODUS is a large-scale, pixel-aligned 15-modality dataset for any-to-any
multimodal training. Every sample aligns 15 modalities covering appearance,
geometry, structure, segmentation, detection, text, and learned features.
Paper: https://huggingface.co/papers/2607.25948
Code: https://github.com/EPFL-VILAB/Modus
Modalities
Group
Modalities
Appearance
rgb, caption
Geometry
depth, normal
Structure
canny, sam_edge… See the full description on the dataset page: https://huggingface.co/datasets/epfl-vilab-modus/MODUS-15Modality.SwissCubecoralscapes
Coralscapes Dataset
The Coralscapes dataset is the first general-purpose dense semantic segmentation dataset for coral reefs. Similar in scope and with the same structure as the widely used Cityscapes dataset for urban scene understanding, Coralscapes allows for the benchmarking of semantic segmentation models in a new challenging domain.
Dataset Structure
The Coralscapes dataset spans 2075 images at 1024×2048px resolution… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/coralscapes.SwissView
Dataset Card for SwissView Dataset
Project Page
https://limirs.github.io/GeoExplorer/
GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
Dataset Summary
This dataset consists of two subsets:
SwissViewMonuments: which includes 15 images of atypical or distinctive scenes, such as unusual buildings and landscapes, with corresponding ground level images.
SwissView100, which comprises 100 images randomly selected from across the Swiss territory… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/SwissView.svi-benchmark
Stable Video Infinity (SVI) Benchmark Dataset
This benchmark dataset is introduced in the paper:
Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
by Wuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao, Alexandre Alahi (2025).
Project page: https://stable-video-infinity.github.io/homepage/
Code: https://github.com/vita-epfl/Stable-Video-Infinity
Abstract
We propose Stable Video Infinity (SVI) that is able to generate infinite-length videos with… See the full description on the dataset page: https://huggingface.co/datasets/epfl-vita/svi-benchmark.ValaisCD
ValaisCD Dataset
High-Resolution Aerial Change Detection (Switzerland, 2017–2023)
Project page: https://manonbechaz.github.io/2Player/
🗺️ Overview
ValaisCD is a high-resolution change detection dataset built from SwissTopo SWISSIMAGE 10 cm aerial imagery, covering several urban and peri-urban regions of the canton of Valais, Switzerland.It provides pairs of aerial images captured in 2017 and 2023, along with automatically generated building-change labels derived from… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/ValaisCD.EcoWikiRS
EcoWikiRS: Learning Ecological Representations of Satellite Images from Weak Supervision with Species Observations and Wikipedia
AuthorsValerie Zermatten · Javiera Castillo-Navarro · Pallavi Jain · Devis Tuia · Diego Marcos
Overview
The WikiRS dataset, composed of triplets of images, species list and Wikipedia sentences :
91k high-resolution aerial images (50cm, RGB bands) from the swissIMAGE product
crowd-sourced species observations from 2745 different… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/EcoWikiRS.LF-Bokeh-EPFL
LF-Bokeh-EPFL
Pares all-in-focus / bokeh para sintese de bokeh e refoco.
Camera: Lytro Illum, grade angular [15, 15]
Cenas: 12 · Alvos: 188
Estrutura
images/<amostra>/aif.png entrada all-in-focus
<alvo>.png alvo com bokeh
disparity.npy disparidade estimada (float16)
meta.csv uma linha por alvo: K, CoC, nitidez, luminancia
bokeh188.csv split
refocus400.csv split
dataset.json configuracao do build… See the full description on the dataset page: https://huggingface.co/datasets/AKCITPixel3/LF-Bokeh-EPFL.CAVE
Dataset Card for CAVE: Commonsense Anomalies in Visual Environments
🏠 Project Page📄 Paper (EMNLP 2025)💻 Code
Dataset Details
Dataset Description
CAVE is the first benchmark of real-world visual anomalies for evaluating Vision-Language Models (VLMs). It is curated from images captured in real-life settings (photographs and screenshots taken by individuals), sourced from Reddit.
The benchmark is grounded in cognitive science literature on how humans detect and… See the full description on the dataset page: https://huggingface.co/datasets/epfl-nlp/CAVE.epfl-vi
