datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CanadaFireSat
Dataset Card for CanadaFireSat 🔥🛰️
In this benchmark, we investigate the potential of deep learning with multiple modalities for high-resolution wildfire forecasting. Leveraging different data settings across two types of model architectures: CNN-based and ViT-based.
📝 Published paper from ISPRS (ArXiv Version)
💿 Dataset repository on GitHub
🤖 Model repository on GitHub & Weights on Hugging Face
🟰 Another "Raw" version of the data with NPY files organized in different… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/CanadaFireSat.nmrexpepfl-enterprise-osai-adoption-research-data
EPFL Enterprise Open-Source AI Adoption Research Dataset
Dataset Summary
This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption.
Dataset Structure
This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.MassSpecGym
MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from MS/MS spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems.
Please refer to the MassSpecGym GitHub page and the paper for details.
epfl-computer-science-mcqamultimodal-brain-scaling
Multimodal Scaling Laws for Task & Data-Optimized Models of Visual Cortex
This repository hosts the result tables and accompanying metadata released
with our ICML 2026 paper
Multimodal Scaling Laws for Task & Data-Optimized Models of Visual Cortex.
Paper summary
Task-optimized deep networks are the leading in-silico models of sensory
cortex, but progress is fragmented across datasets, modalities, and
evaluation protocols, making it hard to identify which… See the full description on the dataset page: https://huggingface.co/datasets/epfl-neuroai/multimodal-brain-scaling.epfl_english_guidelines_translated_to_dutch_with_gpt4omini
Dataset Card for Epfl English Guidelines Translated To Dutch With Gpt4Omini
This dataset was created by the EPFL, and can found in it original form here
The source language: English
The original data source: Original Data Source
Data description
Translation of the English medical guidelines that are part of the Meditron corpus, using the LLM GPT 4o mini
Acknowledgement
This is part of the DT4H project with attribution [Cite the paper].
Doi and… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/epfl_english_guidelines_translated_to_dutch_with_gpt4omini.zip2zip-wikitext-repeat-phi35
Zip2Zip Repeated WikiText Stress Tests (Phi-3.5)
This repository contains two controlled evaluation corpora for studying
merge-size transfer in Zip2Zip models. They are derived from the document-level
WikiText-2 raw test split and built specifically with the
microsoft/Phi-3.5-mini-instruct tokenizer.
Configurations
Config
Repetitions per source block
Rows
Repeated base tokens
SHA-256 of test.jsonl
repeat4
4
1,329
1,297,250… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/zip2zip-wikitext-repeat-phi35.qm9
QM9 for the structure project
This dataset was cloned from https://huggingface.co/datasets/yairschiff/qm9 .
It contains additional properties computed via rdkit on top of the base QM9.
Info
QM9 dataset from Ruddigkeit et al., 2012;
Ramakrishnan et al., 2014.
Original data downloaded from: http://quantum-machine.org/datasets.
Additional annotations (QED, logP, SA score, NP score, bond and ring counts) added using rdkit library.
Quick start usage:
from… See the full description on the dataset page: https://huggingface.co/datasets/structure-epflai/qm9.epfl_guidelines_dutch_marianmt
Dataset Card for Epfl English Guidelines Translated To Dutch With MariaNMT
This dataset was created by the EPFL, and can found in it original form here
The source language: English
The original data source: Original Data Source
The MariaNMT model used can be found: here
Data description
Translation of the English medical guidelines that are part of the Meditron corpus, using the LLM GPT 4o mini
Acknowledgement
This is part of the DT4H project with… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/epfl_guidelines_dutch_marianmt.m1_epfl_data
