CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01epfl-vilab-modus /MODUS-15Modality MODUS — 15-Modality Aligned Dataset MODUS is a large-scale, pixel-aligned 15-modality dataset for any-to-any multimodal training. Every sample aligns 15 modalities covering appearance, geometry, structure, segmentation, detection, text, and learned features. Paper: https://huggingface.co/papers/2607.25948 Code: https://github.com/EPFL-VILAB/Modus Modalities Group Modalities Appearance rgb, caption Geometry depth, normal Structure canny, sam_edge… See the full description on the dataset page: https://huggingface.co/datasets/epfl-vilab-modus/MODUS-15Modality.imageimage-to-image10M<n<100M2 likes7k downloads2mo agoHugging Face02epfl-dlab /JSONSchemaBench JSONSchemaBench JSONSchemaBench is a benchmark of real-world JSON schemas designed to evaluate structured output generation for Large Language Models (LLMs). It contains approximately 10,000 JSON schemas, capturing diverse constraints and complexities. import datasets from datasets import load_dataset def main(): # Inspect the available subsets of the datasetall_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench") print("Available subsets:"… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/JSONSchemaBench.texttext-generation10K<n<100K12 likes4.1k downloads1y agoHugging Face03EPFL-ECEO /CanadaFireSat Dataset Card for CanadaFireSat 🔥🛰️ In this benchmark, we investigate the potential of deep learning with multiple modalities for high-resolution wildfire forecasting. Leveraging different data settings across two types of model architectures: CNN-based and ViT-based. 📝 Published paper from ISPRS (ArXiv Version) 💿 Dataset repository on GitHub 🤖 Model repository on GitHub & Weights on Hugging Face 🟰 Another "Raw" version of the data with NPY files organized in different… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/CanadaFireSat.tabularimage-segmentation100K<n<1M2 likes3.9k downloads21d agoHugging Face04epfl-llm /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.texttext-generation10K<n<100K158 likes3k downloads3y agoHugging Face05epfl-vita /svi-benchmark Stable Video Infinity (SVI) Benchmark Dataset This benchmark dataset is introduced in the paper: Stable Video Infinity: Infinite-Length Video Generation with Error Recycling by Wuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao, Alexandre Alahi (2025). Project page: https://stable-video-infinity.github.io/homepage/ Code: https://github.com/vita-epfl/Stable-Video-Infinity Abstract We propose Stable Video Infinity (SVI) that is able to generate infinite-length videos with… See the full description on the dataset page: https://huggingface.co/datasets/epfl-vita/svi-benchmark.imageimage-to-videon<1K7 likes766 downloads11mo agoHugging Face06structure-epflai /neurips-spectraThe dataset from Albert's et al, downloaded from zenodo. It's on here for easier access and organisation. text100K<n<1M0 likes695 downloads6mo agoHugging Face07EPFLiGHT /fully-open-meditron Fully Open Meditron Corpus 👋 Join our LiGHT community. 📖 Check out the MeditronFO blog and MeditronFO preprint. 🔜 If you are a clinician join the MOOVE initiative here. [Hugging Face] [Preprint] [GitHub] [Dataset] License: Apache 2.0 | Authors: LiGHT [!Note] A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs. The… See the full description on the dataset page: https://huggingface.co/datasets/EPFLiGHT/fully-open-meditron.textquestion-answering100K<n<1M8 likes327 downloads3mo agoHugging Face08EPFL-ECEO /EcoWikiRS EcoWikiRS: Learning Ecological Representations of Satellite Images from Weak Supervision with Species Observations and Wikipedia AuthorsValerie Zermatten · Javiera Castillo-Navarro · Pallavi Jain · Devis Tuia · Diego Marcos Overview The WikiRS dataset, composed of triplets of images, species list and Wikipedia sentences : 91k high-resolution aerial images (50cm, RGB bands) from the swissIMAGE product crowd-sourced species observations from 2745 different… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/EcoWikiRS.image10K<n<100K1 likes260 downloads9mo agoHugging Face09epfl-dlab /zip2zip-1Btext100K<n<1M0 likes256 downloads1y agoHugging Face10structure-epflai /nmrshiftdb2text10K<n<100K0 likes216 downloads7mo agoHugging Face11pablovela5620 /epfl-smart-kitchen-av1 EPFL-Smart-Kitchen AV1 SimpleCV Mirror This is an AV1-transcoded SimpleCV-compatible mirror of the EPFL-Smart-Kitchen-30 dataset. Original data: Collected videos/data: https://zenodo.org/records/15535461 Poses/annotations: https://zenodo.org/records/15551913 GitHub: https://github.com/amathislab/EPFL-Smart-Kitchen Contents RGB and HoloLens videos transcoded to AV1 MP4 Depth videos transcoded to AV1 MP4 Metadata, timestamps, IMUs, poses, and annotations preserved… See the full description on the dataset page: https://huggingface.co/datasets/pablovela5620/epfl-smart-kitchen-av1.textn<1K0 likes187 downloads4mo agoHugging Face12epfl-dlab /llaza-20B Llaza Mixture 20B This dataset is a 20B-token pretraining subset built for zip2zip language-model pretraining. It is derived from the full Llaza mixture, which is byte-balanced across four top-level domains: Domain Source Target byte ratio General HuggingFaceFW/fineweb-edu, sample-100BT 50% Code bigcode/the-stack-dedup 20% Math HuggingFaceTB/finemath, finemath-3plus 10% Multilingual epfml/FineWeb2-HQ, 20 language subsets 20% The subset was created from remixed… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/llaza-20B.texttext-generation10M<n<100M0 likes157 downloads5mo agoHugging Face13epfl-dlab /zip2zip-1B-no-split HuggingFaceFW/fineweb-edu (20%) (common knowledge) devngho/the-stack-llm-annotations-v2 (25%) (code) AI-MO/NuminaMath-1.5 (20%) (math) HuggingFaceH4/ultrachat_200k (20%) (chat) HuggingFaceFW/fineweb-2 (15%) (multilingual: [cmn_Hani, deu_Latn, jpn_Jpan, spa_Latn, fra_Latn, ita_Latn, por_Latn, nld_Latn, arb_Arab]) text100K<n<1M0 likes153 downloads2y agoHugging Face14EPFL-VILAB /K600-MM K600-MM K600-MM is a multimodal video dataset used to pretrain A2A-Video, an any-to-any multimodal model for the video domain. It's built on top of Kinetics-600 (RGB video + class-category annotations) with 10 additional modalities obtained via pseudo-labeling. Refer to A2A-Video's README_DATA.md for additional details on dataset construction. In total, there are ~392K training and ~30K validation/test video clips with 12 aligned modalities per clip. The provided train/test data… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-VILAB/K600-MM.textn<1K1 likes136 downloads24d agoHugging Face15structure-epflai /nmrexptabular1M<n<10M0 likes92 downloads5mo agoHugging Face16vanek-epfl /MNLP_M3_mcqa_dataset Tulu 3 SFT Mixture (Sampled) This dataset is a sampled and filtered subset of the allenai/tulu-3-sft-mixture, curated and rebalanced for structured instruction fine-tuning. The goal is to support research and model development in math reasoning, coding, knowledge recall, instruction following (IF), and conversational alignment, while explicitly excluding safety, multilingual, and certain task-specific sources. 📦 Dataset Structure Source: Filtered from… See the full description on the dataset page: https://huggingface.co/datasets/vanek-epfl/MNLP_M3_mcqa_dataset.text100K<n<1M0 likes91 downloads1y agoHugging Face17epfl-dlab /llaza-200B Llaza Mixture Full (200B) This dataset is the full Llaza pretraining-data mixture for zip2zip language-model pretraining. It combines general web text, code, math, and multilingual web text with byte-based top-level mixture ratios. Domain Source Target byte ratio General HuggingFaceFW/fineweb-edu, sample-100BT 50% Code bigcode/the-stack-dedup 20% Math HuggingFaceTB/finemath, finemath-3plus 10% Multilingual epfml/FineWeb2-HQ, 20 language subsets 20%… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/llaza-200B.texttext-generation100M<n<1B0 likes83 downloads5mo agoHugging Face18itseffi /epfl-enterprise-osai-adoption-research-data EPFL Enterprise Open-Source AI Adoption Research Dataset Dataset Summary This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption. Dataset Structure This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.tabulartext-classificationn<1K0 likes68 downloads1y agoHugging Face19structure-epflai /MassSpecGym MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from MS/MS spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems. Please refer to the MassSpecGym GitHub page and the paper for details. tabular100K<n<1M0 likes53 downloads5mo agoHugging Face20PJMixers /epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining. texttext-generation10K<n<100K0 likes48 downloads3y agoHugging Face21gabrieljimenez /epfl-computer-science-mcqatabularn<1K0 likes40 downloads1y agoHugging Face22Hugues898 /epfl-math-step-10ktext10K<n<100K0 likes38 downloads1y agoHugging Face23vanek-epfl /MNLP_M2_mcqa_dataset Smol-SmalTalk This is a subset of SmolTalk dataset adapted for smol models with less than 1B parameters. We used it to build SmolLM2-360M-Instruct and SmolLM2-135M-Instruct. We do SFT on this dataset and then DPO on UltraFeedback. Compared to SmolTalk: The conversations from Smol-Magpie-Ultra are shorter in this dataset We include less task specific data compared to SmolTalk (e.g no function calling and less rewriting and summarization examples) since these smaller models have… See the full description on the dataset page: https://huggingface.co/datasets/vanek-epfl/MNLP_M2_mcqa_dataset.text100K<n<1M0 likes35 downloads1y agoHugging Face24fh1628 /stack_exchange_epfl_data_2048text10K<n<100K0 likes35 downloads1y agoHugging Face25epfl-nlp /include-89text100K<n<1M0 likes35 downloads4mo agoHugging Face26vanek-epfl /tulu3-sft-mixture-sampled Tulu 3 SFT Mixture (Sampled) This dataset is a sampled and filtered subset of the allenai/tulu-3-sft-mixture, curated and rebalanced for structured instruction fine-tuning. The goal is to support research and model development in math reasoning, coding, knowledge recall, instruction following (IF), and conversational alignment, while explicitly excluding safety, multilingual, and certain task-specific sources. 📦 Dataset Structure Source: Filtered from… See the full description on the dataset page: https://huggingface.co/datasets/vanek-epfl/tulu3-sft-mixture-sampled.text100K<n<1M0 likes34 downloads1y agoHugging Face27epfl-neuroai /multimodal-brain-scaling Multimodal Scaling Laws for Task & Data-Optimized Models of Visual Cortex This repository hosts the result tables and accompanying metadata released with our ICML 2026 paper Multimodal Scaling Laws for Task & Data-Optimized Models of Visual Cortex. Paper summary Task-optimized deep networks are the leading in-silico models of sensory cortex, but progress is fragmented across datasets, modalities, and evaluation protocols, making it hard to identify which… See the full description on the dataset page: https://huggingface.co/datasets/epfl-neuroai/multimodal-brain-scaling.tabular1M<n<10M0 likes34 downloads3mo agoHugging Face28EPFL-VILAB /TST-ProcTHOR Dataset Card for TST-ProcTHOR Dataset Summary This custom TST-ProcTHOR dataset is used in research work "Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality". pretrain/ is a multimodal pretraining dataset collected using ProcTHOR environment. It contains RGB images, and 9 additional tokenized modalities. segmentation/train is the associated downstream dataset used to finetune TST pretrained models on semantic… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-VILAB/TST-ProcTHOR.text100K<n<1M0 likes32 downloads5mo agoHugging Face29berquetR /dpo_preference_dataset_EPFL_shptext10K<n<100K0 likes30 downloads2y agoHugging Face30MarceauBBB /epfl_mnlp_dpo_evaluation_datasettext1K<n<10K0 likes29 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.