datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fava-flagged-demo
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/abhika-m/fava-flagged-demo.fav_db_test_7FaVOS_examplesFAVOR
A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
🔥 News
2025.09.18 🎉 FAVOR-Bench has been accepted by NeurIPS 2025 Datasets and Benchmarks Track!
2025.03.19 🌟 We released Favor-Bench, a new benchmark for fine-grained video motion understanding that spans both ego-centric and third-person perspectives with comprehensive evaluation including both close-ended QA tasks and open-ended descriptive tasks!
Introduction
Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/FAVOR-Bench/FAVOR.fav_db_test_1fav_db_test_9fav_db_test_19hku-viirs-nighttime-lights-global
🌍 Global 500m-Resolution Monthly VIIRS Nighttime Lights (1992–2024)
Dataset Description
This repository is a cloud-optimized mirror of the world's first global long-term monthly VIIRS-like nighttime light (NTL) dataset. It features a unified spatial resolution of ~500 meters (15 arc-seconds) and includes ocean masking. It contains images in GeoTIFF format organized temporally, making it ideal for studies on socioeconomic development, urbanization processes… See the full description on the dataset page: https://huggingface.co/datasets/faviolc/hku-viirs-nighttime-lights-global.kaggle-api-testsome personal dev files
fav_db_test_4fava-data
FAVA Datasets
FAVA datasets include: annotation data and training data.
Dataset Details
Annotation Data
The annotation dataset includes 460 annotated passages identifying and editing errors using our hallucination taxonomy. This dataset was used for the fine-grained error detection task, using the annotated passages as the gold passages.
Training Data
The training data includes 35k training instances of erroneous input and corrected output pairs… See the full description on the dataset page: https://huggingface.co/datasets/fava-uw/fava-data.fav_db_test_11BlackMarbleLeivaProject
Black Marble Project: Daily VIIRS Nighttime Lights for selected countries
The Black Marble Project: Daily VIIRS Nighttime Lights for selected countries is an open dataset repository for daily nighttime lights panels built from NASA Black Marble VIIRS VNP46A2 products and aggregated to subnational administrative units.
This repository is the public data layer of the project. It stores figures, metadata, final Parquet outputs, diagnostics, and documentation for country-level daily… See the full description on the dataset page: https://huggingface.co/datasets/faviolc/BlackMarbleLeivaProject.fav_db_test_17monthlypm25
🌍 Global 0.01°-Resolution Monthly PM₂.₅ (1998–2024) — SatPM V6.GL.03
Dataset Description
This repository is a Hugging Face mirror of the SatPM V6.GL.03 high-resolution scientific dataset. It provides global monthly estimates of ground-level fine particulate matter concentration, PM₂.₅, at a fine spatial resolution of 0.01° × 0.01°.
The estimates were generated by combining Aerosol Optical Depth (AOD) retrievals from multiple satellite-based instruments, including… See the full description on the dataset page: https://huggingface.co/datasets/faviolc/monthlypm25.fav_db_test_12fav_db_test_16fav_db_test_15cdataset
LeRobot aims to provide models, datasets, and tools for real-world robotics in PyTorch. The goal is to lower the barrier to entry so that everyone can contribute to and benefit from shared datasets and pretrained models.
🤗 A hardware-agnostic, Python-native interface that standardizes control across diverse platforms, from low-cost arms (SO-100) to humanoids.
🤗 A standardized, scalable LeRobotDataset format (Parquet + MP4 or images) hosted on the Hugging Face Hub, enabling… See the full description on the dataset page: https://huggingface.co/datasets/FAVL/cdataset.wdfav_db_test_8favoritafav_db_test_1310k-human-favorite-number
What Is Your Favorite Number?
We asked 11,631 people in 124 countries one question:
What is your favorite number?
It took about 30 minutes to collect on the Rapidata network.
If you enjoy this dataset, consider liking it. We will keep asking humanity fun questions.
What people said
7 is the world's favorite number. 18% of all answers, more than twice the runner-up (5). It ranks first in every one of the top ten countries and in every age group. Japan is the… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/10k-human-favorite-number.ansm-favismeSmart-shower-favorites-dataset
Smart Spa User Favorites Configuration Dataset
Dataset Description
This dataset contains 200 anonymized user favorite configuration files from a smart spa system, capturing diverse user preferences for water temperature, lighting, music, steam settings, and outlet configurations. Each file represents a unique user's personalized spa experience settings, making this dataset valuable for recommendation systems, personalization algorithms, and user preference modeling… See the full description on the dataset page: https://huggingface.co/datasets/Zachzzz33/Smart-shower-favorites-dataset.favorita_stores
favorita_stores (TsFile format)
This repository contains time-series forecasting data stored in Apache TsFile format.
Summary
FEV subset: favorita_stores
Unified source collection: autogluon/fev_datasets
Original source: https://www.kaggle.com/competitions/store-sales-time-series-forecasting
Paper / citation: [8]
Series: 1,579
Modalities: Time-series
TsFile rows (flattened observations): 12,054,086
Frequencies: 1D, 1M, 1W
TsFile files: 5
Time precision:… See the full description on the dataset page: https://huggingface.co/datasets/THULab/favorita_stores.favorita_transactions
favorita_transactions (TsFile format)
This repository contains time-series forecasting data stored in Apache TsFile format.
Summary
FEV subset: favorita_transactions
Unified source collection: autogluon/fev_datasets
Original source: https://www.kaggle.com/competitions/store-sales-time-series-forecasting
Paper / citation: [8]
Series: 51
Modalities: Time-series
TsFile rows (flattened observations): 288,252
Frequencies: 1D, 1M, 1W
TsFile files: 3
Time precision:… See the full description on the dataset page: https://huggingface.co/datasets/THULab/favorita_transactions.faviqfava_bean_stomata_imprint
Fava Bean Stomata Imprint
This dataset provides high-resolution RGB images of stomata imprints from faba bean leaves, captured in a field environment at Taastrup campus, Denmark. Images were acquired using a fixed platform equipped with a Leica DM750 light microscope and ICC50 HD digital microscope camera during the 2021-2022 growing season. The dataset contains 2,064 images with no classification, segmentation, or bounding-box annotations.
This dataset is indexed on… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/fava_bean_stomata_imprint.
