datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WxC-Bench
Dataset Card for WxC-Bench
WxC-Bench primary goal is to provide a standardized benchmark for evaluating the performance of AI models in Atmospheric and Earth Sciences across various tasks.
Dataset Details
WxC-Bench contains datasets for six key tasks:
Nonlocal Parameterization of Gravity Wave Momentum Flux
Prediction of Aviation Turbulence
Identifying Weather Analogs
Generation of Natural Language Weather Forecasts
Long-Term Precipitation Forecasting
Hurricane Track and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/WxC-Bench.llava-video-178k-siglip-tokens-ftov-new
LLaVA-Video-178K SigLIP Token Cache (LLaVA-OV fine-tuned vision tower)
Derived data (vision-encoder features of video frames), not a
redistribution of the source videos. Source:
lmms-lab/LLaVA-Video-178K -- its card
restricts use to academic research and education, and its annotations come
from GPT-4-class models (see the OpenAI usage policy).
Complete: 85000 clips.
Subset
Folders: 0_30_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_academic_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new.Landslide4sense
Landslide4Sense
Dataset Description
This dataset is originally introduced in GitHub repo Landslide4Sense-2022.
The Landslide4Sense dataset has three splits, training/validation/test, consisting of 3799, 245, and 800 image patches, respectively. Each image patch is a composite of 14 bands that include:
Multispectral data from Sentinel-2: B1, B2, B3, B4, B5, B6, B7, B8, B9, B10, B11, B12.
Slope data from ALOS PALSAR: B13.
Digital elevation model (DEM) from ALOS… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/Landslide4sense.hls_burn_scarsThis dataset contains Harmonized Landsat and Sentinel-2 imagery of burn scars and the associated masks for the years 2018-2021 over the contiguous United States. There are 804 512x512 scenes. Its primary purpose is for training geospatial machine learning models.nasdaq_2013_2023data for the paper arxiv.org/abs/2502.07393
NASA-Power-Daily-Weather
NASA Power Weather Data over North, Central, and South America from 1984 to 2022
This dataset contains daily solar and meteorological data downloaded from the NASA Power API
Overview
The dataset includes solar and meteorological variables collected from January 1st, 1984, to December 31st, 2022.
We downloaded 28 variables directly and estimated an additional 3 from the collected data. The data spans a 5 x 8 grid covering
the United States, Central America, and South… See the full description on the dataset page: https://huggingface.co/datasets/notadib/NASA-Power-Daily-Weather.core-sdo
ML-Ready Multi-Modal Image Dataset from SDO
Overview
This dataset provides machine learning (ML)-ready solar data curated from NASA’s Solar Dynamics Observatory (SDO), covering observations from May 13, 2010, to Dec 31, 2024. It includes Level-1.5 processed data from: Atmospheric Imaging Assembly (AIA)
and Helioseismic and Magnetic Imager (HMI).
The dataset is designed to facilitate large-scale learning applications in heliophysics, such as space weather forecasting… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/core-sdo.clevrer-siglip-tokens-ftov
CLEVRER SigLIP Token Cache (LLaVA-OV fine-tuned vision tower)
Derived data (vision-encoder features of video frames), not a redistribution
of the source videos. Source: CLEVRER --
"CLEVRER: CoLlision Events for Video REpresentation and Reasoning" (Yi et al.,
ICLR 2020) -- official release from MIT CSAIL under CC0.
Complete: 11000 videos encoded (10000 train, 1000 validation), 0 failed (see manifest.json).
Subset
train: video_00000 ... video_09999 (10000 of the… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/clevrer-siglip-tokens-ftov.BioMassters
BioMassters: A Benchmark Dataset for Forest Biomass Estimation using Multi-modal Satellite Time-series https://nascetti-a.github.io/BioMasster/
The objective of this repository is to provide a deep learning ready dataset to predict yearly Above Ground Biomass (AGB) for Finnish forests using multi-temporal satellite imagery from
the European Space Agency and European Commission's joint Sentinel-1 and Sentinel-2 satellite missions, designed to collect a rich array of Earth observation… See the full description on the dataset page: https://huggingface.co/datasets/nascetti-a/BioMassters.Surya-1.0_validation_data
Validation data for Surya 1.0
This dataset comprises imagery from NASA's Solar Dynamics Observatory (SDO). The data can and should be used to validate a local installation of the Surya Foundation Model for Heliophysics. The data is compressed; you should use the hdf5plugin to read it directly.
aurora_rollout_betanasle-mana-clean-chunked-30s
Nasl-e-Mana Clean Speech Corpus — Sentence-Safe 30s Chunks
Training-oriented WAV chunks derived from the public Nasl-e-Mana magazine audio corpus. Chunks target approximately 30 seconds and are cut at detected acoustic pauses; the labeled configuration additionally assigns only complete source-text sentences to each chunk.
Configuration
Rows
Columns
Meaning
labeled (train/)
9,886
audio, label
Sentence-grouped text/audio pairs from duration-compatible source-text… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s.multi-temporal-crop-classification
Dataset Card for Multi-Temporal Crop Classification
Dataset Summary
This dataset contains temporal Harmonized Landsat-Sentinel imagery of diverse land cover and crop type classes across the Contiguous United States for the year 2022. The target labels are derived from USDA's Crop Data Layer (CDL). It's primary purpose is for training segmentation geospatial machine learning models.
Dataset Structure
TIFF Files
Each tiff file covers a… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/multi-temporal-crop-classification.nasle-mana-clean-chunked-30s-avasanj
Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks
Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here.
Splits
Split
Rows
Audio
Columns
labeled
4,981
41.41 hours
audio, label
to_transcribe
11,127
92.72 hours
audio
The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.Sombench-Ice-Prospectivity-Regression
SomBench Benchmark: Polar Ice Prospectivity Regression
Science theme: Polar volatiles
Task: Regression
Dataset Summary
A polar, multi-layer benchmark for predicting near-surface water-ice
prospectivity within ~10° latitude of each pole at 240 m/pixel. Following
the ice-prospectivity workflow of Coyan et al. (2025), the dataset includes a
group of physically motivated evidential layers (thermophysical,
illumination, and terrain) alongside a continuous prospectivity… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-Ice-Prospectivity-Regression.dol-visas-database
DOL Visas Database (H-1B LCA + PERM)
Every H-1B/H-1B1/E-3 Labor Condition Application and PERM permanent labor
certification application disclosed by the DOL Office of Foreign Labor
Certification, FY2015 to present, as a single queryable DuckDB database.
8,812,639 rows across 2 tables.
Table
Description
Row Count
Column Count
Date Range
lca
H-1B/H-1B1/E-3 Labor Condition Applications, one row per application per disclosure file, FY2015-present
7,479,697
110
FY2015 to… See the full description on the dataset page: https://huggingface.co/datasets/Nason/dol-visas-database.mls-en-6kh-nast100nasa-exoplanets
NASA Exoplanet Archive
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
Confirmed exoplanets with orbital, stellar, and discovery parameters from the NASA Exoplanet Archive.
The NASA Exoplanet Archive is the authoritative database of confirmed exoplanets, maintained by Caltech/IPAC under contract with NASA. Each entry represents a confirmed planet with its best-available physical and orbital parameters, host star… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/nasa-exoplanets.FNSPID_nasdaq@misc{dong2024fnspid,
title={FNSPID: A Comprehensive Financial News Dataset in Time Series},
author={Zihan Dong and Xinyu Fan and Zhiyuan Peng},
year={2024},
eprint={2402.06698},
archivePrefix={arXiv},
primaryClass={q-fin.ST}
}
pashto-emoji-dataset
Pashto Emoji Dataset
This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text.
The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label.
Dataset Structure
The dataset is provided in the following format:
text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.Sombench-pretraining-data
SomBench Pre-training Corpus: Multimodal Lunar Tiles
Dataset Summary
This includes a small sample from SomBench: a corpus of co-registered, multimodal lunar image tiles built for
large-scale self-supervised (foundation-model) pre-training. It contains a subset of modalities from the
low-resolution (WAC-anchored) and high-resolution (NAC-anchored) tracks specifically used in pretraining.
Tiles are anchored to individual LROC Experiment Data Record (EDR) image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-pretraining-data.scdb-database
Supreme Court Database (SCDB)
Every U.S. Supreme Court decision 1791-present from Washington University's Supreme Court
Database: the modern (1946-, release 2026_01) and Legacy (1791-1945) databases at the case
and justice level, all four units of analysis, plus the complete online code book as a
codes table so every integer code decodes in SQL. v_cases is pre-decoded.
677,452 rows across 9 tables.
Table
Description
Row Count
Column Count
justice_votes
One row per… See the full description on the dataset page: https://huggingface.co/datasets/Nason/scdb-database.hurricane
Data Format Description for Hurricane Evaluation on Prithvi WxC
Overview
To evaluate the performance of Prithvi WxC on hurricanes, the surface and pressure data from the MERRA-2 dataset, comprising 160 variables used in training, is required. The complete evaluation dataset includes 75 different initial conditions for hurricanes that formed in the Atlantic Ocean between 2017 and 2023.
The scientific objective is to assess the zero-shot performance of Prithvi WxC in… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/hurricane.Surya-bench-solarwind
Solar Wind Forecasting Dataset
Dataset Summary
This dataset provides hourly solar wind plasma and interplanetary magnetic field (IMF) parameters at L1, derived from NASA’s OMNI dataset. The primary forecasting target is the solar wind speed (V), while additional parameters are included for completeness:
Solar wind speed (V)
IMF Bx (GSE)
IMF By (GSM)
IMF Bz (GSM)
Proton number density (N)
The dataset is structured for machine learning experiments, particularly… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Surya-bench-solarwind.eoir-database
EOIR Immigration Court Database
A clean, queryable DuckDB database built from the EOIR FOIA data dump -- the most comprehensive public dataset on U.S. immigration court proceedings.
173,278,686 rows across 97 tables covering every immigration court case since the 1970s.
Built with eoir-database.
Quick Start
DuckDB CLI
INSTALL httpfs;
LOAD httpfs;
ATTACH 'https://huggingface.co/datasets/Nason/eoir-database/resolve/main/eoir.duckdb' AS eoir… See the full description on the dataset page: https://huggingface.co/datasets/Nason/eoir-database.nasdaq_dataSombench-NAC-Crater-Detection
SOMBench Benchmark: Hand-Labeled Crater Detection in NAC data (nac_craters_dataset)
Science theme: Impact cratering
Task: Single-class object detection (COCO bounding boxes)
Dataset Summary
A crater-detection benchmark built from expert hand-labeled crater outlines
on LROC NAC orthophotos across six lunar NAC PHO sites (Apollo 15 SIVB impact,
Apollo 17, Highland Photom, King Ejecta, March 17 Impact, Reiner Gamma). Subject matter expert (SME) annotators… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-NAC-Crater-Detection.FNSPID_nasdaq_sorted@misc{dong2024fnspid,
title={FNSPID: A Comprehensive Financial News Dataset in Time Series},
author={Zihan Dong and Xinyu Fan and Zhiyuan Peng},
year={2024},
eprint={2402.06698},
archivePrefix={arXiv},
primaryClass={q-fin.ST}
}
openpayments-database
Open Payments Database (CMS Sunshine Act)
Every disclosed industry payment to physicians, non-physician practitioners,
and teaching hospitals, program years 2013-2025, as a single queryable
DuckDB database. 172,450,995 rows.
Table
Description
Rows
Cols
Coverage
general_payments
General (non-research) payments and transfers of value to covered recipients, PY2013-2025
148,797,140
107
PY2013 to PY2025
research_payments
Research payments (wide form, up to 5 principal… See the full description on the dataset page: https://huggingface.co/datasets/Nason/openpayments-database.hls_merra2_gppFlux
Dataset Summary:
This dataset consists of Harmonized Landsat and Sentinel-2 multispectral reflectance imagery and MERRA-2 observations centered around eddy covariance flux towers and the corresponding Gross Primary Productivity (GPP) data at the towers. Its purpose is to serve as a finetuning dataset for geospatial foundation models for the task of regressing GPP flux observations from HLS and MERRA-2 data.
Dataset Structure:
The dataset consists of:
(1) HLS 6-band Tiff… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/hls_merra2_gppFlux.
