datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sombench-Ice-Prospectivity-Regression
SomBench Benchmark: Polar Ice Prospectivity Regression
Science theme: Polar volatiles
Task: Regression
Dataset Summary
A polar, multi-layer benchmark for predicting near-surface water-ice
prospectivity within ~10° latitude of each pole at 240 m/pixel. Following
the ice-prospectivity workflow of Coyan et al. (2025), the dataset includes a
group of physically motivated evidential layers (thermophysical,
illumination, and terrain) alongside a continuous prospectivity… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-Ice-Prospectivity-Regression.multi-temporal-crop-classification
Dataset Card for Multi-Temporal Crop Classification
Dataset Summary
This dataset contains temporal Harmonized Landsat-Sentinel imagery of diverse land cover and crop type classes across the Contiguous United States for the year 2022. The target labels are derived from USDA's Crop Data Layer (CDL). It's primary purpose is for training segmentation geospatial machine learning models.
Dataset Structure
TIFF Files
Each tiff file covers a… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/multi-temporal-crop-classification.nasle-mana-clean-chunked-30s
Nasl-e-Mana Clean Speech Corpus — Sentence-Safe 30s Chunks
Training-oriented WAV chunks derived from the public Nasl-e-Mana magazine audio corpus. Chunks target approximately 30 seconds and are cut at detected acoustic pauses; the labeled configuration additionally assigns only complete source-text sentences to each chunk.
Configuration
Rows
Columns
Meaning
labeled (train/)
9,886
audio, label
Sentence-grouped text/audio pairs from duration-compatible source-text… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s.nasle-mana-clean-chunked-30s-avasanj
Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks
Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here.
Splits
Split
Rows
Audio
Columns
labeled
4,981
41.41 hours
audio, label
to_transcribe
11,127
92.72 hours
audio
The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.mls-en-6kh-nast100nasa-exoplanets
NASA Exoplanet Archive
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
Confirmed exoplanets with orbital, stellar, and discovery parameters from the NASA Exoplanet Archive.
The NASA Exoplanet Archive is the authoritative database of confirmed exoplanets, maintained by Caltech/IPAC under contract with NASA. Each entry represents a confirmed planet with its best-available physical and orbital parameters, host star… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/nasa-exoplanets.FNSPID_nasdaq@misc{dong2024fnspid,
title={FNSPID: A Comprehensive Financial News Dataset in Time Series},
author={Zihan Dong and Xinyu Fan and Zhiyuan Peng},
year={2024},
eprint={2402.06698},
archivePrefix={arXiv},
primaryClass={q-fin.ST}
}
Sombench-pretraining-data
SomBench Pre-training Corpus: Multimodal Lunar Tiles
Dataset Summary
This includes a small sample from SomBench: a corpus of co-registered, multimodal lunar image tiles built for
large-scale self-supervised (foundation-model) pre-training. It contains a subset of modalities from the
low-resolution (WAC-anchored) and high-resolution (NAC-anchored) tracks specifically used in pretraining.
Tiles are anchored to individual LROC Experiment Data Record (EDR) image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-pretraining-data.pashto-emoji-dataset
Pashto Emoji Dataset
This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text.
The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label.
Dataset Structure
The dataset is provided in the following format:
text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.Sombench-WAC-Crater-Detection
SomBench Benchmark: Robbins Crater Detection, WAC
Science theme: Impact processes
Task: Object detection
Dataset Summary
An impact-crater object-detection benchmark built from the
Robbins (2019) global lunar crater
catalog, a manually compiled, near-complete census of
lunar impact craters (≥ ~1–2 km). Catalog crater centers and diameters are
converted to bounding boxes and packaged over LROC WAC visible tiles drawn
from the pre-training corpus test split, in COCO… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-WAC-Crater-Detection.Surya-bench-solarwind
Solar Wind Forecasting Dataset
Dataset Summary
This dataset provides hourly solar wind plasma and interplanetary magnetic field (IMF) parameters at L1, derived from NASA’s OMNI dataset. The primary forecasting target is the solar wind speed (V), while additional parameters are included for completeness:
Solar wind speed (V)
IMF Bx (GSE)
IMF By (GSM)
IMF Bz (GSM)
Proton number density (N)
The dataset is structured for machine learning experiments, particularly… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Surya-bench-solarwind.nasdaq_dataFNSPID_nasdaq_sorted@misc{dong2024fnspid,
title={FNSPID: A Comprehensive Financial News Dataset in Time Series},
author={Zihan Dong and Xinyu Fan and Zhiyuan Peng},
year={2024},
eprint={2402.06698},
archivePrefix={arXiv},
primaryClass={q-fin.ST}
}
surya-bench-ar-segmentation
A Dataset of Binary Maps of Active Regions with Polarity Inversion Lines
Dataset Summary
This dataset provides hourly binary segmentation maps (4096×4096 resolution) derived from Solar Dynamics Observatory (SDO) / Helioseismic and Magnetic Imager (HMI) line-of-sight magnetograms. The maps highlight regions containing Active Regions (ARs) and Polarity Inversion Lines (PILs). The dataset spans observations from May 13, 2010 to December 31, 2024 and is intended for image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/surya-bench-ar-segmentation.DarijaDz
DarijaDZ
DarijaDZ is a large-scale corpus of user-generated text collected from Algerian YouTube and TikTok channels. The corpus contains approximately 22.5 million comment documents and 259.22 million word-level tokens, with content written primarily in Algerian Darija script alongside Latin/Arabizi writing and mixed-script content.
Dataset Description
Motivation
Algerian Darija is an under-resourced language variety with comparatively limited… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDz.NASA_Nearest_Earth_Objects_1910-2024CONTEXT:
There are many dangerous bodies in space, one of them is N.E.O. - "Nearest Earth Objects". Some such bodies really pose a danger to the planet Earth, NASA classifies them as "is_hazardous". This dataset contains ALL NASA observations of similar objects from 1910 to 2024!!!
There are 338,199 records of N.E.O. in the Dataset!
Try to predict "is_hazardous" as accurately as possible! (otherwise we will not be ready for an asteroid attack)
SOURCES:
NASA Open API: https://api.nasa.gov/… See the full description on the dataset page: https://huggingface.co/datasets/IvanSher/NASA_Nearest_Earth_Objects_1910-2024.nasa
NASA Image Library (Recaptioned)
A public-domain subset of the NASA Image and Video Library, filtered and captioned with a Qwen vision-language model. Prepared by the Swiss AI Initiative vision team.
The set contains 36,997 images, all public domain, covering spaceflight, astronomy, planetary science, Earth observation and aeronautics, from NASA centers including JPL, KSC, MSFC, JSC and GSFC.
How it was made
Images were sourced directly from NASA's public API… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/nasa.Sombench-IMP-Segmentation
SomBench Benchmark: Irregular Mare Patch (IMP) Segmentation
Science theme: Volcanic history
Task: Binary semantic segmentation
Dataset Summary
A binary semantic-segmentation benchmark for irregular mare patches
(IMPs): rare, morphologically subtle features interpreted as unusually young
volcanic landforms. Each sample is an LROC NAC image tile paired with a
binary IMP mask (IMP vs. background). The set is derived from published IMP
polygon annotations, framed as a… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-IMP-Segmentation.icrm-hitek-full-db-mixed
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/Nasskeke/icrm-hitek-full-db-mixed.coco2TawfikPashto-Free-Hand-Reasoning-Dataset
Pashto Free-Hand Reasoning SFT Dataset 🧠♻️
This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses.
🔄 The 3R Approach (Recycle, Reuse, Reason)
Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy:
Recycle: Taking older, simple, or raw legacy Pashto questions.
Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.nasdaq-external-data-processedCOCO_AI
VISUAL COUNTER TURING TEST - COCO DATASET
The Visual Counter Turing Test (VCT²) dataset is introduced in the paper“Visual Counter Turing Test (VCT²): Discovering the Challenges for AI-Generated Image Detection and Introducing Visual AI Index (V_AI)”,accepted at IJCNLP–AACL 2025 and available on arXiv:2411.16754.
This dataset aims to benchmark and analyze the challenges of AI-generated image detection (AGID) in the era of advanced text-to-image models.It provides a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/NasrinImp/COCO_AI.NASA-C-MAPSS-Turbofan-Engine
----------------------------------------------------------------------------------------------------------------------------------------------------
Remaining useful life prediction
|
Predictive maintenance
|
Digital twin research
|
Turbofan engine degradation
----------------------------------------------------------------------------------------------------------------------------------------------------
1. Project Introduction… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/NASA-C-MAPSS-Turbofan-Engine.nasa-neows
NASA NeoWs — Near Earth Asteroid Close Approaches
Every asteroid close approach to Earth from 2000 to today, as tracked by NASA's Center for Near Earth Object Studies (CNEOS). 70,944 close-approach events involving 20,159 unique asteroids — with size estimates, speed, miss distance, and hazard classification for each pass.
Source
API: NASA NeoWs API
Data refreshed: periodically via NASA NeoWs public API
Notes
Potentially Hazardous Asteroid (PHA):… See the full description on the dataset page: https://huggingface.co/datasets/Hari5115/nasa-neows.nasa-exoplanet
NASA Exoplanet Archive — Confirmed Planets
Every confirmed exoplanet in the NASA Exoplanet Archive — 6,316 planets across 31+ years of discovery, from the first pulsar-timing detections in 1992 to the latest TESS and JWST findings. Includes orbital parameters, planetary size/mass, stellar properties, and discovery metadata.
Source
API: NASA Exoplanet Archive TAP Service
Table: pscomppars (Planetary Systems Composite Parameters)
Data refreshed: periodically via… See the full description on the dataset page: https://huggingface.co/datasets/Hari5115/nasa-exoplanet.NASADEM_DATASET
NASADEM_DATASET/
README.md
train/metadata.csv
train.zip
details_Naseej__noon-7b_v2
Dataset Card for Evaluation run of Naseej/noon-7b
Dataset automatically created during the evaluation run of model Naseej/noon-7b.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Naseej__noon-7b_v2.NASA_IMPACT_TCB_OBJ
Transverse Cirrus Bands (TCB) Dataset
Dataset Overview
This dataset contains manually annotated satellite imagery of Transverse Cirrus Bands (TCBs), a type of cloud formation often associated with atmospheric turbulence. The dataset is formatted for object detection tasks using the YOLO and COCO annotation formats, making it suitable for training deep learning models for automated TCB detection.
Data Collection
Source: NASA-IMPACT Data Share
Satellite Sensors:… See the full description on the dataset page: https://huggingface.co/datasets/viknesh1211/NASA_IMPACT_TCB_OBJ.nasle-mana-clean
Nasl-e-Mana Speech Corpus
Clean, playable Persian speech audio collected from the public Nasl-e-Mana magazine WordPress site. The export contains two intentionally different collections:
Split
Rows
Columns
Meaning
labeled configuration (train/)
809
audio, label
Audio with recovered article text. These clips were identified as the consistent female source-text voice and are kept together.
to_transcribe configuration (to_transcribe/)
626
audio
Playable audio for which… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean.
