datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OceanDepths
OceanDepths GeoTIFF Raster and Aligned ARGO Dataset
This dataset package contains the model-ready Ocean variables (ARGO submarine data, sea surface height, sea surface temperature and
salinity, as well as GLORYS reanalysis information for 50 depth levels. The ARGO data has been projected onto the GLORYS grid in order
to build a ML-ready dataset. The intention is that users can create tensors easily for CV-inspired ML approaches to ocean-variable
reconstruction. While… See the full description on the dataset page: https://huggingface.co/datasets/ESA-philab/OceanDepths.OceanTACO
Dataset Card: OceanTACO
Dataset Summary
This dataset is a multi-source collection of global ocean sea surface measurements, integrating numerical model reanalysis, L4 gap-filled products, L3 satellite observations, and in-situ data. The collection includes sea surface height (SSH), temperature (SST), salinity (SSS), wind speed, and other variables.
The L3 SWOT data has been processed onto a consistent regular grid through irreversible interpolation and coordinate… See the full description on the dataset page: https://huggingface.co/datasets/nilsleh/OceanTACO.planktonzilla-17M
Planktonzilla-17M Dataset
Overview
planktonzilla-17M is a large-scale, comprehensive dataset combining 17 million plankton images from all publicly available -to the best
of our knowledge- labeled plankton datasets. This unified collection enables researchers to train robust deep learning models for plankton
identification and classification across diverse imaging systems and oceanographic environments.
Each image includes a standardized taxonomic hierarchy… See the full description on the dataset page: https://huggingface.co/datasets/project-oceania/planktonzilla-17M.OceanInstruction
OceanInstruction Dataset
1. Dataset Description
OceanInstruction is a specialized instruction-tuning dataset designed for multimodal large language models (MLLMs) in the marine domain. The data has been rigorously curated, deduplicated, and standardized. It encompasses a diverse range of tasks, spanning from text-only encyclopedic QA and sonar image-based QA to RGB natural image QA (covering biological specimens and scientific diagrams).
2. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/OceanInstruction.ocean_soundscape
Dataset Card for Koh Man Marine Soundscape (Raw Audio)
Dataset Summary
The Koh Man Marine Soundscape dataset is an unannotated, raw underwater acoustic dataset collected via Passive Acoustic Monitoring (PAM) around the Man Islands, Rayong Province, Thailand. The primary objective is to study and monitor coral reef ecosystem health by comparing a relatively pristine reef area (Ao Ton Liab / MN-GOOD) with a degraded reef area (Na Ban Beach / MN-DEGRADED).
This… See the full description on the dataset page: https://huggingface.co/datasets/WasuratS/ocean_soundscape.OceanBenchmark
OceanBenchmark
1. Dataset Description
OceanBenchmark is a benchmark dataset designed to evaluate the comprehensive capabilities of marine-focused large models. It encompasses a diverse range of tasks, spanning from unimodal marine science knowledge question answering to complex multimodal visual question answering.
2. Sub-datasets
Subset Directory
Task Type
Sample Size
Description
Example
VQA
816
Combined multimodal examples with Sonar… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/OceanBenchmark.Ocean_R1_collected_visual_dataSymVAE
Geometric Logic-Form Dataset (SymParser / SymVAE / SymHPR)
Dataset for the CVPR 2026 paper Hierarchical Process Reward Models are Symbolic Vision Learners — Shan Zhang, Aotian Chen, Kai Zou, Jindong Gu, Yuan Xue, Anton van den Hengel.
🔗 Project page: vi-ocean.github.io/projects/SymVAE — SymVAE stands for Symbolic Variational Auto-Encoder.
A multimodal dataset for training and evaluating MLLMs on geometric diagram → logic-form extraction. Given a geometry image, a model must… See the full description on the dataset page: https://huggingface.co/datasets/vi-ocean/SymVAE.IndustryCorpus2_water_resources_ocean
IndustryCorpus2: Water Resources & Marine
This repository contains the IndustryCorpus2: Water Resources & Marine domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_water_resources_ocean.oceania-gov-open-data-catalog
Oceania Government Open Data — Combined Catalogue (hourly snapshot)
Combined regional catalogue of Oceania (Australia + New Zealand) public-service open
data harvested from both data.gov.au and data.govt.nz portals, including state,
territory and local-council publishers.
License declaration
License: other (see below). Records in this catalogue inherit the licence of their
source dataset. Where the source declares a standard open licence the record is tagged
with… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/oceania-gov-open-data-catalog.oceania-landuse-11class-v1
MapSpace LSG 2026: Oceania 11-Class Land Use
Dataset Details
This dataset contains POI-based land-use estimates for LandScan Global grid cells in Oceania. It was generated and is shared by the Geospatial Science and Engineering Division at Oak Ridge National Laboratory.
Curated by: Geospatial Science and Engineering Division, Oak Ridge National Laboratory
Version: v1
License: CC BY 4.0
Geographic coverage: Oceania
Spatial grid: 30 arc seconds (1/120 degree)
CRS:… See the full description on the dataset page: https://huggingface.co/datasets/MapSpaceORNL/oceania-landuse-11class-v1.Ocean_R1_visual_data_stage1This repository contains the data presented in Ocean-R1: An Open and Generalizable Large Vision-Language Model enhanced by Reinforcement Learning.
PlasticInWaterThis model was trained on images of different types of plastic placed inside a tank connected to a 1080p webcam, located in the Ocean Technology Center at the University of Washington. The goal of this project is to enhance the detection and classification of various types of plastic debris commonly found in marine environments, providing valuable tools for environmental monitoring and research.
license: MIT
Dataset Details
The dataset consists of 4,511 images, capturing… See the full description on the dataset page: https://huggingface.co/datasets/OceanCV/PlasticInWater.marine_ocean_mammal_sound
Marine Ocean Mammal Sound Dataset
Sound files on this website are free to download for personal or academic (not commercial) use.
In this database version, the audio archive includes sounds of 32 species:
Atlantic_Spotted_Dolphin
Bearded_Seal
Beluga,_White_Whale
Bottlenose_Dolphin
Bowhead_Whale
Clymene_Dolphin
Common_Dolphin
False_Killer_Whale
Fin,_Finback_Whale
Frasers_Dolphin
Grampus,_Rissos_Dolphin
Harp_Seal
Humpback_Whale
Killer_Whale
Leopard_Seal
Long-Finned_Pilot_Whale… See the full description on the dataset page: https://huggingface.co/datasets/ardavey/marine_ocean_mammal_sound.Ocean_R1_visual_data_stage2This repository contains the data presented in Ocean-R1: An Open and Generalizable Large Vision-Language Model enhanced by Reinforcement Learning.
oceanscout-sft-v1-duplicate
OceanScout SFT
Maritime SFT samples from Sentinel-2. The pipeline searches a temporal STAC pair (for robust scene choice / metadata) but each training row uses a single post-scene RGB chip per tile. NDWI defines water; bright targets on water suggest vessel candidates. Each tile has a maritime caption row and, when detections exist, a grounding row.
Record counts (this build)
Split
JSONL lines
train
38
validation
16
test
15
total
69
Tiles… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/oceanscout-sft-v1-duplicate.isro-space-ocean-dataset
ISRO Multimodal Space & Ocean Telemetry Dataset
Official open-source scientific dataset curated for the National Space Day 2026 Hackathon and ISRO/IN-SPACe research submissions.
Dataset Structure
rain.jsonl: 1,204 high-precision instruction-tuning pairs mapping 6-band multispectral satellite telemetry (Coastal, Blue, Green, Red, NIR, SWIR) to atmospheric composition (O2 %, N2 %, Water Vapor g/m3) and oceanographic parameters (SST deg C, Salinity PSU).… See the full description on the dataset page: https://huggingface.co/datasets/Anoopsingh53/isro-space-ocean-dataset.OceanInstruct-v0.1We release OceanInstruct, which is part of the instruction data for training OceanGPT.
🛠️ How to use OceanInstruct
We provide the example and you can modify the input according to your needs.
from datasets import load_dataset
dataset = load_dataset("zjunlp/OceanInstruct")
🚩Citation
Please cite the following paper if you use OceanInstruct in your work.
@article{bi2023oceangpt,
title={OceanGPT: A Large Language Model for Ocean Science Tasks},
author={Bi, Zhen and… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/OceanInstruct-v0.1.OceanInstruction
OceanInstruction Dataset
1. Dataset Description
OceanInstruction is a specialized instruction-tuning dataset designed for multimodal large language models (MLLMs) in the marine domain. The data has been rigorously curated, deduplicated, and standardized. It encompasses a diverse range of tasks, spanning from text-only encyclopedic QA and sonar image-based QA to RGB natural image QA (covering biological specimens and scientific diagrams).
2. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/cola60/OceanInstruction.Global-Ocean-Science-Corpus
🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned)
A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography
Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institutes.
Global-Ocean-Science-Corpus, derin okyanus bilimleri, deniz biyolojisi… See the full description on the dataset page: https://huggingface.co/datasets/tilikumotp/Global-Ocean-Science-Corpus.oceanscout-sft-v1
OceanScout SFT
Maritime SFT samples from Sentinel-2. The pipeline searches a temporal STAC pair (for robust scene choice / metadata) but each training row uses a single post-scene RGB chip per tile. NDWI defines water; bright targets on water suggest vessel candidates. Each tile has a maritime caption row and, when detections exist, a grounding row.
Record counts (this build)
Split
JSONL lines
train
1336
validation
166
test
194
total
1696
Tiles… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/oceanscout-sft-v1.kittirewrite_geneval_t2icompbenchOCEANOceanBenchmark
OceanBenchmark
1. Dataset Description
OceanBenchmark is a benchmark dataset designed to evaluate the comprehensive capabilities of marine-focused large models. It encompasses a diverse range of tasks, spanning from unimodal marine science knowledge question answering to complex multimodal visual question answering.
2. Sub-datasets
Subset Directory
Task Type
Sample Size
Description
Example
VQA
816
Combined multimodal examples with Sonar… See the full description on the dataset page: https://huggingface.co/datasets/cola60/OceanBenchmark.TOA-Ultrafeedback-DPO-TOA-model-num-4OceanBenchWe release OceanBench, which is part of the dataset for evaluating OceanGPT.
🛠️ How to use OceanGPT
We provide the example and you can modify the input according to your needs.
from datasets import load_dataset
dataset = load_dataset("zjunlp/OceanBench")
🚩Citation
Please cite the following paper if you use OceanBench in your work.
@article{bi2023oceangpt,
title={OceanGPT: A Large Language Model for Ocean Science Tasks},
author={Bi, Zhen and Zhang, Ningyu and Xue… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/OceanBench.TOA-Ultrafeedback-SFT-SeqRefine-model-num-4baseline_train_split
Baseline Training Data
This link provides processed training data, which differs from the evaluation data shown in another dataset card here. Specifically, this data can be used with different splits to train your baselines or develop new unlearning algorithms.
For example, if you want to implement the Gradient Ascent (GA) algorithm with a 5% split of the data, you can use the provided data to update the gradient in the opposite direction of the forgotten data. Alternatively, if… See the full description on the dataset page: https://huggingface.co/datasets/oceanoceanna/baseline_train_split.vietnam-real-estates
🏠 Tinix Vietnam Real Estate Listings (2025-2026)
Tinix Vietnam Real Estate Listings 2025-2026 là bộ dữ liệu bất động sản Việt Nam quy mô lớn được thu thập và xử lý bởi TiniX AI, bao gồm đúng 3.500.744 tin đăng bán/cho thuê bất động sản từ tháng 6/2025 đến tháng 3/2026 sau khi đã qua bước lọc loại hình nghiêm ngặt (loại bỏ Nhà mặt phố, Nhà trong ngõ). Đây là tài nguyên phục vụ nghiên cứu về thị trường bất động sản, xây dựng mô hình định giá nhà, phân tích xu hướng thị trường tại… See the full description on the dataset page: https://huggingface.co/datasets/oceanNG/vietnam-real-estates.
