datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OceanDepths
OceanDepths GeoTIFF Raster and Aligned ARGO Dataset
This dataset package contains the model-ready Ocean variables (ARGO submarine data, sea surface height, sea surface temperature and
salinity, as well as GLORYS reanalysis information for 50 depth levels. The ARGO data has been projected onto the GLORYS grid in order
to build a ML-ready dataset. The intention is that users can create tensors easily for CV-inspired ML approaches to ocean-variable
reconstruction. While… See the full description on the dataset page: https://huggingface.co/datasets/ESA-philab/OceanDepths.OceanCorpus
OceanCorpus
Dataset Description
OceanCorpus is a large-scale, multimodal dataset designed to inject structured marine domain knowledge into Large Language Models (LLMs). It aggregates data from three primary sources to support text generation, instruction tuning, and vision-language alignment:
Web Knowledge (Text-Only): A dataset of 113,626 instruction-style QA pairs extracted from Wikipedia and authoritative marine websites, available in Web/data.csv.
Paper… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/OceanCorpus.planktonzilla-17M
Planktonzilla-17M Dataset
Overview
planktonzilla-17M is a large-scale, comprehensive dataset combining 17 million plankton images from all publicly available -to the best
of our knowledge- labeled plankton datasets. This unified collection enables researchers to train robust deep learning models for plankton
identification and classification across diverse imaging systems and oceanographic environments.
Each image includes a standardized taxonomic hierarchy… See the full description on the dataset page: https://huggingface.co/datasets/project-oceania/planktonzilla-17M.OceanInstruction
OceanInstruction Dataset
1. Dataset Description
OceanInstruction is a specialized instruction-tuning dataset designed for multimodal large language models (MLLMs) in the marine domain. The data has been rigorously curated, deduplicated, and standardized. It encompasses a diverse range of tasks, spanning from text-only encyclopedic QA and sonar image-based QA to RGB natural image QA (covering biological specimens and scientific diagrams).
2. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/OceanInstruction.10132025_human_demonstrations_tiktokmug_oceantablecloth_reorientThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 151,
"total_frames": 6731,"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:151"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/10132025_human_demonstrations_tiktokmug_oceantablecloth_reorient.OceanBenchmark
OceanBenchmark
1. Dataset Description
OceanBenchmark is a benchmark dataset designed to evaluate the comprehensive capabilities of marine-focused large models. It encompasses a diverse range of tasks, spanning from unimodal marine science knowledge question answering to complex multimodal visual question answering.
2. Sub-datasets
Subset Directory
Task Type
Sample Size
Description
Example
VQA
816
Combined multimodal examples with Sonar… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/OceanBenchmark.Ocean_R1_collected_visual_datalensless
Dataset Lensless Freshwater Plankton Dataset
A mixture of ten freshwater plankton species imaged with a lensless microscope designed for in situ data collection.
Original dataset available online at: https://ibm.ent.box.com/v/PlanktonData.
Original dataset license: <cc-by-4.0>.
Details
train split means (RGB): [0.4140324021223468, 0.43283753018640264, 0.4201652068819248]
train split standard deviations (RGB): [0.12368459767878866, 0.1250351586616661… See the full description on the dataset page: https://huggingface.co/datasets/project-oceania/lensless.OceanInstruct-oWe release OceanInstruct-o, a bilingual Chinese-English multimodal instruction dataset of approximately 50K samples in the ocean domain, constructed from publicly available corpora and data collected via ROV (recent update 20250506). Part of the instruction data is used for training zjunlp/OceanGPT-o.
❗ Please note that the models and data in this repository are updated regularly to fix errors. The latest update date will be added to the README for your reference.
🛠️ How to use… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/OceanInstruct-o.PlasticInWaterThis model was trained on images of different types of plastic placed inside a tank connected to a 1080p webcam, located in the Ocean Technology Center at the University of Washington. The goal of this project is to enhance the detection and classification of various types of plastic debris commonly found in marine environments, providing valuable tools for environmental monitoring and research.
license: MIT
Dataset Details
The dataset consists of 4,511 images, capturing… See the full description on the dataset page: https://huggingface.co/datasets/OceanCV/PlasticInWater.zoolake
Dataset ZooLake Plankton Dataset
Plankton images annotated into 35 classes over 17900 images of zooplankton and large phytoplankton colonies, detected in Lake Greifensee (Switzerland) with the Dual Scripps Plankton Camera.
Original dataset available online at: https://opendata.eawag.ch/dataset/52b6ba86-5ecb-448c-8c01-eec7cb209dc7/resource/1cc785fa-36c2-447d-bb11-92ce1d1f3f2d/download/data.zip.
Original dataset license: <cc-by-4.0>.
Details
train split means (RGB):… See the full description on the dataset page: https://huggingface.co/datasets/project-oceania/zoolake.oceanscout-sft-v1-duplicate
OceanScout SFT
Maritime SFT samples from Sentinel-2. The pipeline searches a temporal STAC pair (for robust scene choice / metadata) but each training row uses a single post-scene RGB chip per tile. NDWI defines water; bright targets on water suggest vessel candidates. Each tile has a maritime caption row and, when detections exist, a grounding row.
Record counts (this build)
Split
JSONL lines
train
38
validation
16
test
15
total
69
Tiles… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/oceanscout-sft-v1-duplicate.OceanInstruction
OceanInstruction Dataset
1. Dataset Description
OceanInstruction is a specialized instruction-tuning dataset designed for multimodal large language models (MLLMs) in the marine domain. The data has been rigorously curated, deduplicated, and standardized. It encompasses a diverse range of tasks, spanning from text-only encyclopedic QA and sonar image-based QA to RGB natural image QA (covering biological specimens and scientific diagrams).
2. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/cola60/OceanInstruction.oceanscout-sft-v1
OceanScout SFT
Maritime SFT samples from Sentinel-2. The pipeline searches a temporal STAC pair (for robust scene choice / metadata) but each training row uses a single post-scene RGB chip per tile. NDWI defines water; bright targets on water suggest vessel candidates. Each tile has a maritime caption row and, when detections exist, a grounding row.
Record counts (this build)
Split
JSONL lines
train
1336
validation
166
test
194
total
1696
Tiles… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/oceanscout-sft-v1.kittiOceanBenchmark
OceanBenchmark
1. Dataset Description
OceanBenchmark is a benchmark dataset designed to evaluate the comprehensive capabilities of marine-focused large models. It encompasses a diverse range of tasks, spanning from unimodal marine science knowledge question answering to complex multimodal visual question answering.
2. Sub-datasets
Subset Directory
Task Type
Sample Size
Description
Example
VQA
816
Combined multimodal examples with Sonar… See the full description on the dataset page: https://huggingface.co/datasets/cola60/OceanBenchmark.baseline_train_split
Baseline Training Data
This link provides processed training data, which differs from the evaluation data shown in another dataset card here. Specifically, this data can be used with different splits to train your baselines or develop new unlearning algorithms.
For example, if you want to implement the Gradient Ascent (GA) algorithm with a 5% split of the data, you can use the provided data to update the gradient in the opposite direction of the forgotten data. Alternatively, if… See the full description on the dataset page: https://huggingface.co/datasets/oceanoceanna/baseline_train_split.MLLMU-Bench
Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench
Abstract
Generative models such as Large Language Models (LLM) and Multimodal Large Language models (MLLMs) trained on massive web corpora can memorize and disclose individuals' confidential and private data, raising legal and ethical concerns. While many previous works have addressed this issue in LLM via machine unlearning, it remains largely unexplored for MLLMs. To tackle this challenge, we… See the full description on the dataset page: https://huggingface.co/datasets/oceanoceanna/MLLMU-Bench.10132025_human_demonstrations_pinkbowltaupebowl_oceantablecloth_stackThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 151,
"total_frames": 10824,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:151"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/10132025_human_demonstrations_pinkbowltaupebowl_oceantablecloth_stack.10132025_pick_tiktokmug_oceantablecloth_reorient_50demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 8,
"total_frames": 2385,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/10132025_pick_tiktokmug_oceantablecloth_reorient_50demos.OceanCorpus
OceanCorpus
Dataset Description
OceanCorpus is a large-scale, multimodal dataset designed to inject structured marine domain knowledge into Large Language Models (LLMs). It aggregates data from three primary sources to support text generation, instruction tuning, and vision-language alignment:
Web Knowledge (Text-Only): A dataset of 113,626 instruction-style QA pairs extracted from Wikipedia and authoritative marine websites, available in Web/data.csv.
Paper… See the full description on the dataset page: https://huggingface.co/datasets/cola60/OceanCorpus.10132025_human_demonstrations__bluespoon_oceantablecloth_pickThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 151,
"total_frames": 10952,"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:151"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/10132025_human_demonstrations__bluespoon_oceantablecloth_pick.09182025_realsense_oceanstripedtable_neutronsmugThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 10,
"total_frames": 8049,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/09182025_realsense_oceanstripedtable_neutronsmug.alldata_public10132025_pinkbowltaupebowl_oceantablecloth_stack_50demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 15,
"total_frames": 3587,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/10132025_pinkbowltaupebowl_oceantablecloth_stack_50demos.planktonzilla-frepj
FREPJ-Z: Freshwater Plankton in Japanese Lakes and Reservoirs (I. Zooplankton)
intermediate validation build (v1.2) — this repository is the FREPJ-only intermediate testing/validation
build. It is deliberately distinct from the frozen, immutable
project-oceania/planktonzilla-17M composite (which is NOT modified by this build) and from
the forthcoming full composite planktonzilla-v1.2.
Attribution
FREPJ-Z (Freshwater Plankton in Japanese Lakes and Reservoirs, I.… See the full description on the dataset page: https://huggingface.co/datasets/project-oceania/planktonzilla-frepj.10132025_pick_bluespoon_oceantablecloth_pick_50demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 13520,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/10132025_pick_bluespoon_oceantablecloth_pick_50demos.mllm_attack_datajojo-stone-ocean-blip-captions-512
Dataset Card for "jojo-stone-ocean-blip-captions-512"
JoJo's Bizarre Adventure: Stone Ocean with Blip captions.
Dataset contains 512x512 cropped images whose source is jojowiki
UBC-OCEAN-Thumbnails
UBC-OCEAN
UBC Ovarian Cancer Subtype Classification and Outlier Detection [UBC-OCEAN] is the world's most extensive ovarian cancer dataset of histopathology images obtained from more than 20 medical centers.
Navigating Ovarian Cancer: Unveiling Common Histotypes and Unearthing Rare Variants
Citation
@misc{UBC-OCEAN,
author = {Ali Bashashati, Hossein Farahani, OTTA Consortium, Anthony Karnezis, Ardalan Akbari, Sirim Kim, Ashley Chow, Sohier Dane, Allen Zhang, Maryam… See the full description on the dataset page: https://huggingface.co/datasets/NouRed/UBC-OCEAN-Thumbnails.
