datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SymVAE
Geometric Logic-Form Dataset (SymParser / SymVAE / SymHPR)
Dataset for the CVPR 2026 paper Hierarchical Process Reward Models are Symbolic Vision Learners — Shan Zhang, Aotian Chen, Kai Zou, Jindong Gu, Yuan Xue, Anton van den Hengel.
🔗 Project page: vi-ocean.github.io/projects/SymVAE — SymVAE stands for Symbolic Variational Auto-Encoder.
A multimodal dataset for training and evaluating MLLMs on geometric diagram → logic-form extraction. Given a geometry image, a model must… See the full description on the dataset page: https://huggingface.co/datasets/vi-ocean/SymVAE.oceanscout-sft-v1-duplicate
OceanScout SFT
Maritime SFT samples from Sentinel-2. The pipeline searches a temporal STAC pair (for robust scene choice / metadata) but each training row uses a single post-scene RGB chip per tile. NDWI defines water; bright targets on water suggest vessel candidates. Each tile has a maritime caption row and, when detections exist, a grounding row.
Record counts (this build)
Split
JSONL lines
train
38
validation
16
test
15
total
69
Tiles… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/oceanscout-sft-v1-duplicate.isro-space-ocean-dataset
ISRO Multimodal Space & Ocean Telemetry Dataset
Official open-source scientific dataset curated for the National Space Day 2026 Hackathon and ISRO/IN-SPACe research submissions.
Dataset Structure
rain.jsonl: 1,204 high-precision instruction-tuning pairs mapping 6-band multispectral satellite telemetry (Coastal, Blue, Green, Red, NIR, SWIR) to atmospheric composition (O2 %, N2 %, Water Vapor g/m3) and oceanographic parameters (SST deg C, Salinity PSU).… See the full description on the dataset page: https://huggingface.co/datasets/Anoopsingh53/isro-space-ocean-dataset.OceanInstruct-v0.1We release OceanInstruct, which is part of the instruction data for training OceanGPT.
🛠️ How to use OceanInstruct
We provide the example and you can modify the input according to your needs.
from datasets import load_dataset
dataset = load_dataset("zjunlp/OceanInstruct")
🚩Citation
Please cite the following paper if you use OceanInstruct in your work.
@article{bi2023oceangpt,
title={OceanGPT: A Large Language Model for Ocean Science Tasks},
author={Bi, Zhen and… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/OceanInstruct-v0.1.oceanscout-sft-v1
OceanScout SFT
Maritime SFT samples from Sentinel-2. The pipeline searches a temporal STAC pair (for robust scene choice / metadata) but each training row uses a single post-scene RGB chip per tile. NDWI defines water; bright targets on water suggest vessel candidates. Each tile has a maritime caption row and, when detections exist, a grounding row.
Record counts (this build)
Split
JSONL lines
train
1336
validation
166
test
194
total
1696
Tiles… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/oceanscout-sft-v1.OceanInstruct-v0.2We release OceanInstruct-v0.2, a bilingual Chinese-English dataset of approximately 50K ocean domain text instructions constructed from publicly available corpora, which includes synthetic data and may therefore contain errors (recent update 20250506). Part of the instruction data is used for training OceanGPT.
❗ Please note that the models and data in this repository are updated regularly to fix errors. The latest update date will be added to the README for your reference.
🛠️ How… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/OceanInstruct-v0.2.alpaca_user_preference
Preference-Guided Reflective Sampling for Aligning Language Models
We released the dataset used in our EMNLP2024 paper PRS. Our project website is here.
1. Usage of the dataset
This dataset can be used for personalized model alignment, which means the model is trained to generate personalized outputs, such as the user preference can be "I expect the response to be humorous" or "I prefer the response to provide support evidence".
2. Dataset Information
The… See the full description on the dataset page: https://huggingface.co/datasets/oceanpty/alpaca_user_preference.OCEAN-Chatoceanguard-marine-debris-eval-1000
OceanGuard AI — Marine Debris Evaluation Hold-out (annotations only)
The held-out evaluation split used to report the LoRA adapter delta in the
OceanGuard AI Kaggle Gemma 4 Good Hackathon submission
(Global Resilience track + Unsloth bonus track).
Important — this repository contains only the annotations and metadata.
The 1 000 underwater / coastal images are not redistributed here. They
come from three pre-existing third-party datasets, each with its own
license. Reviewers and… See the full description on the dataset page: https://huggingface.co/datasets/asferrer/oceanguard-marine-debris-eval-1000.oceanservice-noaa-factsoceanjust add ocean data into alpaca-cleaned
ocean_onlyeew_statusSocio-oceanography-no-hypergraphoceansSocio-oceanography
LLM training dataset for socio-oceanography auto-modeling
This file aims to build the first high-quality dataset for socio-oceanography auto-modeling. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/T20091125276/Socio-oceanography.lima150_selfcurated_q45ocean-sysprompt-e-plus
