CoolFace
Datasetpublic

cola60/OceanCorpus

OceanCorpus Dataset Description OceanCorpus is a large-scale, multimodal dataset designed to inject structured marine domain knowledge into Large Language Models (LLMs). It aggregates data from three primary sources to support text generation, instruction tuning, and vision-language alignment: Web Knowledge (Text-Only): A dataset of 113,626 instruction-style QA pairs extracted from Wikipedia and authoritative marine websites, available in Web/data.csv. Paper… See the full description on the dataset page: https://huggingface.co/datasets/cola60/OceanCorpus.

sourceHugging Facemitupdated 4d agoView on Hugging Face
0likes291downloads
Dataset Card

OceanCorpus

Dataset Description

OceanCorpus is a large-scale, multimodal dataset designed to inject structured marine domain knowledge into Large Language Models (LLMs). It aggregates data from three primary sources to support text generation, instruction tuning, and vision-language alignment:

  1. 1.Web Knowledge (Text-Only): A dataset of 113,626 instruction-style QA pairs extracted from Wikipedia and authoritative marine websites, available in Web/data.csv.
  2. 2.Paper Knowledge (Multimodal): A collection of approximately 300 academic PDFs, including 278 successfully processed papers converted into Markdown and extracted images using the MinerU pipeline. A structured index containing paper identifiers, titles, relative file paths, and image counts is available in Paper/CleanedData/papers.csv. The processed outputs are located in Paper/CleanedData/ProcessedData/.
  3. 3.Open-Dataset (Imagery and Reference Instructions): A collection of 35,323 domain-specific images covering coral species, wild fish, and sonar targets. Reference instruction JSON files are provided in Open-Dataset/OceanSona/ and Open-Dataset/OceanVision/.

Dataset Statistics

MetricValue
Total Structured Entries113,626
LanguageEnglish
Entity Types8 categories (Location, Organism, Process, etc.)
Source PDFs~300
Auxiliary Images35,323
Image BreakdownCoral (3,680), Sonar (9,081), Fish (22,562)

Dataset Structure

The repository is organized as follows:

text
OceanCorpus/
├── Web/
│   └── wiki_updated_full.csv        # Text-only version (113,626 rows)
├── Paper/
│   ├── Pdf/                         # ~300 raw source PDFs
│   └── CleanedData/
│       ├── papers.csv               # Index of 278 successfully processed papers
│       └── ProcessedData/           # MinerU outputs (Markdown files + image folders)
└── Open-Dataset/
    ├── CoralData/                   # 3,680 coral images
    ├── FishData/                    # 22,562 fish images
    ├── SonarData/                   # 9,081 sonar images
    ├── OceanSona/                   # Reference sonar instruction JSON files
    └── OceanVision/                 # Reference vision instruction JSON files

Open-Dataset instruction files

The instruction data can be adapted for supervised fine-tuning (SFT). However, the uploaded JSON files are provided as usage references only and cannot be used for training directly. Their image paths and data organization do not map directly to every image currently hosted in this repository. Users must revise the JSON files according to their own database layout and training pipeline before use.

Data Fields

Paper

The paper index is located at Paper/CleanedData/papers.csv. It contains one row for each successfully processed paper with a matching source PDF and Markdown file. Papers without extracted images are retained with an image_count of 0 and an empty images_dir. |Field| Type| Description| |---|---|---| |paperid|string| Unique paper identifier derived from the processed directory name. |title|string| Paper title extracted from the first level-one heading in the Markdown file. |pdfpath|string| Relative path to the raw source PDF under Paper/Pdf/. |mdpath|string| Relative path to the MinerU-generated Markdown file. |imagesdir|string| Relative path to the extracted image directory; empty when no images were generated. |image_count|integer| Number of extracted images associated with the paper.

Note : Web/data.csv contains only input, output, and entity_type fields.

entity_type Categories:

The dataset categorizes marine knowledge into eight scientifically grounded types: |Type |Description |Examples| |---|---|---| |Location |Marine geographic features and regions |Trenches, currents, reserves, passages| |Instrument |Research and engineering equipment |CTD profilers, ROVs, multibeam sonars| |Organism |Marine species and biological entities |Fish, mammals, corals, phytoplankton| |Process |Oceanographic and ecological mechanisms |Upwelling, thermohaline circulation| |Phenomenon |Observable marine events |Red tides, rogue waves, marine heatwaves| |Substance |Chemical compounds in seawater |Methane hydrates, microplastics| |Property |Physical/chemical parameters |Salinity, pH, density, turbidity| |Theory |Scientific models and frameworks |Ocean conveyor belt, niche theory|

Usage

Loading with Hugging Face Datasets

python
from datasets import load_dataset

# Load the default Web corpus
dataset = load_dataset("zjunlp/OceanCorpus", split="train")

Loading Locally with Pandas

python
import pandas as pd

# Load the index of successfully processed papers
df_papers = pd.read_csv("Paper/CleanedData/papers.csv")

# Load the text-only web version
df_web = pd.read_csv("Web/data.csv")

print(f"Processed papers: {len(df_papers)}")

License

This dataset is released under the MIT License.

Citation

If you use OceanCorpus in your work, please cite:

bibtex
@article{xue2026oceanpile,
  title={OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models},
  author={Xue, Yida and Zhang, Ningyu and Wu, Tingwei and Ma, Zhe and Ji, Daxiong and Wang, Zhao and Zheng, Guozhou and Chen, Huajun},
  journal={arXiv preprint arXiv:2605.00877},
  year={2026}
}
cola60/OceanCorpus · CoolFace