datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
marine-species-zh
海纳海洋生物数据集 | Marine Species Dataset (Chinese)
4662 个海洋物种的结构化数据集:学名、中文俗名、完整分类阶元、命名人、OBIS 分布记录数、图片链接(逐行标注授权协议与作者署名)、中文简介。
A structured dataset of 4,662 marine species: scientific names, Chinese vernacular names, full taxonomy, authorities, OBIS occurrence counts, image links (with per-row license and author attribution), and 2,120 Chinese descriptions translated/organized from Chinese Wikipedia.
数据来自海洋科普公益平台 海纳 · hainahub.cn 的物种图鉴底层,随图鉴扩容滚动更新。
数据规模 / Stats
指标… See the full description on the dataset page: https://huggingface.co/datasets/hainahub/marine-species-zh.code-disciplinaire-penal-marine-marchande
Code disciplinaire et pénal de la marine marchande, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-disciplinaire-penal-marine-marchande.marine-dataset-qa
Marine Biology - Instruction Fine-Tuning Dataset (Q&A)
Description
A question-answer dataset on marine biology topics, generated from Wikipedia
articles using the Groq API (LLaMA 3.3 70B). Intended for supervised
fine-tuning (SFT) of language models to answer marine science questions.
Content
Q&A pairs generated from Wikipedia articles across the following categories:
Marine Biology
Marine Ecology
Ocean
Coral Reefs
Marine Mammals
Oceanography
Fisheries Science… See the full description on the dataset page: https://huggingface.co/datasets/Meriem-DH/marine-dataset-qa.marine-dataset-cpt
Marine Biology - Continued Pre-Training Dataset
Description
A corpus of Wikipedia articles covering marine biology and related domains, intended for continued pre-training (CPT) of language models on marine science knowledge.
Content
Plain text articles scraped from Wikipedia across the following categories:
Marine Biology
Marine Ecology
Ocean
Coral Reefs
Marine Mammals
Oceanography
Fisheries Science
Marine Conservation
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Meriem-DH/marine-dataset-cpt.
