datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dirbe-zsma
COBE/DIRBE Zodi-Subtracted Mission Average maps
This dataset contains the ten COBE/DIRBE Zodi-Subtracted Mission Average maps
served by LAMBDA at 1.25, 2.2, 3.5, 4.9, 12, 25, 60, 100, 140, and 240
microns. Each configuration is one source FITS table with 393,216 rows in the
COBE resolution-9 quadrilateralized spherical-cube pixelisation.
python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/dirbe-zsma.fineweb-2_zsm-filtered-0.99chatgpt-openqa-zsm-qaretrieval
chatgpt-openqa-zsm-qaretrieval
Deduplicated copy of kornwtp/chatgpt-openqa-zsm-qaretrieval,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/chatgpt-openqa-zsm-qaretrieval
Deduplicated on: 2026-09-04
Task type: qa_retrieval
Splits: train
What changed
Kept in this dataset's ORIGINAL schema -- same columns, same nesting, same extra fields (ids, titles, answers) -- so it is a drop-in replacement for the source repo. Documents differing only in… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/chatgpt-openqa-zsm-qaretrieval.alt-fil-zsm-bitextmining
alt-fil-zsm-bitextmining
Deduplicated copy of kornwtp/alt-fil-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-fil-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-fil-zsm-bitextmining.massive-scenario-zsm-classification
MassiveScenario_zsm_Classification
Deduplicated copy of kornwtp/massive-scenario-zsm-classification.
Splits
split
rows
test
2,930
train
11,151
validation
2,014
sea-translationese-resampled-zsm-classificationalt-khm-zsm-bitextmining
alt-khm-zsm-bitextmining
Deduplicated copy of kornwtp/alt-khm-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-khm-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-khm-zsm-bitextmining.software-documentation-zsm-bitextmining
software-documentation-zsm-bitextmining
Deduplicated copy of kornwtp/software-documentation-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/software-documentation-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-zsm-bitextmining.alt-zsm-lao-bitextmining
alt-zsm-lao-bitextmining
Deduplicated copy of kornwtp/alt-zsm-lao-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-zsm-lao-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-zsm-lao-bitextmining.sea-translationese-resampled-zsm-classification
SEATranslationeseResampled_zsm_Classification
Deduplicated copy of kornwtp/sea-translationese-resampled-zsm-classification.
Splits
split
rows
test
3,633
train
14,693
massive-intent-zsm-classification
MassiveIntent_zsm_Classification
Deduplicated copy of kornwtp/massive-intent-zsm-classification.
Splits
split
rows
test
2,930
train
11,151
validation
2,014
alt-vie-zsm-bitextmining
alt-vie-zsm-bitextmining
Deduplicated copy of kornwtp/alt-vie-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-vie-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-vie-zsm-bitextmining.alt-ind-zsm-bitextmining
alt-ind-zsm-bitextmining
Deduplicated copy of kornwtp/alt-ind-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-ind-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-ind-zsm-bitextmining.xnli-zsm-pairclassification
XNLI_zsm_PairClassification
Deduplicated copy of kornwtp/xnli-zsm-pairclassification.
Splits
split
rows
test
3,339
validation
1,660
alt-mya-zsm-bitextmining
alt-mya-zsm-bitextmining
Deduplicated copy of kornwtp/alt-mya-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-mya-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-mya-zsm-bitextmining.ted2020-zsm-bitextmining
ted2020-zsm-bitextmining
Deduplicated copy of kornwtp/ted2020-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/ted2020-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/ted2020-zsm-bitextmining.tweets-zsm-classification
Tweets_zsm_Classification
Deduplicated copy of kornwtp/tweets-zsm-classification.
Splits
split
rows
train
6,552
alt-tha-zsm-bitextmining
alt-tha-zsm-bitextmining
Deduplicated copy of kornwtp/alt-tha-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-tha-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-tha-zsm-bitextmining.qed-zsm-bitextmining
qed-zsm-bitextmining
Deduplicated copy of kornwtp/qed-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-zsm-bitextmining.news-sentiment-zsm-classification
NewsSentiment_zsm_Classification
Deduplicated copy of kornwtp/news-sentiment-zsm-classification.
Splits
split
rows
train
3,673
alt-zsm-bitextmining
alt-zsm-bitextmining
Deduplicated copy of kornwtp/alt-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-zsm-bitextmining.talpco-zsm-bitextmining
talpco-zsm-bitextmining
Deduplicated copy of kornwtp/talpco-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/talpco-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/talpco-zsm-bitextmining.stsbenchmark-zsm-sts
stsbenchmark-zsm-sts
Deduplicated copy of kornwtp/stsbenchmark-zsm-sts,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/stsbenchmark-zsm-sts
Deduplicated on: 2026-09-04
Task type: sts
Splits: test
What changed
Kept in this dataset's ORIGINAL schema (sentence1/sentence2/score). Identical sentence pairs are collapsed to one row -- a repeat is counted twice in the rank correlation and so carries double weight for no reason -- taking the mean of the… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/stsbenchmark-zsm-sts.massive-scenario-zsm-classification
Dataset Card for "ms-scenario-classification"
More Information needed
shippinglaw
Shipping Law Q&A Dataset Sample
The Shipping Law Q&A Dataset is a curated collection of approximately 1500 question and answer pairs on various topics within shipping law (Using ChatGPT and Claude). Each entry is structured to facilitate training of language models (LLaMA Chat) for the legal domain, particularly within the maritime law context.
Data Structure
Entries in the dataset are presented as JSON objects, each containing a text field with instructional tokens… See the full description on the dataset page: https://huggingface.co/datasets/zsmail/shippinglaw.chatgpt-openqa-zsm-qaretrievalmassive-intent-zsm-classification
Dataset Card for "ms-intent-classification"
More Information needed
news-sentiment-zsm-classificationref: https://github.com/mesolitica/malaysian-dataset/tree/master/sentiment/news-sentiment
xnli-zsm-pairclassificationalt-fil-zsm-bitextmining
