datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mb-s5mars
mb-s5mars
A segmentation dataset for planetary science applications.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-10-24
Cite As: TBD
Classes
This dataset contains the following classes:
0: Background
1: Bedrock
2: Hole
3: Ridge
4: Rock
5: Rover
6: Sand / Soil
7: Sky
8: Track
Directory Structure
The dataset follows this structure:
dataset/
├── train/
│ ├── images/ #… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-s5mars.grpo-dapo_shuffled-01_offline-grpo-dapo-qwen3-1.7B-Base-mbs128-n4-mbs128-n4_mathevalgrpo-dapo_ordered-0_offline-grpo-dapo-qwen3-4B-Base-mbs128-n4-mbs128-n4_mathevalDAPO-Math-17k-grpo-dapo-qwen3-4B-Base-mbs128-n4_sft_mathevalgrpo-dapo_ordered-0_offline-grpo-dapo-qwen3-1.7B-Base-mbs128-n4-mbs128-n4_mathevalgrpo-dapo_shuffled-0_offline-grpo-dapo-qwen3-1.7B-Base-mbs128-n4-mbs128-n4_mathevalDAPO-Math-17k-grpo-dapo-qwen3-1.7B-Base-mbs128-n4_sft_mathevalaustralian-mbs-source-archive
Australian MBS source archive
Public, content-addressed preservation of the exact legacy MBS XML, P7 workbook, and complete Git ancestry through the pinned donor commits. The donor code licences and publication authorization are included; no single software licence is asserted over government source payloads. MBS service-benefit evidence is not medicine regulatory or PBS funding evidence.
grpo-dapo-qwen2.5math-1.5B-base-mbs256-n8_actor_mathevalgrpo-dapo-01_offline-qwen2.5math-1.5B-base-mbs256-n8_actor_mathevalpg-dapo-01_offline-qwen2.5math-1.5B-base-mbs256-n8_actor_mathevalgrpo-dapo_offline-qwen2.5math-1.5B-base-mbs256-n8_actor_mathevalmbspeech-mn
MBSpeech Mongolian (cleaned)
A quality-filtered, normalised subset of MBSpeech Mongolian (Bible read-speech), prepared for training
Mongolian (Khalkha Cyrillic) text-to-speech with
oron-tts.
Built by oron-cleaner. Every threshold
was calibrated on this corpus, and every number and column on this page is read
from the shipped data rather than asserted.
from datasets import load_dataset
ds = load_dataset("btsee/mbspeech-mn-clean", split="train")
print(ds[0]["text"]… See the full description on the dataset page: https://huggingface.co/datasets/btsee/mbspeech-mn.pg-dapo_offline-qwen2.5math-1.5B-base-mbs256-n8_actor_mathevalgrpo-dapo_shuffled-0_offline-grpo-dapo-qwen3-4B-Base-mbs128-n4-mbs128-n4_mathevalaustralian-mbs-utilisation-archivegrpo-dapo_ordered-01_offline-grpo-dapo-qwen3-1.7B-Base-mbs128-n4-mbs128-n4_mathevalpg-dapo_ordered-01_offline-grpo-dapo-qwen3-4B-Base-mbs128-n4-mbs128-n4_mathevalmb-surface_cls
mb-surface_cls
A Mars image classification dataset for planetary science research.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-05-14
Cite As: TBD
Classes
This dataset contains the following classes:
0: apx
1: act
2: arm
3: art
4: cct
5: cio
6: clr
7: dls
8: dri
9: drh
10: drp
11: drt
12: flr
13: gro
14: hor
15: inl
16: lar
17: ltv
18: mah
19: mct
20: mas
21: mca
22: nsk
23: obt
24:… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-surface_cls.pg-dapo_ordered-0_offline-grpo-dapo-qwen3-4B-Base-mbs128-n4-mbs128-n4_mathevalmb-surface_multi_label_cls
MER - Mars Exploration Rover Dataset
A multi-label classification dataset containing Mars images from the Mars Exploration Rover (MER) mission for planetary science research.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-10-23
Cite As: TBD
Classes
This dataset uses multi-label classification, meaning each image can have multiple class labels.
The dataset contains the following classes:… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-surface_multi_label_cls.pg-dapo_shuffled-0_offline-pg-dapo-qwen3-4B-Base-mbs128-n4_kl-grpo_mathevalXSUMUrdu-DW_BBC
Urdu_DW-BBC-512
Dataset Summary
Urdu Summarization Dataset containining 76,637 records of Article + Summary pairs scrapped from BBC Urdu and DW Urdu News Websites.
Preprocessed Version: upto 512 tokens (~words); removed URLs, Pic Captions etc
Supported Tasks and Leaderboards
Summarization: Extractive and Abstractive
urT5 adapted from mT5 having monolingual vocabulary only; 40k tokens of Urdu.
Fine-tuned version @… See the full description on the dataset page: https://huggingface.co/datasets/mbshr/XSUMUrdu-DW_BBC.mbspeech
mbspeech
Mongolian speech recognition dataset
Dataset Statistics
Total samples: 3,846Total duration: 6h 37m 43s (6.63 h)
Per-split breakdown
Split
Samples
Total Duration
Avg Duration
train
3,846
6h 37m 43s (6.63 h)
6.20 s
DAPO-Math-17k-pg-dapo-qwen3-4B-Base-mbs128-n4_shuffledpg-dapo_shuffled-0_offline-pg-dapo-qwen3-4B-Base-mbs128-n4_kl_behavior_mathevalgrpo-dapo_shuffled-01_offline-grpo-dapo-qwen3-4B-Base-mbs128-n4-mbs128-n4_mathevalgrpo-dapo_shuffled-005_offline-grpo-dapo-qwen3-1.7B-Base-mbs128-n4-mbs128-n4_mathevalsynthetic_mbspeech_dataset_edgettsDAPO-Math-17k-grpo-dapo-qwen3-1.7B-Base-mbs128-n4_ordered
