datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mammalMammaSyn
MammaSyn
MAMMA training dataset
Project site: https://mamma.is.tue.mpg.de/
Synthetic training data, each scene rendered from 8 views, in WebDataset format.
Dataset Groups
MammaSyn-Interactions
Includes Harmony4D, Inter-X, InteractionCouple, LatinDance10
Hi4D is currently not included but will be added when we receive permission
MammaSyn-Singles
Includes BEDLAM, MOYO
MammaSyn-Hands
Includes InterHand
SignAvatars is currently not included but will be added… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/MammaSyn.MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M
🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo
Introduction
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.
The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.mammoth_vl_sea_shard_5mammogps
MammoGPS
Dataset Summary
MammoGPS is a benchmark for evaluating vision-language model spatial understanding on 2D mammography. The benchmark is designed for analysis-oriented evaluation rather than single-number leaderboard reporting: the goal is to separate failures of generic localization, medically relevant finding recognition, and landmark-grounded spatial reasoning.
This repository currently includes benchmark task views for:
finding localization
finding… See the full description on the dataset page: https://huggingface.co/datasets/mammovlmbench/mammogps.mammoth_vl_sea_shard_4TurkmenSpeech
Turkmen Speech Dataset (ASR)
This dataset contains 251 hours of Turkmen speech audio with transcriptions, intended for training and evaluating Automatic Speech Recognition (ASR) models.
It is one of the largest publicly available Turkmen speech datasets.
Dataset Overview
Property
Value
Total clips
119,847
Total duration
251.86 hours
Sampling rate
16,000 Hz
Language
Turkmen (tk)
Split
train
Each item includes:
audio: waveform + sampling… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenSpeech.NIDS-Thesis-Experimental-Evidence
thesis_pipeline
Pipeline rebuilt following the supervisor's conditional review (13/07/2026). See SCOPE_FROZEN.md at the parent project root for the frozen scientific scope.
Structure
config/: centralized configuration (paths, seeds)
manifests/: data audits and temporal split manifests
src/data/: dataset preparation and cleaning
src/models/: model training and comparison
src/evaluation/: aggregation and global comparison of results
tests/: leakage tests and… See the full description on the dataset page: https://huggingface.co/datasets/MamadouSY-NIDS-Thesis/NIDS-Thesis-Experimental-Evidence.multilingual-wikipedia-paragraphsgenomes-v4-genome_set-mammals-intervals-v1_255_128-id1_cov1mammalgenomes-v2-genome_set-mammals-intervals-v2_512_256marine_mammal_labelsgenomes-v4-genome_set-mammals-intervals-v5_256_128genomes-v4-genome_set-mammals-intervals-v16_254_127-id0.3_cov0.3vertebrate-v1-cds_mammals_only
marin-dna/vertebrate-v1-cds_mammals_only
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment. This
draft covers the cds region cohort with mammals_only species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases preserve source non-repeat-masked sequence, and… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-cds_mammals_only.genomes-v5-genome_set-mammals-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-mammals-intervals-v5_255_128
Mammals CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
41,848,032 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals-intervals-v5_255_128.medical-reasoningPrompting-MammAlps
Dataset Card for MammAlps-S2 & Prompting-MammAlps
MammAlps-S2 is a camera-trap video dataset of wildlife monitoring in the Swiss National Park, covering two summer field seasons (2023 and 2024).
It contains video clips of seven Alpine mammal species annotated with per-frame bounding boxes, individual tracks, behavioural labels, and frame-level attributes (age for deer species, and sex for adult deer species). It extends MammAlps (see MammAlps) with a new season of data, refined… See the full description on the dataset page: https://huggingface.co/datasets/amathislab/Prompting-MammAlps.ref_seq_vertebrate_non_mammal_part_1metric-mamba-ml2021-hungyi-corpus
Dataset Card for "metric-mamba-ml2021-hungyi-corpus"
More Information needed
wikipedia_paragraphs
Description
This dataset consists of English Wikipedia articles, which first have been split by paragraph breaks and subsequently by spaces. For each resulting token, there is a corresponding binary ner_tag, which is 1 if a token was followed by paragraph break in the original text. There are two deliberate exceptions to this, which can be seen in the dataset generation code:
The text is not split if a paragraph break is preceded by a colon (":"), to avoid lists being separated… See the full description on the dataset page: https://huggingface.co/datasets/mamei16/wikipedia_paragraphs.MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M used in MoCa Pre-training
🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper
Introduction
This is a VQA style dataset used in the modality-aware continual pre-training of MoCa models. It is adapted from MAmmoTH-VL-Instruct-12M by concatenating prompts and responses.
The dataset consists of interleaved multimodal examples. text is a string containing text while imagesare image binaries that can be loaded… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/MAmmoTH-VL-Instruct-12M.en_wikipedia_paragraphsgenomes-v4-genome_set-mammals-intervals-v1_256_128ref_seq_vertebrate_non_mammal_part_2VinDr-MammoMammalNetMAMe2
Dataset Card for "MAMe2"
More Information needed
mammoth_vl_sea_shard_3
