datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M
🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo
Introduction
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.
The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.mammoth_vl_sea_shard_5mammogps
MammoGPS
Dataset Summary
MammoGPS is a benchmark for evaluating vision-language model spatial understanding on 2D mammography. The benchmark is designed for analysis-oriented evaluation rather than single-number leaderboard reporting: the goal is to separate failures of generic localization, medically relevant finding recognition, and landmark-grounded spatial reasoning.
This repository currently includes benchmark task views for:
finding localization
finding… See the full description on the dataset page: https://huggingface.co/datasets/mammovlmbench/mammogps.mammoth_vl_sea_shard_4MAMe2
Dataset Card for "MAMe2"
More Information needed
mammoth_vl_sea_shard_3mamahahanotsuregogamotokanodatta
Bangumi Image Base of Mamahaha No Tsurego Ga Motokano Datta
This is the image base of bangumi Mamahaha no Tsurego ga Motokano datta, we detected 40 characters, 3708 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/mamahahanotsuregogamotokanodatta.MAMA-MIA-Lite
About
This is a preprocessed redistribution of MAMA-MIA (Synapse syn60868042), which is released under the CC BY-NC 4.0 license.
Dataset summary: 1506 breast DCE-MRI scans with expert primary-tumour segmentation masks.
Contents of this repository:
Images/ — 1506 files
Masks/ — 1506 files
📝 Landmark annotations, visualization figures and the benchmark plan files live in 🔥MedVision🔥, where you can load the complete images and annotations from dataset configs.… See the full description on the dataset page: https://huggingface.co/datasets/YongchengYAO/MAMA-MIA-Lite.DMID_Breast_Cancer_Mammography_Dataset
DMID Breast Cancer Mammography Dataset 🎗️
Overview 🔬
This dataset, the Digital Mammography Dataset for Breast Cancer Diagnosis Research (DMID), provides a comprehensive collection of mammogram images intended for research and educational purposes. It aims to facilitate the development and evaluation of computer-aided diagnosis systems for breast cancer detection. 👩⚕️
Dataset Contents 📁
The dataset includes the following components:
TIFF Images: Contains… See the full description on the dataset page: https://huggingface.co/datasets/MyTwinLab/DMID_Breast_Cancer_Mammography_Dataset.mammoth_vl_searetirement-invitation-mamtamammosightr-preprocessed
MammosighTR — Preprocessed Mammography Dataset (BI-RADS)
Preprocessed PNG mammograms with image-level BI-RADS labels, derived from
the nationwide Turkish breast-cancer screening dataset (MammosighTR)
released for the TEKNOFEST 2023 Artificial Intelligence in Health Competition
by the Republic of Turkey Ministry of Health. Original DICOMs are cropped to
the breast region with a YOLOX detector and exported as PNG; we add an
image-level metadata mapping built from the official… See the full description on the dataset page: https://huggingface.co/datasets/gulluk/mammosightr-preprocessed.MAMI-datasetarchforge-mamluk-cairo
ArchForge — Mamluk / Islamic Cairo Architecture Dataset
A small, curated, licence-documented image dataset of Mamluk and Islamic Egyptian
architecture, built for LoRA style adaptation on FLUX.1-dev. 39 images at
512x512px, 33 train / 6 validation.
Built as part of ArchForge — a time-boxed proof of concept,
not a production dataset. The evaluation, the LoRA adapter and the comparison grid
are linked from there.
Why this dataset exists
The target is a style, not a… See the full description on the dataset page: https://huggingface.co/datasets/marwantosolve/archforge-mamluk-cairo.DMID_Breast_Cancer_Mammography_Dataset
DMID Breast Cancer Mammography Dataset 🎗️
Overview 🔬
This dataset, the Digital Mammography Dataset for Breast Cancer Diagnosis Research (DMID), provides a comprehensive collection of mammogram images intended for research and educational purposes. It aims to facilitate the development and evaluation of computer-aided diagnosis systems for breast cancer detection. 👩⚕️
Dataset Contents 📁
The dataset includes the following components:
TIFF Images: Contains… See the full description on the dataset page: https://huggingface.co/datasets/swapnilk2701/DMID_Breast_Cancer_Mammography_Dataset.maml-pde-datasetsone
With You Our Love Will Make It Through Anime Full Character Dataset
This dataset contains images of all characters from the anime "With you, Our Love will Make it Through" sorted into different folders.
Original forum post:
https://diffused.to/Thread-Image-Anime-With-You-Our-Love-Will-Make-It-Through-Full-Character-Dataset
Dataset collection date
Dec 2025
Total Characters:
3
Total Images:
109
Dataset structure:
├── 📂 256x256/
│… See the full description on the dataset page: https://huggingface.co/datasets/mamoth/one.tringbracs_datasetMAMe-Dataset
MAMe Dataset: Museum Artworks Medium
The MAMe Dataset is an image classification dataset focused on the recognition of mediums in artworks and heritage held by museums (e.g., Oil on canvas, Bronze or Woodcut).
The classes considered in the MAMe dataset comprise a wide variety of mediums according to both interpretations of the term. These can range from simple material aspects (e.g., Bronze, Silver or Gold) to complex, high-level techniques (e.g., Faience, Woodblock or Woven… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MAMe-Dataset.mamluk-cairo-style-dataset
Mamluk Cairo architecture — style LoRA dataset
42 curated photos of 17 Mamluk-era buildings in Cairo, each paired with a training caption, used to train
Amr292/mmlkcairo-sdxl-lora.
Contents
NNNN.jpg + NNNN.txt — image and caption (kohya / diffusers compatible), numbered 0002–0043 (0001, a skyline shot, was removed).
metadata.csv — per image: building, view type, Wikimedia Commons title, source URL, author, license, license URL.
scripts/collect_images.py —… See the full description on the dataset page: https://huggingface.co/datasets/Amr292/mamluk-cairo-style-dataset.mammoth_vl_testsynthetic_mammography_csaw
Dataset Card for Synthetic CSAW 100k Mammograms
Dataset Description
This is a synthetic mammogram dataset created with the latent diffusion model from Generative AI for Medical Imaging: extending the MONAI Framework paper.
The generative model was trained on the CSAW-M dataset.
**Paper: https://arxiv.org/abs/2307.15208
**Point of Contact: walter.diaz_sanz@kcl.ac.uk
Dataset Summary
Supported Tasks
Classification masking of cancer in… See the full description on the dataset page: https://huggingface.co/datasets/SinKove/synthetic_mammography_csaw.final_mammoth_datasetmhist_binaryimnet1k_green_mambafotos2laion2b6plus_mammalMemoryMambaDatacars-make-model-year-chunk-12
