datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mierukochan
Bangumi Image Base of Mieruko-chan
This is the image base of bangumi Mieruko-chan, we detected 52 characters, 3778 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/mierukochan.mieb-train
mieb-train
CLIP training data built from the MIEB
zero-shot image-classification datasets, with the text side re-written so the
class labels are usable as captions.
4,434,196 training rows across 23 datasets.
column
notes
image
passed through from the source repo, not re-encoded
text
the caption — this is what CLIP trains on
label
source class id (-1 where the source has no class list)
class_name
normalized class name
dataset_name
source dataset key — filter… See the full description on the dataset page: https://huggingface.co/datasets/PumeTu/mieb-train.5_Malaysian_Fish_Species_in_Turbid_Water
5 Malaysian Fish Species in Turbid Water
Dataset Overview
This dataset is designed for underwater fish classification under challenging environmental conditions. It supports domain adaptation research by addressing domain shift caused by variations in water clarity, illumination, turbidity, and image quality degradation.
The dataset is specifically constructed for evaluating deep learning-based domain adaptation methods to improve robustness in real-world… See the full description on the dataset page: https://huggingface.co/datasets/mieraehh/5_Malaysian_Fish_Species_in_Turbid_Water.MieDB-100k
MieDB-100k: A Comprehensive Dataset for Medical Image Editing
📄 Introduction
MieDB-100k is a large-scale, high-quality and diverse dataset for text-guided medical image editing,
which includes 104,267 editing data, covering 63 distinct editing targets and 10 diverse medical image modalities.
We categorize editing tasks into three types: Perception, Modification and Transformation, which consider both model's intrinsic understanding and generation abilities on medical… See the full description on the dataset page: https://huggingface.co/datasets/Laiyf/MieDB-100k.MIEBenchparkingparking_labeledparking_labeled_croppedimnet1k_meerkat_mierkat
