datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cleaned_manga_dataset
Licensing and source material
This repository contains derived/processed images from publicly accessible manga
sources (GANMA! and Comic Walker / カドコミ). The original works are copyrighted by
their respective rights holders. No ownership of the original artwork is claimed.
The CC BY-NC 4.0 license applies only to the annotations and dataset metadata
created by the dataset authors, not to the underlying original artwork.
Unpacking the dataset
The dataset ships as a… See the full description on the dataset page: https://huggingface.co/datasets/Vasyanator2/cleaned_manga_dataset.mnist-cleaned-full
Dataset Card for 2025.11.21.16.40.44.970939
This is a FiftyOne dataset with 69807 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Linus-L/mnist-cleaned-full")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Linus-L/mnist-cleaned-full.cleaned-plotqa-v2cleaned_webtoon_dataset
Licensing and source material
This repository contains derived/processed images from publicly accessible web comics/manga sources. The original works are copyrighted by their respective rights holders. No ownership of the original artwork is claimed.
The CC BY-NC 4.0 license applies only to the annotations and dataset metadata created by the dataset authors, not to the underlying original artwork.
Unpacking the dataset
The dataset ships as a zstd-compressed… See the full description on the dataset page: https://huggingface.co/datasets/Vasyanator2/cleaned_webtoon_dataset.OpenDetection-50K-Remastered-Cleaned
OpenDetection-50K-Remastered-Cleaned
OpenDetection-50K-Remastered-Cleaned is the cleaned version of OpenDetection-50K-Remastered, created by removing every sample that contains no detected objects. This ensures that every image in the dataset includes at least one valid object annotation, making the dataset more suitable for training, evaluation, and benchmarking object detection models. The dataset is built primarily from general, publicly available images, which make up the… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenDetection-50K-Remastered-Cleaned.OpenDetection-80K-Unified-Cleaned
OpenDetection-80K-Unified-Cleaned
OpenDetection-80K-Unified-Cleaned is a large-scale object detection dataset built primarily from general, publicly available images, which make up the majority of the input imagery, together with additional publicly available datasets. This dataset is a unified collection created by combining OpenDetection-15K-Dense-v1.0, OpenDetection-15K-Dense-v2.0, and OpenDetection-50K-Remastered-Cleaned into a single standardized dataset. Every sample… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenDetection-80K-Unified-Cleaned.cc12m-cleaned
CC12m-cleaned
This dataset builds on two others: The Conceptual Captions 12million dataset, which lead to the LLaVa captioned subset done by
CaptionEmporium
(The latter is the same set, but swaps out the (Conceptual Captions 12million) often-useless alt-text captioning for decent ones_
I have then used the llava captions as a base, and used the detailed descrptions to filter out
images with things like watermarks, artist signatures, etc.
I have also manually thrown out all… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned.mapillary_traffic_sign_dataset_cleaned1920-raider-waite-tarot-public-domain-cleanedA cleaned up version of the multimodalart/1920-raider-waite-tarot-public-domain dataset, without the card borders and names
Japanese_Photo_conversation_cleaned
Japanese Photo Conversation (Cleaned)
A cleaned and organized Japanese photo conversation dataset for training vision-language models on Japanese photo description and visual question answering tasks.
Dataset Description
This dataset is a cleaned and reorganized version combining data from:
llm-jp/japanese-photos-conversation
ThePioneer/japanese-photos
We thank the original authors for their excellent work in collecting and annotating these datasets.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/WayBob/Japanese_Photo_conversation_cleaned.mnist-cleaned-up
Dataset Card for cleaned-up-mnist-training-set
This is a FiftyOne dataset with 505 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Linus-L/mnist-cleaned-up")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Linus-L/mnist-cleaned-up.vision-feedback-mix-binarized-cleaned
Dataset Card for Vision-Feedback-Mix-Binarized-Cleaned
Introduction
This dataset represents a cleaned version on wangclnlp/vision-feedback-mix-binarized.
Descriptions of the base datasets, including the data format and the procedure for mixing data, can be found in this link.
Our Methods for Cleaning Vision Feedback Data
Our goal is to select vision feedback samples where the preferred outputs are significantly differentiated from the dispreferred ones, and the… See the full description on the dataset page: https://huggingface.co/datasets/NiuTrans/vision-feedback-mix-binarized-cleaned.downsampled_cleaned_chartQa_plotQaHungarianDocQA_IT_SynQA_ocr_v3_cleanedsynthdog_cleaned
synthdog_cleaned
The synthdog__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
445,694
QA turns
1,613,204
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
547
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/synthdog_cleaned.downsampled_cleaned_chartQa_plotQa_distributedAndStandardizedDoclingMatix_cleaned
DoclingMatix_cleaned
The DoclingMatix__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
588,763
QA turns
6,394,614
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
503
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/DoclingMatix_cleaned.downsampled_cleaned_chartQa_plotQa_colored_standardDocmatix_merged_cleaned
Docmatix_merged_cleaned
The Docmatix_merged family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
547,033
QA turns
6,199,743
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
507
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/Docmatix_merged_cleaned.cleaned-plotqa-v2-difficulty
Cleaned-PlotQA v2 with difficulty tiers (vectorized + calibrated)
This repository augments jrc/cleaned-plotqa-v2 by adding a single column difficulty_tier ∈ {easy, medium, hard} computed with a vectorized, batch‑scored rule set and cutoffs calibrated on a 1,000‑example sample to avoid tier collapse.
Tier counts
easy: 77403
medium: 78521
hard: 43369
total labeled: 199293
Notes
Only one new column is added; original fields remain unchanged.
The scoring runs… See the full description on the dataset page: https://huggingface.co/datasets/Mohamed-Abbas/cleaned-plotqa-v2-difficulty.downsampled_cleaned_chartQa_plotQa_coloredMIT-States-Cleaned-Subset-edited-v4-s1.0-g3.5cardiology-cleaned_datasetCleaned_Augmented_CASIA_FaceV5
Cleaned_Augmented_CASIA_FaceV5
在众多人脸识别数据集中,大多都由白人面孔组成,为了增强人脸识别模型对亚洲面孔的精确度,本数据集在亚洲人脸的重要数据集 CASIA_FaceV5 的基础上进行了清洗和数据增强操作,以增强其质量。
清洗:
通过 RtinaFace 模型进行人脸检测和关键点提取,基于这些提取到的人脸信息,对人脸进行仿射变换对齐和裁剪
数据增强:
通过翻转、旋转、模糊、亮度、遮挡和调色等多种方式扩充了数据集规模,丰富了数据多样性,使其更适用于人脸识别、人脸分析等相关的机器学习和深度学习任务。
数据集规模
500个人
32500张图像(每人65张)
tags:
人脸识别
CASIA_FaceV5
亚洲人脸数据集
数据增强
Reference… See the full description on the dataset page: https://huggingface.co/datasets/JustinLeee/Cleaned_Augmented_CASIA_FaceV5.danbooru-cleaned
What is this?
This hosts a collaborative effort to clean up the mess that is the danbooru dataset.
The Danbooru dataset is a wealth of (mostly) free anime style images,
that are already individually tagged!
The problem being that the images are indiscriminately included.
Some pics are explicitly copyrighted and shouldnt be in it.
Some have watermarks.
Some have legally questionable subject matter.
Some just frankly arent good.
But the rest... are really good
So, this is a public… See the full description on the dataset page: https://huggingface.co/datasets/ppbrown/danbooru-cleaned.CSIRO-Cleaned-Datamnist-cleaneddownsampled_cleaned_chartQa_plotQa_colored_bulkcleaned-figureqa-v2gtsrb-cleaned
