datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmu_manga
mmu_manga HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_manga.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be installed via… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_manga.manga-colorization-datasetmangaManga-Drawings
Dataset Card: MangaDF Text-to-Image Prompts Dataset
Dataset Description
Image Source: Images generated by alvdansen/BandW-Manga weights on ChanY/Stable-Flash-Lighting modelPrompt Source: ChatGPT
Overview
The MangaDF Text-to-Image Prompts Dataset is a collection of text prompts paired with corresponding images. The images in this dataset were generated using the alvdansen/BandW-Manga weights applied to the ChanY/Stable-Flash-Lighting diffusion model. This… See the full description on the dataset page: https://huggingface.co/datasets/Chan-Y/Manga-Drawings.manga2Mmanga exhentai site majorly lang EN, JP, from all (nyaa. si) torrents-exclude huge 300GB+ UBUCA, Exhentai Mothercon Archive with no seeders. Convert to .webp and approximately keep only manga with text using DB_TD500_resnet50;
Intended use for bubble box+OCR training model on 2.2M images.
If proced to label OCR of all image, it would take two months...
April 2026: Added ~150 GB in folder [猫の仓库] 汉化本合集(上半)[3037本] of 243,703 images (WebP q85) across 2332 manga, include no-text images, seed… See the full description on the dataset page: https://huggingface.co/datasets/Zicara/manga2M.mangakasantoassistantsanto
Bangumi Image Base of Mangaka-san To Assistant-san To
This is the image base of bangumi Mangaka-san to Assistant-san to, we detected 10 characters, 3298 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/mangakasantoassistantsanto.manga-syntheticmangadex-images-30kMangaZero
Introduction
MangaZero dataset from paper DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
Please see GitHub repo to get the usage
Project page: https://jianzongwu.github.io/projects/diffsensei
The "type" key in character annotation
Type = 0: Characters with clean faces
Type = 1: Characters with vague faces
Type = 2: Exaggerated, artistic characters with appearances that don't conform to typical standards.
prompt-injection-multilayerMangaSegmentation
Advancing Manga Analysis: Comprehensive Segmentation Annotations for the Manga109 Dataset
License
Please check the LICENSE file for more details.
All images in the segmentation annotations are owned and copyrighted by Minshan Xie.
You are automatically granted permission to use the images for academic and commercial
usages, provided that the image credit "Copyrighted by Minshan Xie" is included in
any form of publication, reproduction, redistribution, or derivatives of… See the full description on the dataset page: https://huggingface.co/datasets/MS92/MangaSegmentation.MangaOCR
MangaOCR Dataset
Dataset Details
This is the MangaOCR dataset. We constructed this dataset by consolidating annotations from the Manga109 dataset and the manga onomatopoeia dataset.
It contains approximately 209K narrative text instances, spanning a wide variety of visual styles and layouts.
Japanese_Mangahausa_aug_lex
title: Lexicon Dataset for the Hausa Language
Dataset with English translation
license: cc-by-nd-4.0
manga-colorization-datasetmanga-covers-text-detection
Manga Covers Text Detection
Manually annotated text boxes for manga covers text detection created by @JustANormalTinkerer.
This dataset contains 100 cover images and 526 manually annotated text rectangles. The image column is a Hugging Face image feature containing the original image bytes. The boxes column contains normalized pixel-coordinate bounding boxes in [x_min, y_min, x_max, y_max] format, the original points, label, and shape metadata. annotation_json preserves the… See the full description on the dataset page: https://huggingface.co/datasets/petersunde/manga-covers-text-detection.MangaliCa_EN-HU
MangaliCa Bilingual Image–Caption Dataset (Hungarian–English)
Dataset Description
The MangaliCa Bilingual Image–Caption Dataset is a large-scale Hungarian–English multimodal dataset containing approximately 70 million aligned image–caption pairs.
Each image is paired with:
an original English caption
a machine-translated Hungarian caption
This dataset was created to address the lack of large-scale multimodal resources for Hungarian, enabling bilingual… See the full description on the dataset page: https://huggingface.co/datasets/Obscure-Entropy/MangaliCa_EN-HU.reforma-tributaria-instructcomics_dataset_512_inv_manga_correctedmanga-whisperer
The Manga Whisperer
Automatically Generating Transcriptions for Comics
Ragav Sachdeva and Andrew Zisserman
University of Oxford
Usage
from transformers import AutoModel
import numpy as np
from PIL import Image
import torch
import os
images = [
"path_to_image1.jpg",
"path_to_image2.png",
]
def read_image_as_np_array(image_path):
with open(image_path, "rb") as file:
image =… See the full description on the dataset page: https://huggingface.co/datasets/MattyMroz/manga-whisperer.manga---
description: '
An IFU dataset from the SDSS-IV MaNGA survey, a wide-field, optical, IFU survey
of ~10,000
nearby galaxies. This dataset contains the following data products for each galaxy:
the 3D data cubes,
and reconstructed griz images from the MaNGA Data Reduction Pipeline (DRP), and
all the derived
analsysis maps from the MaNGA Data Analyis Pipeline (DAP).
'
homepage: https://www.sdss4.org/dr17/manga/
version: 1.0.0
citation: "% % ACKNOWLEDGEMENTS\n% % From:… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/manga.Manga-Encylopedia
📚 Manga Encyclopedia Dataset (ChatML) ✨
This dataset is a comprehensive collection of conversational pairs designed to train an AI model (via LoRA or Full Fine-tuning) to become an expert on manga. It covers over 48,000 unique manga titles with summaries, tags, cover URLs, and cross-manga comparisons.
🚀 Dataset Features
📖 Direct Summaries: Detailed information about individual manga titles.
💡 Smart Recommendations: Responses based on specific genre/theme… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Manga-Encylopedia.hausaBERTdatatrainMangaVQA
MangaVQA
Dataset Details
This is the MangaVQA benchmark, designed to evaluate performance under realistic conditions for manga understanding.
This dataset includes 526 manually created question-answer pairs based on images from Manga109.
malayalam_characters
Malayalam Character Dataset
This dataset contains 136,344 images of Malayalam characters, covering 662 distinct classes.
Dataset Structure
The dataset is organized as a single split (train) with the following columns:
image_path: The character image (128x128 grayscale).
text: The Malayalam character/word represented in the image.
label_id: Integer label ID (0-661).
filename: Original filename.
Classes
The dataset includes:
Basic consonants and vowels… See the full description on the dataset page: https://huggingface.co/datasets/mangalathkedar/malayalam_characters.manga-querymanga_line_generation
Manga Line Generation dataset
Converted from https://github.com/1never/MangaLineGeneration.
Paper: https://aclanthology.org/2023.paclic-1.34.pdf
manga109s-line-annotations
Manga109-s Text Line Annotations
High-precision, line-level bounding box and polygon annotations for the Manga109-s Dataset, supporting both full manga pages and speech bubble crops. Furigana is not labeled and is almost entirely excluded from line labels. Includes 8-point oriented polygons for slanted/rotated text lines. The annotation process is documented in METHODOLOGY.md (WIP).
Notice: This dataset contains zero dialogue text and zero images. It requires your own local… See the full description on the dataset page: https://huggingface.co/datasets/bluolightning/manga109s-line-annotations.media-metadata-jikan-manga
TigreGotico/media-metadata-jikan-manga
Rich entity dataset scraped by metadatarr
scraper jikan_manga.
Rows: 83,790
Fields
mal_id
title
title_english
title_japanese
aliases
type
status
chapters
volumes
published_from
published_to
authors
serializations
genres
themes
demographics
score
scored_by
rank
popularity
members
synopsis
background
approved
Source
Generated by scrapers/jikan_manga.py. See the metadatarr repo for the full
pipeline and scraper… See the full description on the dataset page: https://huggingface.co/datasets/LeData/media-metadata-jikan-manga.One_Piece_Chapter1_Manga_En_Zh_Jp
