datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
manga-colorization-datasetcleaned_manga_dataset
Licensing and source material
This repository contains derived/processed images from publicly accessible manga
sources (GANMA! and Comic Walker / カドコミ). The original works are copyrighted by
their respective rights holders. No ownership of the original artwork is claimed.
The CC BY-NC 4.0 license applies only to the annotations and dataset metadata
created by the dataset authors, not to the underlying original artwork.
Unpacking the dataset
The dataset ships as a… See the full description on the dataset page: https://huggingface.co/datasets/Vasyanator2/cleaned_manga_dataset.mangaManga-Drawings
Dataset Card: MangaDF Text-to-Image Prompts Dataset
Dataset Description
Image Source: Images generated by alvdansen/BandW-Manga weights on ChanY/Stable-Flash-Lighting modelPrompt Source: ChatGPT
Overview
The MangaDF Text-to-Image Prompts Dataset is a collection of text prompts paired with corresponding images. The images in this dataset were generated using the alvdansen/BandW-Manga weights applied to the ChanY/Stable-Flash-Lighting diffusion model. This… See the full description on the dataset page: https://huggingface.co/datasets/Chan-Y/Manga-Drawings.MangaSegmentationmangakasantoassistantsanto
Bangumi Image Base of Mangaka-san To Assistant-san To
This is the image base of bangumi Mangaka-san to Assistant-san to, we detected 10 characters, 3298 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/mangakasantoassistantsanto.manga-syntheticNegima-manga-reference-chapters
Negima Manga Reference Chapters (Public)
Personal reference dataset for AI image & video generation (Artlist Seedance 2.0, Kling, LoRAs, etc.).
Mahou Sensei Negima! (Negima!) by Ken Akamatsu
UQ Holder! (sequel series) — Chapters added
Perfect for consistent characters (Asuna, Setsuna, Konoka, Touta, Kirie, etc.), magic circles, pactio cards, Ensis Exsequens slashes, wind blades, immortal fights, and Ken Akamatsu art style.
english versions coming soon
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/Zentoria/Negima-manga-reference-chapters.mangasMangaImagesDatasetJapanese_Mangamanga-colorization-datasetmanga-covers-text-detection
Manga Covers Text Detection
Manually annotated text boxes for manga covers text detection created by @JustANormalTinkerer.
This dataset contains 100 cover images and 526 manually annotated text rectangles. The image column is a Hugging Face image feature containing the original image bytes. The boxes column contains normalized pixel-coordinate bounding boxes in [x_min, y_min, x_max, y_max] format, the original points, label, and shape metadata. annotation_json preserves the… See the full description on the dataset page: https://huggingface.co/datasets/petersunde/manga-covers-text-detection.MangaliCa_EN-HU
MangaliCa Bilingual Image–Caption Dataset (Hungarian–English)
Dataset Description
The MangaliCa Bilingual Image–Caption Dataset is a large-scale Hungarian–English multimodal dataset containing approximately 70 million aligned image–caption pairs.
Each image is paired with:
an original English caption
a machine-translated Hungarian caption
This dataset was created to address the lack of large-scale multimodal resources for Hungarian, enabling bilingual… See the full description on the dataset page: https://huggingface.co/datasets/Obscure-Entropy/MangaliCa_EN-HU.manga-colorization-mastercomics_dataset_512_inv_manga_correctedmangaba-adsmalayalam_characters
Malayalam Character Dataset
This dataset contains 136,344 images of Malayalam characters, covering 662 distinct classes.
Dataset Structure
The dataset is organized as a single split (train) with the following columns:
image_path: The character image (128x128 grayscale).
text: The Malayalam character/word represented in the image.
label_id: Integer label ID (0-661).
filename: Original filename.
Classes
The dataset includes:
Basic consonants and vowels… See the full description on the dataset page: https://huggingface.co/datasets/mangalathkedar/malayalam_characters.manga-querymanga_line_generation
Manga Line Generation dataset
Converted from https://github.com/1never/MangaLineGeneration.
Paper: https://aclanthology.org/2023.paclic-1.34.pdf
manga_libro_clasOne_Piece_Chapter1_Manga_En_Zh_JpMangaColoringManga109Pose_HAmalayalam_char_consolidatedsenko-mangamanga-datasetManga-Popularis-EN
Dataset Card: Manga-Popularis-EN
Dataset Details
Dataset Description
This dataset contains information about art books, including titles, authors, descriptions, image paths, prices, publication dates, publishers, ISBN numbers, and sizes.
License: GPL-3.0
Dataset Sources
The dataset was compiled from various sources, including online bookstores, publisher websites, and catalogs. On date 4/21/24.
Dataset Structure
The dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/Tsunnami/Manga-Popularis-EN.manga-colorization-testGenshin-Impact-Official-Manga-EN-US
