datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MangaZero
Introduction
MangaZero dataset from paper DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
Please see GitHub repo to get the usage
Project page: https://jianzongwu.github.io/projects/diffsensei
The "type" key in character annotation
Type = 0: Characters with clean faces
Type = 1: Characters with vague faces
Type = 2: Exaggerated, artistic characters with appearances that don't conform to typical standards.
prompt-injection-multilayerMangaSegmentation
Advancing Manga Analysis: Comprehensive Segmentation Annotations for the Manga109 Dataset
License
Please check the LICENSE file for more details.
All images in the segmentation annotations are owned and copyrighted by Minshan Xie.
You are automatically granted permission to use the images for academic and commercial
usages, provided that the image credit "Copyrighted by Minshan Xie" is included in
any form of publication, reproduction, redistribution, or derivatives of… See the full description on the dataset page: https://huggingface.co/datasets/MS92/MangaSegmentation.Japanese_Mangamanga-colorization-masterreforma-tributaria-instructManga-Encylopedia
📚 Manga Encyclopedia Dataset (ChatML) ✨
This dataset is a comprehensive collection of conversational pairs designed to train an AI model (via LoRA or Full Fine-tuning) to become an expert on manga. It covers over 48,000 unique manga titles with summaries, tags, cover URLs, and cross-manga comparisons.
🚀 Dataset Features
📖 Direct Summaries: Detailed information about individual manga titles.
💡 Smart Recommendations: Responses based on specific genre/theme… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Manga-Encylopedia.manga109s-line-annotations
Manga109-s Text Line Annotations
High-precision, line-level bounding box and polygon annotations for the Manga109-s Dataset, supporting both full manga pages and speech bubble crops. Furigana is not labeled and is almost entirely excluded from line labels. Includes 8-point oriented polygons for slanted/rotated text lines. The annotation process is documented in METHODOLOGY.md (WIP).
Notice: This dataset contains zero dialogue text and zero images. It requires your own local… See the full description on the dataset page: https://huggingface.co/datasets/bluolightning/manga109s-line-annotations.manga_chatbot_questionsbbc-news
BBC News Topic Dataset
Dataset on BBC News Topic Classification consisting of 2,225 articles published on the BBC News website corresponding during 2004-2005. Each article is labeled under one of 5 categories: business, entertainment, politics, sport or tech.
Original source for this dataset:
Derek Greene, Pádraig Cunningham, “Practical Solutions to the Problem of Diagonal Dominance in Kernel Document Clustering,” in Proc. 23rd International Conference on Machine learning… See the full description on the dataset page: https://huggingface.co/datasets/mangara/bbc-news.Manga_ita_animeclickList of mangas (snapshot 28.02.2025) listed on Animeclick. The dataset contains the following details:
{
"url": "https://www.animeclick.it/manga/46610/restart-rui-si-yi-kai",
"titolo_originale": "Gui Ling",
"titolo_inglese": "Zona 0: Restart / Return To Zero",
"titolo_kanji": "归零",
"nazionalita": "Cina",
"casa_editrice": "KuaiKan Manhua",
"storia": "Rui Si",
"disegni": "Rui Si",
"anno": "2020",
"stato_patria": "completato",
"stato_italia": "in corso"… See the full description on the dataset page: https://huggingface.co/datasets/WasamiKirua/Manga_ita_animeclick.manga_genres_synopsispop-test_manga6pop-test_mangamanga_prompt_to_synopsis
