Zicara/manga2M
manga exhentai site majorly lang EN, JP, from all (nyaa. si) torrents-exclude huge 300GB+ UBUCA, Exhentai Mothercon Archive with no seeders. Convert to .webp and approximately keep only manga with text using DB_TD500_resnet50; Intended use for bubble box+OCR training model on 2.2M images. If proced to label OCR of all image, it would take two months... April 2026: Added ~150 GB in folder [猫の仓库] 汉化本合集(上半)[3037本] of 243,703 images (WebP q85) across 2332 manga, include no-text images, seed… See the full description on the dataset page: https://huggingface.co/datasets/Zicara/manga2M.
Create manifest.json
Update manifest.json
Add manifest.json
Update README.md
Add "[猫の仓库] 汉化本合集(上半)[3037本]" seed.
Add [猫の仓库] 汉化本合集(上半)[3037本] 2
Add [猫の仓库] 汉化本合集(上半)[3037本] 1
Update README.md
Update README.md
Update README.md
Add manga_webp-text 2.2M
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Create README.md
initial commit
