datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zen-imagezen-multi-imagezendo-synthetic-data
Zendo Synthetic Visual Reasoning Dataset
Synthetic Zendo-style scenes with associated rules and per-scene tensor
representations. Each scene either follows ("positive", label=1) or violates
("negative", label=0) a rule that is given in natural language and as a Prolog
query.
Splits
split
scenes
train
56475
test
3344
rules total
3439
Layout
images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.zenshuu
Bangumi Image Base of Zenshuu.
This is the image base of bangumi Zenshuu., we detected 80 characters, 4266 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/zenshuu.GarmentCodeVTONDataset
GarmentCodeVTONDataset
Simulated garment try-on renders (GarmentCodeV2).
Layout
female/ - female bodies, all outfits (incl. skirts/dresses)
male/ - male bodies, pants outfits only (no skirts/dresses)
Ref/ - per-unit reference renders (19 garment units)
danbooru2023
[Mirror]Danbooru2023: A Large-Scale Crowdsourced and Tagged Anime Illustration Dataset
Danbooru2023 is an extension of Danbooru2021, featuring over 6.8 million anime-style images, totaling more than 8.3 TB.
Each image is accompanied by community-contributed tags that provide detailed descriptions of its content, including characters,
artists, copyright information, concepts, and attire.
This makes it a crucial resource for stylized computer vision tasks and transfer learning.… See the full description on the dataset page: https://huggingface.co/datasets/zenless-archive/danbooru2023.zenodo-second-hand-fashion-v3
Second-Hand Fashion Dataset — wide (one row per garment)
Repack of Zenodo record 10.5281/zenodo.13788681 (Nauman et al., RISE + Wargön Innovation + Myrorna, CC-BY-4.0) into a one-row-per-garment wide layout so the HF dataset viewer shows every attribute — three images plus 25 metadata columns — on a single row.
Previous v3 releases stored one row per (garment, view) with satellite tables that had to be joined manually. That layout is preserved in git history if you need it; the… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/zenodo-second-hand-fashion-v3.zenodo-presentations-openVideoArgusBench
VideoArgusBench
Sample-specific rubric benchmark for conditioned video generation.
VideoArgusBench is the evaluation benchmark for VideoArgus, a framework that scores a generated
video against a rubric written for that specific prompt rather than a fixed global metric. This
dataset ships the inputs (conditioning assets + prompts) and, for each input, a rubric. It does
not contain generated videos — you bring your own model's outputs and score them with the VideoArgus
evaluation… See the full description on the dataset page: https://huggingface.co/datasets/zengziyun/VideoArgusBench.ccpd-subset-30kFireSmokeDatasetNegima-manga-reference-chapters
Negima Manga Reference Chapters (Public)
Personal reference dataset for AI image & video generation (Artlist Seedance 2.0, Kling, LoRAs, etc.).
Mahou Sensei Negima! (Negima!) by Ken Akamatsu
UQ Holder! (sequel series) — Chapters added
Perfect for consistent characters (Asuna, Setsuna, Konoka, Touta, Kirie, etc.), magic circles, pactio cards, Ensis Exsequens slashes, wind blades, immortal fights, and Ken Akamatsu art style.
english versions coming soon
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/Zentoria/Negima-manga-reference-chapters.ccpd-100k-yoloFittingEffectDatasettemprordccpd-ocr-recognitionflux2-consistent-pilot
FLUX.2 AI-consistent control set (200 images) — PRIVATE
An AI-generated-but-consistent control to disentangle "detecting AI generation
artifacts" from "detecting claim-image inconsistency". For each TARA test image we
apply, via the FLUX.2-klein-9B reference pipeline (seed 42, 28 steps), a benign,
context-appropriate edit that does NOT break consistency with the claim — the same
generator and edit magnitude as the inconsistent set, the only difference being whether
the edit… See the full description on the dataset page: https://huggingface.co/datasets/zengrh3/flux2-consistent-pilot.indian-number-platefetal-planes-classification-dataset-zenodofetal-planes-classification-dataset-zenodo
Fetal_Planes_DB
Burgos-Artizzu, X.P., Coronado-Gutiérrez, D., Valenzuela-Alcaraz, B. et al. Evaluation of deep convolutional neural networks for automatic classification of common maternal fetal ultrasound planes. Sci Rep 10, 10200 (2020). https://doi.org/10.1038/s41598-020-67076-5
Data Description
A large dataset of routinely acquired maternal-fetal screening ultrasound images collected from two different hospitals by several operators and ultrasound machines.… See the full description on the dataset page: https://huggingface.co/datasets/ERO26/fetal-planes-classification-dataset-zenodo.zenlesszonezero_chibilarge-scale-multimodal-multilingual-summarization-datasetPlease cite this paper if you use our code or data:
@inproceedings{verma-etal-2023-large,
title = "Large Scale Multi-Lingual Multi-Modal Summarization Dataset",
author = "Verma, Yash and
Jangra, Anubhav and
Verma, Raghvendra and
Saha, Sriparna",
booktitle = "Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics",
month = may,
year = "2023",
address = "Dubrovnik, Croatia",
publisher =… See the full description on the dataset page: https://huggingface.co/datasets/Zenquiorra/large-scale-multimodal-multilingual-summarization-dataset.photo-concept-bucket-wdsfinancial-layout-analysis
📑 Financial Document Layout Analysis (Synthetic)
This dataset contains synthetically generated scanned financial reports intended for Document Layout Analysis. This data helps in training AI models for object detection (such as identifying tables, stamps, and figures) without compromising the confidentiality of real financial information.
📂 Label Classes
The dataset supports 6 main object classes:
0: Title (Section headers)
1: Text (Standard text paragraphs)
2: Table… See the full description on the dataset page: https://huggingface.co/datasets/Zenng2812/financial-layout-analysis.zenodo-presentations-opendanbooru2023-metazerochan-wip-testmaluschatvqa-vietnamese-chartsjujutsu-kaisen-maki-zenin
Dataset Card for "jujutsu-kaisen-maki-zenin"
More Information needed
