datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HBGSDG-30K
SDG-30K — Structured Defect Grounding Dataset
A 30,000-image dataset for structured defect grounding in text-to-image
generations. Each image is annotated with bounding-box-level defects, where
each defect carries:
a category (artifact for visual flaws / misalignment for caption-image
mismatches),
a natural-language description, and
a chain-of-thought reasoning trace.
This is the public release accompanying the NeurIPS 2026 anonymous submission
"SDG: Structured Defect… See the full description on the dataset page: https://huggingface.co/datasets/P1n3/SDG-30K.danbooru-2024Insectsvisually-dependent-ambiguity
VIDA: Visually-Dependent Ambiguity for Multimodal MT
VIDA is an English-Chinese multimodal machine translation dataset for visual ambiguity resolution.Each instance contains an English source sentence, its paired image, and Chinese references that resolve annotated ambiguity spans using visual evidence.
Paper: A Multimodal Dataset for Visually Grounded Ambiguity in Machine Translation
Dataset composition
This release contains four splits:
Split
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/p1k0/visually-dependent-ambiguity.p316-p147-p217-indus-branch-audit-20260609
P316/P147/P217 Indus Branch Audit Article Assets
This dataset repository hosts the Hugging Face-only assets for the article draft
What We Can Say About P316/P147/P217.
The package includes generated article images, QA summaries, route evidence
tables, and memo files. The strict result remains NOT LICENSED; the speculative
functional lane remains confidence 0.36.
Current QA
Checked artifacts: 428
All checks pass: true
False blockers: 0
M-959 ICIT help-layout… See the full description on the dataset page: https://huggingface.co/datasets/junafinity/p316-p147-p217-indus-branch-audit-20260609.pvc
PVC figure products dataset
This dataset contains product information of figure images scraped from multiple Web sites.
Dataset information
Subset
Source
Size
goodsmile-figma
https://www.goodsmile.info/ja/products/category/figma/announced/2023
947
goodsmile-nendoroid
https://www.goodsmile.info/ja/products/category/nendoroid_series/announced/2023
3378
goodsmile-scale
https://www.goodsmile.info/ja/products/category/scale/announced/2023
2203
kotobukiya… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/pvc.military-labeled-clip
Military-Labeled CLIP Crops (DVIDS sourced)
Per-object crops extracted from DVIDS military imagery, each accompanied by a
Gemini-VLM caption suitable for CLIP fine-tuning or zero-shot evaluation.
Files
crops/dvids_image_{id}_{class}_{idx}.jpg — 4,844 cropped objects
captions.jsonl — per-crop metadata: {crop_path, image_id, class_name, class_id, bbox, caption, image_caption, branch, source, source_url}
Classes (12) — Distribution
Class
Crops… See the full description on the dataset page: https://huggingface.co/datasets/p14ton/military-labeled-clip.FractalDB-1k
FractalDB 1k
FractalDB 1k dataset from Pre-training without Natural Images.
Original repo | Project page | arXiv
Citing
@article{KataokaIJCV2022,
author={Kataoka, Hirokatsu and Okayasu, Kazushige and Matsumoto, Asato and Yamagata, Eisuke and Yamada, Ryosuke and Inoue, Nakamasa and Nakamura, Akio and Satoh, Yutaka},
title={Pre-training without Natural Images},
article={International Journal on Computer Vision (IJCV)},
year={2022},
}… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/FractalDB-1k.niji-v5私がnijijourney v5で生成した画像。自由に使えます。(けど詐欺とか犯罪につかうのはやめてね)
おすすめの使い方としては、とりあえず中の画像を見てみて好きなものだけ選んで使うとよいと思います。
全体の注意点として、必ずしもすべての画像にキャプションが付属してるとは限らないのと、キャプションがついていてもそのまま使うと問題が発生する可能性があるのであまり信用しないでください。
また人為的なミスにより、4分割されずに結合している画像があったり、過度に分割されている画像があったりするので注意してください。
vol1
だいたい2000枚くらいで、多分全部デフォルトスタイルのものです。
解答すると中にLAION Aesthetic v2のスコアでいくつかのフォルダに分類されています。aesthetic_50 ならスコア0.5以上のものです。not_aesthetic は0.5未満のものです。
ただし、exceptional フォルダはチェリーピックした画像が入っており、aesthetic_xx の中のものと重複します。exclude… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/niji-v5.tt-nerf-p150-whiteOracle-P15K
Mitigating Long-tail Distribution in Oracle Bone Inscriptions: Dataset, Model, and Benchmark ☯️
The first attempt to apply diffusion model in realistic and controllable OBI generation
Jinhao Li1*,
Zijian Chen2*,
Runze Jiang3,
Tingzhu Chen3†,
Changbo Wang1†,
Guangtao Zhai2
1School of Computer Science and Technology, East China Normal University
2Institute of Image Communication and Information Processing, Shanghai Jiao Tong University… See the full description on the dataset page: https://huggingface.co/datasets/lomljhoax/Oracle-P15K.adl-p13-t2-armsadl-p16-arms
adl-p16-arms
Phase 16e S5 T2 tranche 2 append: 121 new capital_structure snapshot pages (new issuers beyond T1 + T2-t1). Jobs read this manifest only.
FractalDB-60
FractalDB 60
FractalDB 60 dataset from Pre-training without Natural Images.
Original repo | Project page | arXiv
Citing
@article{KataokaIJCV2022,
author={Kataoka, Hirokatsu and Okayasu, Kazushige and Matsumoto, Asato and Yamagata, Eisuke and Yamada, Ryosuke and Inoue, Nakamasa and Nakamura, Akio and Satoh, Yutaka},
title={Pre-training without Natural Images},
article={International Journal on Computer Vision (IJCV)},
year={2022},
}… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/FractalDB-60.manga_line_generation
Manga Line Generation dataset
Converted from https://github.com/1never/MangaLineGeneration.
Paper: https://aclanthology.org/2023.paclic-1.34.pdf
OCRMT30K-refine下载数据集使用:
huggingface-cli download --repo-type dataset --resume-download p1k0/OCRMT30K-refine --local-dir OCRMT30K-refine
original_data:原始标注
whole_image_v2.zip: 图片文件
nobodies
Nobodies
AI-generated human image dataset.
Contents
Face
vol1: 32 photos of women's faces. Generated with WD1.5 beta 2.
Sample:
Portrait
vol1: 31 photos of women's portraits. Generated with WD1.5 beta 2 and the fashion LoCon.
Sample:
vol2: 165 photos of woman's portraits. Generated with WD1.5 beta 2 and the fashion LoCon. Classified with LAION Aesthetic v2.75 hair bun photos
90 medium hair photos
nijijourneyOracle-P15K
Mitigating Long-tail Distribution in Oracle Bone Inscriptions: Dataset, Model, and Benchmark ☯️
The first attempt to apply diffusion model in realistic and controllable OBI generation
Jinhao Li1*,
Zijian Chen2*,
Runze Jiang3,
Tingzhu Chen3†,
Changbo Wang1†,
Guangtao Zhai2
1School of Computer Science and Technology, East China Normal University
2Institute of Image Communication and Information Processing, Shanghai Jiao Tong University… See the full description on the dataset page: https://huggingface.co/datasets/ausirzzn/Oracle-P15K.college-texts-annas-archive-v1
Dataset Card for "college-texts-annas-archive-v1"
More Information needed
LSDIR_colorize_pixels_256_p16LSDIR_colorize_walloc_x16_256_p16DeathSe46_p1_newasyouwere-ocr-p1NBA_StateFarm_splits_p1pexels-1k-randomrandomly sampled 1k image and caption pairs from CaptionEmporium/pexels-568k-internvl2
LSDIR_colorize_walloc_x4_256_p163am-refinebadquality
