datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers_image_dockaminotoukoubousen
Bangumi Image Base of Kami No Tou: Koubou-sen
This is the image base of bangumi Kami no Tou: Koubou-sen, we detected 110 characters, 9146 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kaminotoukoubousen.kamuicode-i2i-images
AI画像編集モデル 総合ベンチマーク v5.4
概要
本ディレクトリは、AI画像編集モデルの性能を総合的に評価するためのベンチマークスイートです。
47種類のテストを4つの難度レベルに分類し、65種類のモデル(旧45モデル + 新20モデル)を比較評価しています。
評価カバレッジ
項目
旧モデル群 (35)
新モデル群 (20)
合計
モデル数
45 (I2I/R2I 22 + T2I 13 + その他 10)
20
65
テスト数
34
47
47
生成画像
3,446
1,301 / 2,403
4,747+
VLM評価済み
完了
1,301 (54%)
進行中
評価方式
VLM-as-a-Judge: Gemini 2.5 Flash による5軸評価(Structure / Identity / Reasoning / Instruction / Quality)
スタイル多様性: photo, anime_flat, anime_cg… See the full description on the dataset page: https://huggingface.co/datasets/yumenojmd/kamuicode-i2i-images.kamitachinihirowaretaotoko2ndseason
Bangumi Image Base of Kami-tachi Ni Hirowareta Otoko 2nd Season
This is the image base of bangumi Kami-tachi ni Hirowareta Otoko 2nd Season, we detected 75 characters, 4264 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamitachinihirowaretaotoko2ndseason.kaminomizoshirusekai
Bangumi Image Base of Kami Nomi Zo Shiru Sekai
This is the image base of bangumi Kami Nomi zo Shiru Sekai, we detected 60 characters, 5684 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kaminomizoshirusekai.kaminotou
Bangumi Image Base of Kami No Tou
This is the image base of bangumi Kami no Tou, we detected 49 characters, 3102 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kaminotou.kamisamakiss
Bangumi Image Base of Kamisama Kiss
This is the image base of bangumi Kamisama Kiss, we detected 50 characters, 2686 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamisamakiss.kamonohashironnokindansuiri
Bangumi Image Base of Kamonohashi Ron No Kindan Suiri
This is the image base of bangumi Kamonohashi Ron no Kindan Suiri, we detected 36 characters, 4169 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamonohashironnokindansuiri.kamitachinihirowaretaotoko
Bangumi Image Base of Kami-tachi Ni Hirowareta Otoko
This is the image base of bangumi Kami-tachi ni Hirowareta Otoko, we detected 75 characters, 4446 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamitachinihirowaretaotoko.kamiwagameniueteiru
Bangumi Image Base of Kami Wa Game Ni Ueteiru
This is the image base of bangumi Kami wa Game ni Ueteiru, we detected 68 characters, 4275 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamiwagameniueteiru.kamikazekaitoujeanne
Bangumi Image Base of Kamikaze Kaitou Jeanne
This is the image base of bangumi Kamikaze Kaitou Jeanne, we detected 43 characters, 3600 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamikazekaitoujeanne.kamierabi
Bangumi Image Base of Kamierabi
This is the image base of bangumi Kamierabi, we detected 55 characters, 5241 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamierabi.forest-fire-annotations
Forest Fire Detection Dataset — Auto-Annotated
Bounding-box annotated version of touati-kamel/forest-fire-dataset,
built for training forest-fire / smoke / fog object detection models.
Overview
This dataset contains video frames auto-labeled with bounding boxes for fire and
smoke-related visual phenomena, using a zero-shot open-vocabulary object detector
(Grounding DINO). It is derived from the original touati-kamel/forest-fire-dataset image
classification dataset… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/forest-fire-annotations.kaminakisekainokamisamakatsudou
Bangumi Image Base of Kaminaki Sekai No Kamisama Katsudou
This is the image base of bangumi Kaminaki Sekai no Kamisama Katsudou, we detected 72 characters, 5020 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kaminakisekainokamisamakatsudou.pochigenesis_joint_position_2000_20fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 2007,
"total_frames": 423477,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:2007"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kaveh-kamali/genesis_joint_position_2000_20fps.image_captions_x
Dataset Card for image_captions_x (URL + Caption)
This dataset provides a lightweight, web-scale resource of image-caption pairs in the form of URLs and their associated textual descriptions (captions). It is designed for training and evaluating vision-language models where users retrieve images independently from the provided links.
This dataset card is based on the Hugging Face dataset card template.
Dataset Details
Dataset Description
This… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/image_captions_x.synthetic-indic-manuscriptsclip-animals-datakami_no_dobutsu_paper_animals
We love origmai at Takara!!
diagram_image_to_text
Dataset Card for "diagram_image_to_text"
More Information needed
tito
Demo Project
Exported from Tito annotator on 2026-09-05. Rows: 3.
load_dataset("Kamatchu/tito") yields one row per image: an image column plus annotation columns.
Filters: none. Annotation schema: caption(textarea), conversations(conversation).
KamonBench
KamonBench
A grammar-based image-to-structure benchmark for evaluating compositional
factor recovery in vision-language models, built around Japanese family crests
(kamon, 家紋).
Each composite crest is paired with:
a formal kamon description language string (KDL, kamon yōgo, 家紋用語),
a segmented Japanese analysis,
an English translation,
a non-linguistic program code over the generator factors.
Because every crest is synthesized from a known triple of generator factors
(container C… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/KamonBench.ghibli-dataset
Ghibli Real vs AI-Generated Dataset
One sample per line
Includes: id, image, label, description
Use this for standard classification or image-text training
Real images sourced from Nechintosh/ghibli (810 images)
AI-generated images created using:
nitrosocke/Ghibli-Diffusion (2637 images)
KappaNeuro/studio-ghibli-style (810 images)
Note: While the KappaNeuro repository does not explicitly state a license, it is a fine-tuned model based on Stable Diffusion XL, which is… See the full description on the dataset page: https://huggingface.co/datasets/Kamala28/ghibli-dataset.kamagra-stickerssonar_datamaml-pde-datasetsgenesis_ee_position_2000_20fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 2000,
"total_frames": 420000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:2000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kaveh-kamali/genesis_ee_position_2000_20fps.genesis_ee_position_40_20fps_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 40,
"total_frames": 8440,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kaveh-kamali/genesis_ee_position_40_20fps_test.slm-sign-dataset
