datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Anime-Art-Multicaptions-v5.0
A large (over 6M) set of various high quality captions for anime/game artworks.
Key features
Features original character names
Only top tier vlms have been used such as Claude, GPT, Gemini
High accuracy and lots of details for the majority
Wide range from simple and innocent to darkest NSFW
Structured formats for versatile tasks
Most complex cases for multiple characters (about 20%) have been rechecked via 2nd pass and corrected
Caption formats and usage
The… See the full description on the dataset page: https://huggingface.co/datasets/Minthy/Anime-Art-Multicaptions-v5.0.Multi-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.flickr-megalith-10m-internvl2-multi-caption
Dataset Card for flickr-megalith-10m-internvl2-multi-caption
Dataset Summary
This is approximately 57.3 million synthetic captions for the images found in madebyollin/megalith-10m.
It includes the following captions:
InternVL2 8B long captions (by CaptionEmporium)
InternVL2 8B short captions (by CaptionEmporium)
Florence2 long captions (by aipicasso)
Florence2 short captions (by CaptionEmporium)
ShareCaptioner long captions (by drawthingsai)
ShareCaptioner short… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/flickr-megalith-10m-internvl2-multi-caption.coco-stuff-captioned-multi
Dataset Card for "coco-stuff-captioned-multi"
More Information needed
BG20K-MulticaptionMultiCaptions-LARGE
MultiCaptions-LARGE
MultiCaptions-LARGE is a lightweight, multi-response image captioning dataset containing 20,592 image-caption pairs, where each image is associated with two complementary captions generated by different vision-language models.
response_1 contains a dense, long-form caption generated using the Qwen3.5 Multimodal model, providing detailed descriptions of scene composition, objects, attributes, spatial relationships, and visual context.
response_2 contains a… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/MultiCaptions-LARGE.coco_captions_1107
Dataset Card for "coco_captions_1107"
More Information needed
multi-view_caption
Multi-View Caption
Per-segment captions for multi-view human (DNA-Rendering, ActorsHQ) and animal
(Artemis / DFA) video datasets. Captions are generated from masked multi-view
composites with Gemini 3 Flash and follow the ActivityNet-style dense
video-captioning layout: one row per video, with parallel captions and
timestamps lists plus a representative thumbnail.
Citation
If you find our caption useful, please cite our paper
Flex4DHuman: Flexible Multi-view Video… See the full description on the dataset page: https://huggingface.co/datasets/andaba/multi-view_caption.flickr30k_captions_1107
Dataset Card for "flickr30k_captions_1107"
More Information needed
BIGstockimage2M-multicaptionanime_multicaptions_v1.1
Varous NL captions for anime pictures
Over 1.7M natural language captions for anime pictures from danbooru, nozomi and other sources.
Made with Claude, Gemini, GPT. Most contain character names, pictures are balanced across different criteria, concepts, characters, etc.
Varoius columns for different caption styles. Pruned is a shortest version made from structured/long with llm. Small/cheap models have been used, for best results consider to perform your own processing from initial… See the full description on the dataset page: https://huggingface.co/datasets/Minthy/anime_multicaptions_v1.1.anime-multicaption-v5-short
Anime character interaction
NLP tools were used to clean up the text.
(Japanese) character names were removed, because BERT models have limited vocabulary.
Discarded rows that included unique words (in the global vocabulary pool).
The captions emphasize the scene and the interactions between the characters. Unlike most LLM-generated captions, these could be written by hand.
References
Captions for anime/game artworks
anime_multicaptions_v1
Varous NL captions for anime pictures
Over 1.5M natural language captions for anime pictures from danbooru, nozomi and other sources. Made with Claude, Gemini, GPT4o. Most contain character names, pictures are balanced across different aspects, concepts, characters, etc.
Dataframe index is md5 for Danbooru and others or numeric id for Nozomi. Varoius columns for different caption styles. "pruned" is result or rewriting from structured/long with llm, for better results consider to… See the full description on the dataset page: https://huggingface.co/datasets/Minthy/anime_multicaptions_v1.civit_v2_multi_captions_500k
