multi-caption
Anime-Art-Multicaptions-v5.0
A large (over 6M) set of various high quality captions for anime/game artworks.
Key features
Features original character names
Only top tier vlms have been used such as Claude, GPT, Gemini
High accuracy and lots of details for the majority
Wide range from simple and innocent to darkest NSFW
Structured formats for versatile tasks
Most complex cases for multiple characters (about 20%) have been rechecked via 2nd pass and corrected
Caption formats and usage
The… See the full description on the dataset page: https://huggingface.co/datasets/Minthy/Anime-Art-Multicaptions-v5.0.Multi-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.flickr-megalith-10m-internvl2-multi-caption
Dataset Card for flickr-megalith-10m-internvl2-multi-caption
Dataset Summary
This is approximately 57.3 million synthetic captions for the images found in madebyollin/megalith-10m.
It includes the following captions:
InternVL2 8B long captions (by CaptionEmporium)
InternVL2 8B short captions (by CaptionEmporium)
Florence2 long captions (by aipicasso)
Florence2 short captions (by CaptionEmporium)
ShareCaptioner long captions (by drawthingsai)
ShareCaptioner short… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/flickr-megalith-10m-internvl2-multi-caption.coco-stuff-captioned-multi
Dataset Card for "coco-stuff-captioned-multi"
More Information needed
BG20K-MulticaptionMultiCaptions-LARGE
MultiCaptions-LARGE
MultiCaptions-LARGE is a lightweight, multi-response image captioning dataset containing 20,592 image-caption pairs, where each image is associated with two complementary captions generated by different vision-language models.
response_1 contains a dense, long-form caption generated using the Qwen3.5 Multimodal model, providing detailed descriptions of scene composition, objects, attributes, spatial relationships, and visual context.
response_2 contains a… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/MultiCaptions-LARGE.
