datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Filtered-COCO-Captions
Dataset Summary
This dataset is derived from the MS COCO caption annotations.
Source
Original annotations: MS COCO / COCO Consortium
License
The original annotation set is licensed under CC BY 4.0.
This repository redistributes a filtered/adapted version of the annotation text only.
No original COCO images are included.
Modifications
Removed captions deemed unsuitable for TOEIC educational materials
Normalized punctuation and whitespace
Filtered for… See the full description on the dataset page: https://huggingface.co/datasets/kknono668/Filtered-COCO-Captions.picture_short_captionIt is used for training to generate short sentence copywriting according to image content, the source of the image dataset is https://unsplash.com/, and the source of short sentence copywriting is Claude3.7
用做图片内容生成短句文案训练,图片数据集来自 https://unsplash.com/,短句文案来自 Claude3.7 模型
