datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
json-coco-format
JSON COCO Format — task-differentiated SFT data
A multi-task supervised fine-tuning dataset that teaches a model to convert
image-synthesis caption prompts into JSON whose structure varies by task.
Built from MS-COCO captions (Karpathy split) with Claude Sonnet 4.6 as the
teacher; designed for training per-task LoRAs on
Qwen/Qwen3.5-0.8B.
Each row is in the Qwen3.5-native tool-call shape: a messages array with an
assistant turn whose tool_calls[0].function.arguments is a dict… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/json-coco-format.coco-karpathy-opus-de
Dataset Card for MS COCO Karpathy in German language
This dataset contains captions that were machine translated using opus-mt-en-de.
Dataset Details
Dataset Sources
The processed MS COCO datasets (Karpathy Split) in this repo are based on the following sources:
Type
MD5
URL
Train
aa31ac474cf6250ebb81d18348a07ed8
https://storage.googleapis.com/sfr-vision-language-research/datasets/coco_karpathy_train.json
Validation
b273847456ef5580e33713b1f7de52a0… See the full description on the dataset page: https://huggingface.co/datasets/Jotschi/coco-karpathy-opus-de.Epilepsy_Synthetics
Epilepsy_Syntheics
This is a cross-languadge dataset for epilepsy-care, support both madarin and english.
It is generated by Qwen 1.5(For mandarin) and LLAMA-3(For English) with the use of self-instruct method.
This dataset contains 1K+1K epilepsy-care data. And it have already been splitted and cleaned.
Have fun and enjoy!
coco-deceptive-clip-llama3.1-8b
COCO-Deceptive-CLIP-LLaMA-3.1-8B Training Dataset
🏆 This work is accepted to ACL 2025 (Main Conference).
Figure: Attack success rate (ASR) and caption diversity of our model on the COCO dataset, illustrating its ability to generate deceptive captions that successfully fool CLIP.
Dataset Details
This dataset provides instruction–response pairs formatted as short two-turn conversations:
The user message contains:
A given image caption.
A set of task… See the full description on the dataset page: https://huggingface.co/datasets/ahnpersie/coco-deceptive-clip-llama3.1-8b.Filtered-COCO-Captions
Dataset Summary
This dataset is derived from the MS COCO caption annotations.
Source
Original annotations: MS COCO / COCO Consortium
License
The original annotation set is licensed under CC BY 4.0.
This repository redistributes a filtered/adapted version of the annotation text only.
No original COCO images are included.
Modifications
Removed captions deemed unsuitable for TOEIC educational materials
Normalized punctuation and whitespace
Filtered for… See the full description on the dataset page: https://huggingface.co/datasets/kknono668/Filtered-COCO-Captions.LLaVA-Instruct-21K-COCO-SubSet
subset from https://huggingface.co/datasets/liuhaotian/LLaVA-Instruct-150K
train: 21000
val seen: 3000
val unseen: 2100
test: 6000
