datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coconot
🥥 CoCoNot: Contextually, Comply Not! Dataset Card
Dataset Details
Dataset Description
Chat-based language models are designed to be helpful, yet they should not comply with every user request.
While most existing work primarily focuses on refusal of "unsafe" queries, we posit that the scope of noncompliance should be broadened. We introduce a comprehensive taxonomy of contextual noncompliance describing when and how models should not comply with user… See the full description on the dataset page: https://huggingface.co/datasets/allenai/coconot.cocoterosCOCOTEROS Dataset V1.1
Dataset Summary: The COCOTEROS dataset is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context which serves as the co-text of the keywords provided. This makes COCOTEROS suitable for tasks where the generated text needs to be related both to a set of specific… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros.coco-captions-pt-br
🎉 COCO Captions Dataset Translation for Portuguese Image Captioning
💾 Dataset Summary
COCO Captions Portuguese Translation, a multimodal dataset for Portuguese image captioning with 123,287 images, each accompanied by five descriptive captions that have been
generated by human annotators for every individual image. The original English captions were rendered into Portuguese
through the utilization of the Google Translator API.
🧑💻 Hot to Get… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/coco-captions-pt-br.json-coco-format
JSON COCO Format — task-differentiated SFT data
A multi-task supervised fine-tuning dataset that teaches a model to convert
image-synthesis caption prompts into JSON whose structure varies by task.
Built from MS-COCO captions (Karpathy split) with Claude Sonnet 4.6 as the
teacher; designed for training per-task LoRAs on
Qwen/Qwen3.5-0.8B.
Each row is in the Qwen3.5-native tool-call shape: a messages array with an
assistant turn whose tool_calls[0].function.arguments is a dict… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/json-coco-format.coco-karpathy-opus-de
Dataset Card for MS COCO Karpathy in German language
This dataset contains captions that were machine translated using opus-mt-en-de.
Dataset Details
Dataset Sources
The processed MS COCO datasets (Karpathy Split) in this repo are based on the following sources:
Type
MD5
URL
Train
aa31ac474cf6250ebb81d18348a07ed8
https://storage.googleapis.com/sfr-vision-language-research/datasets/coco_karpathy_train.json
Validation
b273847456ef5580e33713b1f7de52a0… See the full description on the dataset page: https://huggingface.co/datasets/Jotschi/coco-karpathy-opus-de.Epilepsy_Synthetics
Epilepsy_Syntheics
This is a cross-languadge dataset for epilepsy-care, support both madarin and english.
It is generated by Qwen 1.5(For mandarin) and LLAMA-3(For English) with the use of self-instruct method.
This dataset contains 1K+1K epilepsy-care data. And it have already been splitted and cleaned.
Have fun and enjoy!
cocoteros_vaCOCOTEROS_VA Dataset
Dataset Summary:
The COCOTEROS_VA dataset is a translation of the COCOTEROS dataset, carried out by a linguist specialized in Valencian. It is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context, which serves as the co-text of the keywords provided. This makes… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros_va.coco-deceptive-clip-llama3.1-8b
COCO-Deceptive-CLIP-LLaMA-3.1-8B Training Dataset
🏆 This work is accepted to ACL 2025 (Main Conference).
Figure: Attack success rate (ASR) and caption diversity of our model on the COCO dataset, illustrating its ability to generate deceptive captions that successfully fool CLIP.
Dataset Details
This dataset provides instruction–response pairs formatted as short two-turn conversations:
The user message contains:
A given image caption.
A set of task… See the full description on the dataset page: https://huggingface.co/datasets/ahnpersie/coco-deceptive-clip-llama3.1-8b.Filtered-COCO-Captions
Dataset Summary
This dataset is derived from the MS COCO caption annotations.
Source
Original annotations: MS COCO / COCO Consortium
License
The original annotation set is licensed under CC BY 4.0.
This repository redistributes a filtered/adapted version of the annotation text only.
No original COCO images are included.
Modifications
Removed captions deemed unsuitable for TOEIC educational materials
Normalized punctuation and whitespace
Filtered for… See the full description on the dataset page: https://huggingface.co/datasets/kknono668/Filtered-COCO-Captions.LLaVA-Instruct-21K-COCO-SubSet
subset from https://huggingface.co/datasets/liuhaotian/LLaVA-Instruct-150K
train: 21000
val seen: 3000
val unseen: 2100
test: 6000
coco-captions-pt-br
🎉 COCO Captions Dataset Translation for Portuguese Image Captioning
💾 Dataset Summary
COCO Captions Portuguese Translation, a multimodal dataset for Portuguese image captioning with 123,287 images, each accompanied by five descriptive captions that have been
generated by human annotators for every individual image. The original English captions were rendered into Portuguese
through the utilization of the Google Translator API.
🧑💻 Hot to Get… See the full description on the dataset page: https://huggingface.co/datasets/EliMC/coco-captions-pt-br.coco-captions_marathi
Coco-Captions Marathi Dataset: High-Quality Marathi NLP Corpus
📌 Overview
The Coco-Captions Marathi dataset is a meticulously curated collection of 414010 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness.
This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource for machine learning… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/coco-captions_marathi.
