CoolFace
2 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CodedotAI /code_clippyThis dataset was generated by selecting GitHub repositories from a large collection of repositories. These repositories were collected from https://seart-ghs.si.usi.ch/ and Github portion of [The Pile](https://github.com/EleutherAI/github-downloader) (performed on July 7th, 2021). The goal of this dataset is to provide a training set for pretraining large language models on code data for helping software engineering researchers better understand their impacts on software related tasks such as autocompletion of code. The dataset is split into train, validation, and test splits. There is a version containing duplicates (209GBs compressed) and ones where exact duplicates (132GBs compressed) are removed. Contains mostly JavaScript and Python code, but other programming languages are included as well to various degrees.text-generation12 likes184 downloads4y agoHugging Face02ahnpersie /coco-deceptive-clip-llama3.1-8b COCO-Deceptive-CLIP-LLaMA-3.1-8B Training Dataset 🏆 This work is accepted to ACL 2025 (Main Conference). Figure: Attack success rate (ASR) and caption diversity of our model on the COCO dataset, illustrating its ability to generate deceptive captions that successfully fool CLIP. Dataset Details This dataset provides instruction–response pairs formatted as short two-turn conversations: The user message contains: A given image caption. A set of task… See the full description on the dataset page: https://huggingface.co/datasets/ahnpersie/coco-deceptive-clip-llama3.1-8b.texttext-generation100K<n<1M0 likes30 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.