datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ads_Creative_Ad_Copy_Programmatic
Dataset Summary
The Programmatic Ad Creatives dataset contains 7097 samples of online programmatic ad creatives along with their ad sizes. The dataset includes 8 unique ad sizes, such as (300, 250), (728, 90), (970, 250), (300, 600), (160, 600), (970, 90), (336, 280), and (320, 50). The dataset is in a tabular format and represents a random sample from Project300x250.com's complete creative data set. It is primarily used for training and evaluating natural language processing models… See the full description on the dataset page: https://huggingface.co/datasets/PeterBrendan/Ads_Creative_Ad_Copy_Programmatic.scope_simile_generation
SCOPE Simile
Dataset Summary
This dataset has been created for the purpose of generating similes from literal descriptive sentences.
The process involves a two-step approach: firstly, self-labeled similes are converted into literal sentences using structured common sense knowledge, and secondly, a seq2seq model is fine-tuned on these [literal sentence, simile] pairs to generate similes. The dataset was collected from Reddit, specifically from the subreddits WRITINGPROMPTS… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/scope_simile_generation.ColBERT_Humor_Detection
ColBERT_Humor
Dataset Summary
ColBERT Humor contains 200,000 labeled short texts, equally distributed between humorous and non-humorous content. The dataset was created to overcome the limitations of prior humor detection datasets, which were characterized by inconsistencies in text length, word count, and formality, making them easy to predict with simple models without truly understanding the nuances of humor. The two sources for this dataset are the News Category… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/ColBERT_Humor_Detection.vua20_metaphor
VUA20
Dataset Summary
Creative Language Toolkit (CLTK) Metadata
CL Type: Metaphor
Task Type: detection
Size: 200k
Created time: 2020
VUA20 is (perhaps) the largest dataset of metaphor detection used in Figlang2020 workshop.
For the details of this dataset, we refer you to the release paper.
The annotation method of VUA20 is elabrated in the paper of MIP.
Citation Information
If you find this dataset helpful, please cite:
@inproceedings{Leong2020ARO… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/vua20_metaphor.ad-creative-quality-human-vs-llm
Human Expert vs LLM Judge: Facebook Ad Creative Quality
500 real Facebook ads from 253 advertisers, each rated for creative quality by a human ad expert AND by a vision LLM — with the LLM's full reasoning.
The headline finding baked into this data: the human and the LLM agree on image quality only 26.8% of the time. The LLM judge rates 71.8% of ads "good"; the human expert rates only 20% "good". If you are using an LLM as a judge of ad creative (or any subjective visual quality)… See the full description on the dataset page: https://huggingface.co/datasets/AdControlCenter/ad-creative-quality-human-vs-llm.Ads_Creative_Text_Programmatic
Dataset Summary
The Programmatic Ad Creatives dataset contains 1000 samples of online programmatic ad creatives along with their ad sizes. The dataset includes 8 unique ad sizes, such as (300, 250), (728, 90), (970, 250), (300, 600), (160, 600), (970, 90), (336, 280), and (320, 50). The dataset is in a tabular format and represents a random sample from Project300x250.com's complete creative data set. It is primarily used for training and evaluating natural language processing models… See the full description on the dataset page: https://huggingface.co/datasets/PeterBrendan/Ads_Creative_Text_Programmatic.creative-alarm-b33fc0
creative-alarm-b33fc0
Synthetic weather test data: 37 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/emberloom/creative-alarm-b33fc0.Creative_Stories_Logical_ReasoningFrench_Grammar_Explanations
This dataset contains 1500+ French grammar explanations. It's the one I used to train my finetuned LLM called FrenchLlama-3.2-1B-Instruct.
You can use this dataset for your own training purposes & find the aforementioned model on my HuggingFace profile.
Creative-Writing-Thinking
Creative-Writing-Thinking
Using essays-creative-writing-prompts and Qwen3-14b to generate the reasoning traces and answers. We created this reasoning dataset.
Suitable for LLM post-training, especially RL.
moh_metaphor
MOH Dataset
Creative Language Toolkit (CLTK) Metadata
CL Type: Metaphor
Task Type: detection, intrepretation
Size: 1k~2k
Created time: 2016
Description:
Moh dataset is a dataset for metaphor processing, which was released in the paper.
For more details, please check the original paper.
Citation
If you use this dataset, please cite:
@inproceedings{Mohammad2016MetaphorAA,
title={Metaphor as a Medium for Emotion: An Empirical Study},
author={Saif M. Mohammad and… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/moh_metaphor.maqsaCreative_Idea
