simile
Datasets
All datasets matching “simile”scope_simile_generation
SCOPE Simile
Dataset Summary
This dataset has been created for the purpose of generating similes from literal descriptive sentences.
The process involves a two-step approach: firstly, self-labeled similes are converted into literal sentences using structured common sense knowledge, and secondly, a seq2seq model is fine-tuned on these [literal sentence, simile] pairs to generate similes. The dataset was collected from Reddit, specifically from the subreddits WRITINGPROMPTS… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/scope_simile_generation.wps_chinese_simile
WPS - Chinese Simile
Dataset Summary
Chinese Simile (CS) Dataset
This dataset is constructed and based on the online free-access fictions that are tagged with sci-fi, urban novel, love story, youth, etc.
All similes are extracted by rich regular expression, and the extraction precision is estimated as 92% by labelling 500 random extracted samples. Further data filtering as well as processing is truly encouraged!
The data split in paper is as follows (You could find more… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/wps_chinese_simile.sft_simileDPO_train_simile_balanced
