CoolFace
Datasetpublic

froggeric/creativity

"The only difference between Science and screwing around is writing it down." (Adam Savage) The LLM Creativity benchmark Last benchmark update: 28 May 2024 The goal of this benchmark is to evaluate the ability of Large Language Models to be used as an uncensored creative writing assistant. Human evaluation of the results is done manually, by me, to assess the quality of writing. There are 24 questions, some standalone, other follow-ups to previous questions for a multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/froggeric/creativity.

sourceHugging Faceupdated 2y agoView on Hugging Face
88likes75downloads
50 commits on main
bbcd78a2y ago

2024-05-28 update

froggeric
9040c182y ago

Upload benchmark-results.csv

froggeric
8858bd32y ago

formatting

froggeric
8759e312y ago

Add refusal bypass method, and link to coding leaderboard

froggeric
41ecd6b2y ago

Upload benchmark-results.csv

froggeric
1d091c52y ago

updated screenshot

froggeric
7ed141e2y ago

presentation

froggeric
48439292y ago

typo

froggeric
035f90c2y ago

updated results with new models, and added recommendations

froggeric
ea92d7a2y ago

Upload benchmark-results.csv

froggeric
d7ff0072y ago

corrected info for command-r-plus and added link to prompting guide

froggeric
e78630d2y ago

New results and observations from 2024-04-16

froggeric
fc384ba2y ago

Upload benchmark-results.csv

froggeric
7f490172y ago

removed duplicate

froggeric
115a06b2y ago

Added link to Creative Writing Leaderboard

froggeric
50d99382y ago

added link to UGI leaderboard

froggeric
296a6572y ago

Update README.md

froggeric
17c878f3y ago

Added https://eqbench.com/creative_writing.html

froggeric
26cd4803y ago

Updated results from 2024-03-12

froggeric
e3ca63a3y ago

Update README.md

froggeric
d30f7c23y ago

Update README.md

froggeric
2546bea3y ago

Update README.md

froggeric
292beac3y ago

Update README.md

froggeric
3133a8c3y ago

Update README.md

froggeric
aca84063y ago

Update README.md

froggeric
f50aa393y ago

Delete sample_reply_westlake-v2-7b.md

froggeric
a0a61a93y ago

Delete sample_reply_miqu-1-70b.md

froggeric
44417e13y ago

Delete sample_reply_daybreak-kunoichi-dpo-7b.md

froggeric
cc80e153y ago

Delete sample_prompt.txt

froggeric
045b6d33y ago

Create sample_reply_daybreak-kunoichi-dpo-7b.md

froggeric
4a016dc3y ago

Update sample_reply_miqu-1-70b.md

froggeric
15a56f03y ago

Update sample_reply_westlake-v2-7b.md

froggeric
0f9e4d03y ago

Create sample_reply_westlake-v2-7b.md

froggeric
a0246ca3y ago

Rename sample_prompt.md to sample_prompt.txt

froggeric
1c3a85c3y ago

Rename sample_prompt.txt to sample_prompt.md

froggeric
ae3ae793y ago

Update sample_reply_miqu-1-70b.md

froggeric
17890f13y ago

Rename sample_reply_miqu-1-70b.txt to sample_reply_miqu-1-70b.md

froggeric
ce6bcf23y ago

Create sample_reply_miqu-1-70b.txt

froggeric
77fe82a3y ago

Create sample_prompt.txt

froggeric
aa300d83y ago

Update README.md

froggeric
7a945793y ago

Update README.md

froggeric
78c03ac3y ago

Update README.md

froggeric
fad4ffc3y ago

Update README.md

froggeric
02cb7093y ago

Update README.md

froggeric
348e1143y ago

Upload benchmark-results.csv

froggeric
3c689c23y ago

Update README.md

froggeric
3150c8c3y ago

Update README.md

froggeric
ce9382c3y ago

Update README.md

froggeric
c39d7993y ago

Update README.md

froggeric
91cb36a3y ago

Update README.md

froggeric