datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
19th-century-novelists19th-century novelists' sentences
We constructed the 5-author dataset using texts from Project Gutenberg, focusing on five prominent 19th-century novelists: Charles Dickens, Mark Twain, Herman Melville, Jane Austen, and Louisa May Alcott. This selection balances male and female authors as well as British and American literary traditions, offering a diverse testbed for stylistic analysis. Sentence segmentation was performed with the NLTK library, and tokenization/word counts were… See the full description on the dataset page: https://huggingface.co/datasets/Mosab-Rezaei/19th-century-novelists.novelprompts
NovelPrompts
NovelPrompts is an English safety dataset (194 prompts + completions)
built to test LLM judges' behaviour when evaluating the safety of prompts.
The prompts are designed such that safety can only be assessed if you understand
a novel concept: a recent event, a new word, or a new meaning of an existing word
(e.g. slang). A concept is considered novel if it appeared after July 2024.
This dataset is intended for evaluating LLM-as-judge safety evaluators, especially… See the full description on the dataset page: https://huggingface.co/datasets/anissa218/novelprompts.
