chargoddard/rpguild
Data scraped from roleplayerguild and parsed into prompts with a conversation history and associated character bio. Thanks to an anonymous internet stranger for the original scrape. As usernames can be associated with multiple character biographies, assignment of characters is a little fuzzy. The char_confidence feature reflects how likely this assignment is to be correct. Not all posts in the conversation history necessarily have an associated character name. The column has_nameless reflects… See the full description on the dataset page: https://huggingface.co/datasets/chargoddard/rpguild.
Data scraped from roleplayerguild and parsed into prompts with a conversation history and associated character bio. Thanks to an anonymous internet stranger for the original scrape.
As usernames can be associated with multiple character biographies, assignment of characters is a little fuzzy. The char_confidence feature reflects how likely this assignment is to be correct. Not all posts in the conversation history necessarily have an associated character name. The column has_nameless reflects this.
Each row should fit into 4096 Llama tokens, depending on your prompt format - there's built in slack of 128 tokens + 8 per message.
There are a few configurations available. I highly recommend not using the default configuration as it contains a lot of questionable quality data. The options, in order of increasing usefulness:
default- ocean of garbage with some gemshigh_confidence- only entries with no nameless posts that are highly likely to be assigned a correctchar_name/biopruned- Further filtered fromhigh_confidenceto remove common types of junk repliesgrammar_filtered- run through a grammar checker to remove rows with too many mistakes
The grammar_filtered configuration is almost certainly what you want to be using. (Unless you want to do your own processing and filtering.)
