CoolFace
Datasetpublic

chargoddard/rpguild

Data scraped from roleplayerguild and parsed into prompts with a conversation history and associated character bio. Thanks to an anonymous internet stranger for the original scrape. As usernames can be associated with multiple character biographies, assignment of characters is a little fuzzy. The char_confidence feature reflects how likely this assignment is to be correct. Not all posts in the conversation history necessarily have an associated character name. The column has_nameless reflects… See the full description on the dataset page: https://huggingface.co/datasets/chargoddard/rpguild.

sourceHugging Facecc-by-nc-4.0updated 3y agoView on Hugging Face
33likes94downloads
Dataset Card

Data scraped from roleplayerguild and parsed into prompts with a conversation history and associated character bio. Thanks to an anonymous internet stranger for the original scrape.

As usernames can be associated with multiple character biographies, assignment of characters is a little fuzzy. The char_confidence feature reflects how likely this assignment is to be correct. Not all posts in the conversation history necessarily have an associated character name. The column has_nameless reflects this.

Each row should fit into 4096 Llama tokens, depending on your prompt format - there's built in slack of 128 tokens + 8 per message.

There are a few configurations available. I highly recommend not using the default configuration as it contains a lot of questionable quality data. The options, in order of increasing usefulness:

  • default - ocean of garbage with some gems
  • high_confidence - only entries with no nameless posts that are highly likely to be assigned a correct char_name/bio
  • pruned - Further filtered from high_confidence to remove common types of junk replies
  • grammar_filtered - run through a grammar checker to remove rows with too many mistakes

The grammar_filtered configuration is almost certainly what you want to be using. (Unless you want to do your own processing and filtering.)