CoolFace
Datasetpublic

Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset

Combined Roleplay Dataset This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks. Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays English content with a few Spanish, Portuguese, and Chinese conversations Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset.

sourceHugging Faceupdated 5d agoView on Hugging Face
0likes71downloads
Dataset Card

Combined Roleplay Dataset

This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks.

  • —Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays
  • —English content with a few Spanish, Portuguese, and Chinese conversations
  • —Conversations limited to 4000 tokens using the Llama 3.1 8B tokenizer (but not done yet for the newest batches)
  • —Structured format featuring alternating user and AI messages, starting with a system or user prompt and ending with an AI response

Dataset Structure

  1. 1.Multiturn mix:
  2. 2.1K version: 1,000 conversations
  3. 3.10K version: 10,000 conversations
  4. 4.30K version: 30,000 conversations
  1. 1.Source datasets:
  2. 2.20231206chaiprizerewardmodel_data: 18,574 lines (only keeping the lines with label = 1)
  3. 3.Cakrawala: 13,000 lines
  4. 4.Capybara: 15,996 lines
  5. 5.CreativeWriting: 8,808 lines
  6. 6.Empathetic_dialogues: 19,531 lines
  7. 7.SODA: 1,155,128 lines
  8. 8.Samantha: 5,868 lines
  9. 9.DIALOGSum: 10,883 lines
  10. 10.RPGPT_PublicDomain: 3,032 lines
  11. 11.Synthetic-characters: 17,668 lines
  12. 12.li2017dailydialog: 13,118 lines (all splits merged)
  13. 13.Conversational-Reasoning-Topical-Chat: 10,784 lines (all splits merged)

Dataset Creation

The MultiturnMix data was created by:

  1. 1.Randomly sampling 20,000 lines from the SODA dataset and combining them with the other datasets
  2. 2.For the v2 and v3 datasets, 13,000 random lines from SODA were selected
  3. 3.Embeddings were computed with the agentlans/snowflake-arctic-embed-xs-zyda-2 model
  4. 4.Clustering the lines into 1,000, 10,000, or 30,000 k-means clusters to ensure diversity

Considerations for Using the Data

Intended Uses

This dataset is primarily intended for training and fine-tuning language models for creative writing, roleplaying, and conversational AI tasks.

Social Impact and Biases

  • —The data may exhibit biases in style and formatting due to its synthetic nature.
  • —It might not represent all genres or fandoms equally well.
  • —Limited to two-player dialogues, which could differ from more complex multi-party interactions.

Limitations

  • —Some subsets may be unsuitable for all audiences.
  • —English with limited representation of other languages.
  • —Not comprehensive coverage of creative writing or roleplaying scenarios.
  • —May still contain poor AI-generated content and repetitive conversations.

Additional Information

Dataset Sources

The dataset incorporates data from the following sources:

Licensing and Privacy

  • —The dataset is not known to contain personal or sensitive information
  • —Users should refer to the original sources for specific licensing information