cahlen/cdg-trending_MickeyRourke_JoJoSiwa_100
Mickey Rourke & JoJo Siwa: The blurred lines between celebrity and reality in the age of social media - Generated by Conversation Dataset Generator This dataset was generated using the Conversation Dataset Generator script available at https://cahlen.github.io/conversation-dataset-generator/. Generation Parameters Number of Conversations Requested: 100 Number of Conversations Successfully Generated: 100 Total Turns: 1090 Model ID:… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-trending_MickeyRourke_JoJoSiwa_100.
Mickey Rourke & JoJo Siwa: The blurred lines between celebrity and reality in the age of social media - Generated by Conversation Dataset Generator
This dataset was generated using the Conversation Dataset Generator script available at https://cahlen.github.io/conversation-dataset-generator/.
Generation Parameters
- Number of Conversations Requested: 100
- Number of Conversations Successfully Generated: 100
- Total Turns: 1090
- Model ID:
meta-llama/Meta-Llama-3-8B-Instruct - Generation Mode: Mode: Creative Brief (
--creative-brief "A discussion between Mickey Rourke and JoJo Siwa about the trending controversy on Celebrity Big Brother UK following homophobic remarks and subsequent apologies.") - Persona 1 Search Term:
Mickey Rourke Celebrity Big Brother homophobic comments - Persona 2 Search Term:
JoJo Siwa response apology homophobic remark - Note: Personas were generated once from the brief. Topic/Scenario/Style were varied for each example based on this brief. Parameters below reflect the last successful example.
- Topic:
The blurred lines between celebrity and reality in the age of social media - Scenario:
A heated debate between Mickey and JoJo at a red-carpet event, where they're both promoting their respective projects - Style:
Passionate, yet nuanced, with a focus on exploring the complexities of fame and responsibility - Included Points:
apology, forgiveness, accountability, understanding, empathy, controversy, LGBTQ+ community
Personas
Mickey Rourke
Description: A seasoned Hollywood actor, known for his tough-guy persona and rough-around-the-edges demeanor. Often speaks in a gruff, gravelly tone, with a tendency to downplay or dismiss criticism. Uses colloquialisms and slang, but can come across as insensitive or dismissive. Has a reputation for being outspoken and unapologetic, even in the face of controversy. -> maps to role: human
JoJo Siwa
Description: A young, bubbly pop star, known for her bright personality and unwavering optimism. Often speaks in a cheerful, enthusiastic tone, with a tendency to use catchphrases and soundbites. Has a strong sense of social justice and is unafraid to speak her mind, especially when it comes to issues affecting the LGBTQ+ community. Has a reputation for being kind and compassionate, with a strong connection to her fans. -> maps to role: gpt
Usage
To use this dataset:
1. Clone the repository:
git lfs install
git clone https://huggingface.co/datasets/cahlen/cdg-trending_MickeyRourke_JoJoSiwa_1002. Load in Python:
from datasets import load_dataset
dataset = load_dataset("cahlen/cdg-trending_MickeyRourke_JoJoSiwa_100")
# Access the data (e.g., the training split)
print(dataset['train'][0])LoRA Training Example (Basic)
Below is a basic example of how you might use this dataset to fine-tune a small model like google/gemma-2b-it using LoRA with the PEFT and TRL libraries.
Note: This requires installing additional libraries: pip install -U transformers datasets accelerate peft trl bitsandbytes torch
import torch
from datasets import load_dataset
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments, BitsAndBytesConfig
from trl import SFTTrainer
# 1. Load the dataset
dataset_id = "cahlen/cdg-trending_MickeyRourke_JoJoSiwa_100"
dataset = load_dataset(dataset_id)
# 2. Load Base Model & Tokenizer (using a small model like Gemma 2B)
model_id = "google/gemma-2b-it"
# Quantization Config (optional, for efficiency)
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16 # or torch.float16
)
# Tokenizer
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
# Set padding token if necessary (Gemma's is <pad>)
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_token
tokenizer.pad_token_id = tokenizer.eos_token_id
# Model
model = AutoModelForCausalLM.from_pretrained(
model_id,
quantization_config=quantization_config,
device_map="auto", # Automatically place model shards
trust_remote_code=True
)
# Prepare model for k-bit training if using quantization
model = prepare_model_for_kbit_training(model)
# 3. LoRA Configuration
lora_config = LoraConfig(
r=8, # Rank
lora_alpha=16,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"], # Adjust based on model architecture
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM"
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# 4. Training Arguments (minimal example)
training_args = TrainingArguments(
output_dir="./lora-adapter-Mickey Rourke-JoJo Siwa", # Choose a directory
per_device_train_batch_size=1,
gradient_accumulation_steps=4,
learning_rate=2e-4,
num_train_epochs=1, # Use 1 epoch for a quick demo
logging_steps=10,
save_steps=50, # Save adapter periodically
fp16=False, # Use bf16 if available, otherwise fp16
bf16=torch.cuda.is_bf16_supported(),
optim="paged_adamw_8bit", # Use paged optimizer for efficiency
report_to="none" # Disable wandb/tensorboard for simple example
)
# 5. Create SFTTrainer
trainer = SFTTrainer(
model=model,
train_dataset=dataset['train'], # Assumes 'train' split exists
peft_config=lora_config,
tokenizer=tokenizer,
args=training_args,
max_seq_length=512, # Adjust as needed
dataset_text_field="content", # Use content field directly
packing=True, # Pack sequences for efficiency
)
# 6. Train
print("Starting LoRA training...")
trainer.train()
### 7. Save the LoRA adapter
# Use a fixed string for the example output directory
trainer.save_model("./lora-adapter-output-directory")
print(f"LoRA adapter saved to ./lora-adapter-output-directory")Dataset Format (JSON Lines source)
Each row in the dataset contains the following keys:
- conversation_id: Unique identifier for the conversation
- turn_number: The sequential number of the turn within a conversation
- role: Either 'human' or 'gpt' indicating who is speaking
- speakername: The actual name of the speaker (e.g., '{finalpersona1}' or '{final_persona2}')
- topic: The conversation topic
- scenario: The scenario in which the conversation takes place
- style: The stylistic direction for the conversation
- include_points: Specific points to include in the conversation
- content: The actual text content of the turn
