MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-jealousy
Dataset Summary Synthetic Persian Chatbot Conversational SA – Jealousy is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "jealousy" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-jealousy.
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Jealousy is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "jealousy" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
- Language(s): Persian (Farsi)
- Task(s): Classification (Emotion Classification – Jealousy)
- Source: Synthetic, generated using GPT-4o-mini
- Part of FaMTEB: Yes
Supported Tasks and Leaderboards
The dataset evaluates text embedding models’ ability to detect the presence and intensity of jealousy in chatbot dialogues. Benchmark results are reported on the Persian MTEB Leaderboard on Hugging Face Spaces (filter by language: Persian).
Construction
This dataset was created using the following procedure:
- 175 conversation topics were predefined
- User and chatbot tones were selected from 9 combinations (e.g., formal, casual, childish)
- The emotion "jealousy" was chosen and assigned an intensity: neutral, moderate, or high
- GPT-4o-mini was used to generate conversations based on these inputs
Labeling Strategy:
- Positive Label: Emotion intensity is moderate or high
- Negative Label: Emotion intensity is neutral (indicating no significant jealousy)
According to the FaMTEB paper (Table 1), the broader parent dataset (covering all emotions) received a 93% human validation accuracy for label quality.
Data Splits
The aggregate dataset (SynPerChatbotConvSAClassification) has:
- Train: 4,496 samples
- Dev: 0 samples
- Test: 1,499 samples
This jealousy-specific subset co
