CoolFace
Datasetpublic

MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-sadness

Dataset Summary Synthetic Persian Chatbot Conversational SA – Sadness is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "sadness" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-sadness.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes69downloads
Dataset Card

Dataset Summary

Synthetic Persian Chatbot Conversational SA – Sadness is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "sadness" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.

  • —Language(s): Persian (Farsi)
  • —Task(s): Classification (Emotion Classification – Sadness)
  • —Source: Synthetic, generated using GPT-4o-mini
  • —Part of FaMTEB: Yes

Supported Tasks and Leaderboards

This dataset evaluates text embedding models’ ability to detect the presence and intensity of sadness in chatbot dialogues. Benchmark results are reported on the Persian MTEB Leaderboard on Hugging Face Spaces (filter by language: Persian).

Construction

This dataset was created using the following procedure:

  • —175 conversation topics were predefined
  • —User and chatbot tones were selected from 9 combinations (e.g., formal, casual, childish)
  • —The emotion "sadness" was chosen and assigned an intensity: neutral, moderate, or high
  • —GPT-4o-mini was used to generate conversations based on these inputs

Labeling Strategy:

  • —Positive Label: Emotion intensity is moderate or high
  • —Negative Label: Emotion intensity is neutral (indicating no significant sadness)

According to the FaMTEB paper (Table 1), the broader parent dataset (covering all emotions) received a 93% human validation accuracy for label quality.

Data Splits

The aggregate dataset (SynPerChatbotConvSAClassification) has:

  • —Train: 4,496 samples
  • —Dev: 0 samples
  • —Test: 1,499 samples

This sadness-specific subset contains 436 examples, but exact train/test splits are not separately listed in the FaMTEB paper. It is assumed to be part of the aggregated totals above.