MCINext/synthetic-persian-chatbot-tone-user-classification
Dataset Summary Synthetic Persian Chatbot Tone User Classification (SynPerChatbotToneUserClassification) is a Persian (Farsi) dataset built for the Classification task. It focuses on identifying the tone of user input in a chatbot conversation. This dataset is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically created using the GPT-4o-mini language model. Language(s): Persian (Farsi) Task(s): Classification (Tone Detection) Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-tone-user-classification.
Dataset Summary
Synthetic Persian Chatbot Tone User Classification (SynPerChatbotToneUserClassification) is a Persian (Farsi) dataset built for the Classification task. It focuses on identifying the tone of user input in a chatbot conversation. This dataset is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically created using the GPT-4o-mini language model.
- Language(s): Persian (Farsi)
- Task(s): Classification (Tone Detection)
- Source: Synthetically generated using GPT-4o-mini
- Part of FaMTEB: Yes
Supported Tasks and Leaderboards
The dataset evaluates a model's ability to classify user input tone (e.g., polite, aggressive, neutral) in chatbot environments. This is essential for emotion-aware and adaptive dialogue systems. Results can be found on the Persian MTEB Leaderboard under the classification category.
Construction
- Conversation snippets and user messages were generated in various tones using GPT-4o-mini.
- A predefined tone taxonomy (e.g., polite, rude, enthusiastic, sad) was used.
- Messages were labeled accordingly based on prompt conditioning.
- A sample set was manually validated to ensure accuracy and consistency.
Data Splits
- Train: 87,208 samples
- Development (Dev): 0 samples
- Test: 21,799 samples
