CoolFace
Datasetpublic

MCINext/synthetic-persian-text-tone-classification

Dataset Summary Synthetic Persian Text Tone Classification (SynPerTextToneClassification) is a Persian (Farsi) dataset created for the Classification task, focusing on identifying the emotional or tonal content of text. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using the GPT-4o-mini model, providing examples across multiple tones like formal, informal, positive, negative, neutral, etc. Language(s): Persian (Farsi)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-text-tone-classification.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes19downloads
Dataset Card

Dataset Summary

Synthetic Persian Text Tone Classification (SynPerTextToneClassification) is a Persian (Farsi) dataset created for the Classification task, focusing on identifying the emotional or tonal content of text. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using the GPT-4o-mini model, providing examples across multiple tones like formal, informal, positive, negative, neutral, etc.

  • Language(s): Persian (Farsi)
  • Task(s): Classification (Tone Identification)
  • Source: Synthetically generated using GPT-4o-mini
  • Part of FaMTEB: Yes

Supported Tasks and Leaderboards

This dataset benchmarks the ability of language models to detect tone in Persian text, which is a critical task for applications in sentiment analysis, conversational agents, and content moderation. Model results are shown on the Persian MTEB Leaderboard under classification tasks.

Construction

  1. 1.A taxonomy of tones was defined, including sentiment-based (positive, negative, neutral) and stylistic (formal, informal) labels.
  2. 2.GPT-4o-mini was used to generate Persian sentences labeled with these tones.
  3. 3.Prompts were designed to vary the sentence structure, vocabulary, and context to enhance generalization.
  4. 4.Human evaluators reviewed a sample to confirm tone alignment, and >95% agreement was observed.

Data Splits

  • Train: 80,415 samples
  • Development (Dev): 0 samples
  • Test: 19,739 samples