CoolFace
Datasetpublic

afnankhth/AEConvs

Dataset Card for AEConvs AEConvs (Arabic Empathetic Conversations) is a genuine Arabic conversational dataset that features more than 4K open-domain dyadic empathetic conversations. Each conversation in the dataset was conducted between two humans, resulting in authentic and natural conversations. The dataset is written in Modern Sandard Arabic (MSA). which is the formal and standardized form of the Arabic language used across the Arabic-speaking world. The AEConvs dataset is… See the full description on the dataset page: https://huggingface.co/datasets/afnankhth/AEConvs.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes36downloads
Dataset Card

Dataset Card for AEConvs

<!-- Provide a quick summary of the dataset. -->

AEConvs (Arabic Empathetic Conversations) is a genuine Arabic conversational dataset that features more than 4K open-domain dyadic empathetic conversations. Each conversation in the dataset was conducted between two humans, resulting in authentic and natural conversations. The dataset is written in Modern Sandard Arabic (MSA). which is the formal and standardized form of the Arabic language used across the Arabic-speaking world. The AEConvs dataset is constructed with the aim of facilitating research into open-domain conversations in general and, in par- ticular, building and training empathetic conversational models. This dataset provides a valuable resource that captures nuanced emotional and empathetic cues in the Arabic language.

Dataset Language

The dataset is written in Modern Standard Arabic (MSA), which is the universal, written, and formal spoken variety of the Arabic language used across the Arab world.

Dataset Statistics

<!-- Provide a longer summary of what this dataset is. -->

**Total conversations**4120
Speaker Utterances9869
Listnere Utterances9869
Average utterances per conversation4.8
Average words per conversation70
Average words per utterance14.4

Most of the dataset (96%) consists of short conversations between 4 and 6 turns.

Dataset Sources

<!-- Provide the basic links for the dataset. -->

  • —Paper: [More Information Needed]

Dataset Structure

**Column name****Data type****Description**
conv_idintegerunique identifier for each conversation
conv_txtstringconversation context spanning multiple rows between the speaker and the listener
rolestringrole of the conversation utterance: 'speaker' or 'listener'
Personal and Sensitive Information

<!-- State whether the dataset contains data that might be considered personal, sensitive, or private (e.g., data that reveals addresses, uniquely identifiable names or aliases, racial or ethnic origins, sexual orientations, religious beliefs, political opinions, financial or health data, etc.). If efforts were made to anonymize the data, describe the anonymization process. -->

We believe that all conversations in AEConvs do not contain any personally identifiable information or any sensitive or harmful content, and we expect that models trained using the dataset to be safe and will not generate inappropriate responses that include discrimination, abuse, bias, etc.

Citation

Alkhathlan A, Mirza AA. AEConvs: A Novel Dataset and Benchmark for Evaluating Empathetic Response Generation in Arabic LLMs. Data. 2026; 11(4):85. https://doi.org/10.3390/data11040085

Dataset License

This dataset is licensed under CC BY 4.0.