afnankhth/AEConvs
Dataset Card for AEConvs AEConvs (Arabic Empathetic Conversations) is a genuine Arabic conversational dataset that features more than 4K open-domain dyadic empathetic conversations. Each conversation in the dataset was conducted between two humans, resulting in authentic and natural conversations. The dataset is written in Modern Sandard Arabic (MSA). which is the formal and standardized form of the Arabic language used across the Arabic-speaking world. The AEConvs dataset is… See the full description on the dataset page: https://huggingface.co/datasets/afnankhth/AEConvs.
Dataset Card for AEConvs
<!-- Provide a quick summary of the dataset. -->
AEConvs (Arabic Empathetic Conversations) is a genuine Arabic conversational dataset that features more than 4K open-domain dyadic empathetic conversations. Each conversation in the dataset was conducted between two humans, resulting in authentic and natural conversations. The dataset is written in Modern Sandard Arabic (MSA). which is the formal and standardized form of the Arabic language used across the Arabic-speaking world. The AEConvs dataset is constructed with the aim of facilitating research into open-domain conversations in general and, in par- ticular, building and training empathetic conversational models. This dataset provides a valuable resource that captures nuanced emotional and empathetic cues in the Arabic language.
Dataset Language
The dataset is written in Modern Standard Arabic (MSA), which is the universal, written, and formal spoken variety of the Arabic language used across the Arab world.
Dataset Statistics
<!-- Provide a longer summary of what this dataset is. -->
Most of the dataset (96%) consists of short conversations between 4 and 6 turns.
Dataset Sources
<!-- Provide the basic links for the dataset. -->
- Paper: [More Information Needed]
Dataset Structure
Personal and Sensitive Information
<!-- State whether the dataset contains data that might be considered personal, sensitive, or private (e.g., data that reveals addresses, uniquely identifiable names or aliases, racial or ethnic origins, sexual orientations, religious beliefs, political opinions, financial or health data, etc.). If efforts were made to anonymize the data, describe the anonymization process. -->
We believe that all conversations in AEConvs do not contain any personally identifiable information or any sensitive or harmful content, and we expect that models trained using the dataset to be safe and will not generate inappropriate responses that include discrimination, abuse, bias, etc.
Citation
Alkhathlan A, Mirza AA. AEConvs: A Novel Dataset and Benchmark for Evaluating Empathetic Response Generation in Arabic LLMs. Data. 2026; 11(4):85. https://doi.org/10.3390/data11040085
Dataset License
This dataset is licensed under CC BY 4.0.
