datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
escwa
🗣️ ESCWA-CS Corpus
The ESCWA-CS Corpus was collected over two days of meetings of the United Nations Economic and Social Commission for Western Asia (ESCWA) held in 2019.It contains intra-sentential code-switching between Arabic and English, with some speakers—particularly from Algeria, Tunisia, and Morocco—alternating between Arabic and French.
The dataset spans approximately 2.8 hours of speech, featuring dialectal Arabic and a Code Mixing Index (CMI) of around 28%.It serves as a… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/escwa.escrow-smart-contract-functions
