datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tunisian-derja-unified-raw-corpusTunisian Derja Unified Raw Corpus
Dataset Description
Repository: hamzabouajila/tunisian-derja-unified-raw-corpus
Paper: Not yet published; dataset card serves as primary documentation
Point of Contact: Hamza Bouajila
License: CC-BY-SA-4.0
Dataset Summary
The Tunisian Derja Unified Raw Corpus is a comprehensive collection of ~802,659 text examples in Tunisian Arabic (Derja), a low-resource dialect of Arabic widely spoken in Tunisia. This raw corpus aggregates data from multiple sources… See the full description on the dataset page: https://huggingface.co/datasets/hamzabouajila/tunisian-derja-unified-raw-corpus.derja_to_msa_dataset
Dataset Card
MADAR (License): We construct 12,000 translation instructions using the Multi-Arabic Dialect Applications and Resources (MADAR) corpus is a collection of parallel sentences covering the dialects of 25 Arab cities. We select the dialect of Tunis city as Derja, along with MSA resulting into two translation directions.
