IbrahimAmin/arz-en-parallel-corpus
Egyptian Arabic - English Parallel Corpus 🇪🇬✨🇬🇧 Dataset Description This dataset is a cleaned and filtered merge of multiple Egyptian Arabic - English parallel corpora, containing ~27,000 aligned sentence pairs. It’s designed for researchers and developers working on machine translation, speech translation, and other NLP tasks involving Egyptian Arabic and English. Sources 📚 This dataset integrates and refines data from the following publicly… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimAmin/arz-en-parallel-corpus.
357
