MohamedRashad/mgb2-arabic
MGB-2: Arabic Multi-Dialect Broadcast Media Recognition Dataset Description Dataset Summary The Arabic Multi-Genre Broadcast (MGB-2) dataset is a large-scale speech recognition corpus containing 1,200 hours of Arabic broadcast audio from Aljazeera Arabic TV channel. The dataset spans recordings from March 2005 to December 2015 and covers 19 distinct programme series. It was originally created for the MGB-2 Challenge at SLT-2016, focusing on handling… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/mgb2-arabic.
71.3k
