BrunoHays/wikitongues-darija
Wikitongues-Darija This is a small test dataset for Automatic Speech Recognition in Darija language, built from 2 captioned videos of the WikiTongues project: nawal anass Process: each webm video has been converted to monochannel 16khz wav files with ffmpeg : ffmpeg -i WIKITONGUES-_Nawal_speaking_Moroccan_Arabic.webm.1080p.vp9.webm -ar 16000 -ac 1 nawal.wav each audio has been cut in samples of less than 30 seconds audio according to the captions timestamps. The script may… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/wikitongues-darija.
Update README.md
Upload 41 files
Delete data
Upload metadata.csv
Update prepare_dataset.py
Upload 42 files
Delete data
Update README.md
script used to build the dataset
Upload 74 files
initial commit
