CoolFace
Datasetpublic

BrunoHays/wikitongues-darija

Wikitongues-Darija This is a small test dataset for Automatic Speech Recognition in Darija language, built from 2 captioned videos of the WikiTongues project: nawal anass Process: each webm video has been converted to monochannel 16khz wav files with ffmpeg : ffmpeg -i WIKITONGUES-_Nawal_speaking_Moroccan_Arabic.webm.1080p.vp9.webm -ar 16000 -ac 1 nawal.wav each audio has been cut in samples of less than 30 seconds audio according to the captions timestamps. The script may… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/wikitongues-darija.

sourceHugging Facecc-by-sa-4.0updated 1y agoView on Hugging Face
1likes23downloads
11 commits on main
01049441y ago

Update README.md

BrunoHays
88e13ec2y ago

Upload 41 files

BrunoHays
e4cd1182y ago

Delete data

BrunoHays
cb43bcf2y ago

Upload metadata.csv

BrunoHays
acf97bd2y ago

Update prepare_dataset.py

BrunoHays
1d20c5f2y ago

Upload 42 files

BrunoHays
2fd95f42y ago

Delete data

BrunoHays
b5824992y ago

Update README.md

BrunoHays
f3281cb2y ago

script used to build the dataset

BrunoHays
771fa8b2y ago

Upload 74 files

BrunoHays
0b370f22y ago

initial commit

BrunoHays