CoolFace
Datasetpublic

TBOGamer22/BrahuiSpeech-70H-V2

BrahuiSpeech-70H V2 BrahuiSpeech-70H V2 is an approximately 70-hour automatic speech recognition dataset containing 15,626 audio-transcription pairs and 69 hours, 32 minutes, 12 seconds of real-world Brahui (Brahvi) speech. Brahui (brh) is a low-resource Dravidian language spoken primarily in Balochistan, Pakistan. The dataset covers naturally occurring speech across varied speakers, speaking styles, media domains, and acoustic conditions. Transcriptions use the Perso-Arabic… See the full description on the dataset page: https://huggingface.co/datasets/TBOGamer22/BrahuiSpeech-70H-V2.

sourceHugging Facemitupdated 12d agoView on Hugging Face
1likes195downloads
8 commits on main
6b534b612d ago

Update README.md

TBOGamer22
3b0354c12d ago

Add BrahuiSpeech-70H V2 dataset card

TBOGamer22
7f51c1812d ago

Publish BrahuiSpeech-70H V2 audio and adjudicated transcripts

TBOGamer22
9f5be5212d ago

Update README.md

TBOGamer22
a8118af12d ago

Add V2 release statistics

TBOGamer22
ea4b77b12d ago

Add BrahuiSpeech-70H V2 dataset card

TBOGamer22
2f45a4612d ago

Publish BrahuiSpeech-70H V2 audio and adjudicated transcripts

TBOGamer22
46acf1912d ago

initial commit

TBOGamer22