datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgricultureVideosTranscriptThis dataset consists of agriculture videos in hindi and oriya.
The dataset consists of mulitple xls files and each xls file has column of video urls (youtube video links) and corresponding transcripts.
The transcripts are:
Generated by ASR models (for the purpose of benchmarking)
Manual transcripts
Time stamps
Manual translations
This dataset can be used for training and benchmarking domain specific models for ASR and translation. The time stamps serve as the soruce of audio and the… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/AgricultureVideosTranscript.KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers.
Process of collection of data:
Selected users were given the option of doing a task and getting paid for it.
The users were supposed to record the sentence as it appeared on the screen.
The audio file thus obtained was validated matched with the sentences to fine tune the model.
Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/KikuyuASR_trainingdataset.
