datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgricultureVideosQnAThe dataset is in XLS format with multiple sheets named for different languages.
The dataset is primarily used for training and ground truth of answers that can be generated for agriculture related queries from the videos.
Each sheet has list of video urls (youtube links) and the question that can be asked, corresponding answers that can be generated from the videos, source of information in the answer and time stamps.
The sources of information could be:
Transcript: based on what one hears… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/AgricultureVideosQnA.AgricultureVideosTranscriptThis dataset consists of agriculture videos in hindi and oriya.
The dataset consists of mulitple xls files and each xls file has column of video urls (youtube video links) and corresponding transcripts.
The transcripts are:
Generated by ASR models (for the purpose of benchmarking)
Manual transcripts
Time stamps
Manual translations
This dataset can be used for training and benchmarking domain specific models for ASR and translation. The time stamps serve as the soruce of audio and the… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/AgricultureVideosTranscript.
