CoolFace
Datasetpublicgated

ARTPARK-IISc/Vaani

VAANI is an India-representative multi-modal multi-lingual dataset. The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages. From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts. Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.

sourceHugging Facecc-by-4.0updated 7d agoView on Hugging Face
157likes19kdownloads

No commit history came back for main. The revision may not exist, or the source declined the request.

ARTPARK-IISc/Vaani · CoolFace