CoolFace
Datasetpublicgated

ai4bharat/UGCE-Resources

BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages Overview BhasaAnuvaad, is the largest Indic-language AST dataset spanning over 44,400 hours of speech and 17M text segments for 13 of 22 scheduled Indian languages and English. This repository consists of parallel data for Speech Translation from UGCE-Resources, a subset of BhasaAnuvaad. How to use The datasets library allows you to load and pre-process your dataset in… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/UGCE-Resources.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
0likes29downloads

ai4bharat/UGCE-Resources · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.