nptel
Datasets
All datasets matching “nptel”nptel_hciindian-english-nptel-v0NPTEL_TECH_ENG_50NPTEL
BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages
Overview
BhasaAnuvaad, is the largest Indic-language AST dataset spanning over 44,400 hours of speech and 17M text segments for 13 of 22 scheduled Indian languages and English.
This repository consists of parallel data for Speech Translation from NPTEL, a subset of BhasaAnuvaad.
How to use
The datasets library allows you to load and pre-process your dataset in pure Python, at… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/NPTEL.indian-english-nptel-testGAPS-nptel
GAPS: Golden-Aligned Parallel Speech Corpus
Overview
GAPS (Golden-Aligned Parallel Speech) is a multi-corpus dataset designed for foreign accent conversion.
The dataset provides parallel speech triplets consisting of:
Original non-native speech
Parallel native speech
Golden speaker speech — synthetic speech that preserves the non-native speaker’s timbre and timing (including pauses) while exhibiting native pronunciation
along with the corresponding text transcript.
GAPS… See the full description on the dataset page: https://huggingface.co/datasets/warisqr007/GAPS-nptel.
