CoolFace
Datasetpublicgated

sajalmadan0909/hindi_and_english_stt_tts_master_data

Hindi and English STT/TTS Master Data Combined speech dataset for Hindi and Indian English automatic speech recognition (ASR) and text-to-speech (TTS) training. Parquet shards embed WAV audio bytes with transcripts. Dataset structure hindi/<source>/train-*.parquet english/<source>/train-*.parquet Each config loads one source independently (~3.24M total rows, ~1.9 TB). Features Column Type Description audio Audio WAV bytes embedded in… See the full description on the dataset page: https://huggingface.co/datasets/sajalmadan0909/hindi_and_english_stt_tts_master_data.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes18downloads

sajalmadan0909/hindi_and_english_stt_tts_master_data · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.