grt
Datasets
All datasets matching “grt”grthyjukipouygrthyuiupou567ne-asr-dataset-grt
Garo (grt) — ASR dataset
A small Garo (grt) speech-to-text dataset for automatic speech recognition
(ASR) of a low-resource North-East India language. Each example pairs a short audio
clip with its Romanized (Latin-script) transcript.
Source
Derived from the ARTPARK-IISc Vaani project (https://vaani.iisc.ac.in/)
Splits
Split
Samples
train
33,480
validation
4,253
test
4,101
Data fields
Each example has:
audio — the… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-grt.ne-asr-dataset-grt-aug
NE ASR Augmented Dataset -- Garo (grt)
Augmented automatic speech recognition dataset for Garo (grt),
a Tibeto-Burman language spoken in Meghalaya, India.
Source
Augmented from sulabhkatiyar/ne-asr-grt
(original transcribed speech data from the ARTPARK-IISc Vaani project).
Language Information
Property
Value
Language
Garo
ISO 639-3
grt
Family
Tibeto-Burman
Region
Meghalaya, India
Tonal
No
Tier
E (47.33h original data)… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-grt-aug.ne-tts-f5-grt
NE-TTS F5 Garo (grt)
F5-TTS training dataset for Garo (grt). Contains 24,772 clips at 24kHz (SNR >= 20dB) from the cleaned NE-TTS dataset, formatted for F5-TTS training.
Stats
Metric
Value
Clips
24,772
Hours
29.6h
Sample rate
24kHz
SNR filter
>= 20dB
Source
ne-tts-grt
Schema
Column
Type
Description
audio
Audio
24kHz WAV audio
text
string
Cleaned transcript
Note
Audio is upsampled from 16kHz… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-f5-grt.grtc-1
