rushilrawat/garhwali-speech
Garhwali Speech Companion to Garhwali Corpus, which is currently private while its redacted catalog and rights-reviewed text release are reconciled. This dataset packages the Garhwali subset of Project VAANI audio with its provider transcripts, source metadata, and clearly separated experimental SraVaani drafts. Contents 110,436 source audio rows (14.54 GiB audio). 5,894 provider-transcribed rows, including the provider's train/validation/test splits. 104,542… See the full description on the dataset page: https://huggingface.co/datasets/rushilrawat/garhwali-speech.
Garhwali Speech
Companion to Garhwali Corpus, which is currently private while its redacted catalog and rights-reviewed text release are reconciled. This dataset packages the Garhwali subset of Project VAANI audio with its provider transcripts, source metadata, and clearly separated experimental SraVaani drafts.
Contents
- 110,436 source audio rows (14.54 GiB audio).
- 5,894 provider-transcribed rows, including the provider's train/validation/test splits.
- 104,542 rows with SraVaani draft output, of which 104,508 have non-empty text; all are marked as unreviewed hypotheses rather than reference transcripts.
- 742 records retain differing transcript values from the two upstream VAANI repositories, with a conflict flag and both supplied text values preserved.
- Audio is embedded in Parquet shards; file paths from the upstream dataset and reference images are not republished.
Source and license
Audio and provider transcripts are from ARTPARK-IISc/VAANI and ARTPARK-IISc/VAANI-transcription-part. The upstream Hub card marks VAANI CC BY 4.0 and gates access behind acceptance of its terms and contact-information form. This package preserves attribution, source revisions, and row-level license metadata. Please cite the VAANI paper and review upstream access conditions before use. SraVaani hypotheses are generated by ARTPARK-IISc/SraVaani-1.0 revision f5dd5358325a5208775b91dad98918e079ea2b27 and are not human references.
Splits and privacy
The transcription-part split is authoritative for every labeled audio row, including the transcription-only files absent from the main repository. The complete release contains 110,436 source rows and 110,428 unique audio hashes, totaling 135.510 hours. The 104,542 remaining VAANI rows stay in train without a provider transcript; where available, machine drafts are separately labeled. We omit exact source filenames, reference images, speaker IDs, stay-duration, and fine-grained location fields. District and gender labels are retained as supplied by VAANI. No Garhwali native-speaker adjudication has been completed; model drafts must not be treated as validated text.
Load with datasets.load_dataset("rushilrawat/garhwali-speech", "garhwali_speech"). For the text/resource companion, use datasets.load_dataset("rushilrawat/garhwali-corpus").
