surindersinghssj/gurbani-sehajpath-yt-captions-canonical
Gurbani Sehajpath — Canonical-aligned ASR corpus Stage-1 + Stage-2 canonical pipeline output for sehaj-path (calm recitation of the Guru Granth Sahib). Built from publicly available audio recordings with aligned transcripts, chunked by caption timing and aligned against the canonical Guru Granth Sahib Ji text (SGGS). Columns Schema is auto-inferred from the parquet shards. Primary columns: audio — 16 kHz mono waveform final_text — canonical Gurmukhi transcription… See the full description on the dataset page: https://huggingface.co/datasets/surindersinghssj/gurbani-sehajpath-yt-captions-canonical.
Gurbani Sehajpath — Canonical-aligned ASR corpus
Stage-1 + Stage-2 canonical pipeline output for sehaj-path (calm recitation of the Guru Granth Sahib). Built from publicly available audio recordings with aligned transcripts, chunked by caption timing and aligned against the canonical Guru Granth Sahib Ji text (SGGS).
Columns
Schema is auto-inferred from the parquet shards. Primary columns:
audio— 16 kHz mono waveformfinal_text— canonical Gurmukhi transcription (post Stage-2 alignment, recommended for training)text/raw_text— earlier-stage text fields for comparison / ablationclip_id,video_id,start_s,end_s,duration_s— provenance + timingsggs_line,canonical_shabad_id,canonical_line_ids— SGGS alignment targetscanonical_match_score,canonical_retrieval_margin,canonical_op_counts— alignment quality signalsis_simran,decision— segment classificationscaption_lang,caption_offset_s,n_cues,clip_mode— caption-pipeline metadata
Intended use
Training / fine-tuning automatic speech recognition models for Gurbani sehaj-path audio → Gurmukhi transcription. Used as one of the primary training sources for `surindersinghssj/surt-small-v3`.
Related
- Eval split (held-out): `gurbani-sehajpath-yt-captions-eval-canonical`
- Companion kirtan corpus: `gurbani-kirtan-yt-captions-300h-canonical`
- Older studio sehaj corpus: `gurbani-sehajpath`
License
CC BY 4.0.
