CoolFace
Datasetpublic

surindersinghssj/gurbani-sehajpath-yt-captions-canonical

Gurbani Sehajpath — Canonical-aligned ASR corpus Stage-1 + Stage-2 canonical pipeline output for sehaj-path (calm recitation of the Guru Granth Sahib). Built from publicly available audio recordings with aligned transcripts, chunked by caption timing and aligned against the canonical Guru Granth Sahib Ji text (SGGS). Columns Schema is auto-inferred from the parquet shards. Primary columns: audio — 16 kHz mono waveform final_text — canonical Gurmukhi transcription… See the full description on the dataset page: https://huggingface.co/datasets/surindersinghssj/gurbani-sehajpath-yt-captions-canonical.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes295downloads
Dataset Card

Gurbani Sehajpath — Canonical-aligned ASR corpus

Stage-1 + Stage-2 canonical pipeline output for sehaj-path (calm recitation of the Guru Granth Sahib). Built from publicly available audio recordings with aligned transcripts, chunked by caption timing and aligned against the canonical Guru Granth Sahib Ji text (SGGS).

Columns

Schema is auto-inferred from the parquet shards. Primary columns:

  • —audio — 16 kHz mono waveform
  • —final_text — canonical Gurmukhi transcription (post Stage-2 alignment, recommended for training)
  • —text / raw_text — earlier-stage text fields for comparison / ablation
  • —clip_id, video_id, start_s, end_s, duration_s — provenance + timing
  • —sggs_line, canonical_shabad_id, canonical_line_ids — SGGS alignment targets
  • —canonical_match_score, canonical_retrieval_margin, canonical_op_counts — alignment quality signals
  • —is_simran, decision — segment classifications
  • —caption_lang, caption_offset_s, n_cues, clip_mode — caption-pipeline metadata

Intended use

Training / fine-tuning automatic speech recognition models for Gurbani sehaj-path audio → Gurmukhi transcription. Used as one of the primary training sources for `surindersinghssj/surt-small-v3`.

Related

License

CC BY 4.0.