CoolFace
Datasetpublic

surindersinghssj/gurbani-sehajpath-yt-captions-canonical

Gurbani Sehajpath — Canonical-aligned ASR corpus Stage-1 + Stage-2 canonical pipeline output for sehaj-path (calm recitation of the Guru Granth Sahib). Built from publicly available audio recordings with aligned transcripts, chunked by caption timing and aligned against the canonical Guru Granth Sahib Ji text (SGGS). Columns Schema is auto-inferred from the parquet shards. Primary columns: audio — 16 kHz mono waveform final_text — canonical Gurmukhi transcription… See the full description on the dataset page: https://huggingface.co/datasets/surindersinghssj/gurbani-sehajpath-yt-captions-canonical.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes295downloads
7 commits on main
4d18b975mo ago

Rewrite dataset card without YouTube references; describe as publicly available recordings with aligned transcripts

surindersinghssj
283080d5mo ago

fix: strip stale dataset_info YAML so HF re-infers schema from parquet shards

surindersinghssj
c5394a65mo ago

Upload dataset

surindersinghssj
ce71cf25mo ago

Upload dataset

surindersinghssj
c7d7f1e5mo ago

Upload dataset

surindersinghssj
a4c70695mo ago

Upload dataset

surindersinghssj
b5b3e465mo ago

initial commit

surindersinghssj