CoolFace
Datasetpublic

Paytmlabs/S2R_Shrutilipi_hindi

Paytmlabs/S2R_Shrutilipi_hindi Hindi speech dataset prepared from ai4bharat/Shrutilipi for Ultravox training. Viewing samples on Hugging Face The hindi config stores audio inside Parquet. The website dataset viewer often cannot decode that and shows no rows. To inspect examples in the browser, open the Subset (config) drop-down and choose hindi_text_samples — text and continuation only (~2000 rows). Ultravox training should keep using subset hindi (full audio).… See the full description on the dataset page: https://huggingface.co/datasets/Paytmlabs/S2R_Shrutilipi_hindi.

sourceHugging Faceunknownupdated 6mo agoView on Hugging Face
0likes351downloads
Dataset Card

Paytmlabs/S2RShrutilipihindi

Hindi speech dataset prepared from ai4bharat/Shrutilipi for Ultravox training.

Viewing samples on Hugging Face

The `hindi` config stores audio inside Parquet. The website dataset viewer often cannot decode that and shows no rows.

To inspect examples in the browser, open the Subset (config) drop-down and choose `hindi_text_samples` — text and continuation only (~2000 rows).

Ultravox training should keep using subset `hindi` (full audio).

Schema (config: hindi)

ColumnTypeDescription
audioAudioSpeech audio
textstringVerbatim transcript
continuationstringLLM-generated continuation (≤50 words)

Progress

  • Train chunks: 73/73
  • Validation: done