CoolFace
Datasetpublic

sarvamai/sarvam-dub-benchmark-set

Sarvam Dubbing Benchmark Dataset Dataset Description Multilingual evaluation dataset for real-time dubbing and voice cloning benchmarking with focus on speaker similarity preservation across same-lingual and cross-lingual scenarios. This dataset was used to benchmark production dubbing systems. Internal evaluations showed higher speaker similarity than ElevenLabs v3 and Cartesia Sonic under an identical scoring protocol. Check out the Sarvam Dub blog for more… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/sarvam-dub-benchmark-set.

sourceHugging Faceotherupdated 8mo agoView on Hugging Face
2likes31downloads
Dataset Card

language:

  • —en
  • —hi
  • —bn
  • —ta
  • —te
  • —kn
  • —ml
  • —mr
  • —gu
  • —or
  • —pa license: other task_categories:
  • —text-to-speech
  • —audio-to-audio prettyname: Sarvam Dubbing Benchmark Dataset sizecategories:
  • —1K<n<10K ---

Sarvam Dubbing Benchmark Dataset

Dataset Description

Multilingual evaluation dataset for real-time dubbing and voice cloning benchmarking with focus on speaker similarity preservation across same-lingual and cross-lingual scenarios.

This dataset was used to benchmark production dubbing systems. Internal evaluations showed higher speaker similarity than ElevenLabs v3 and Cartesia Sonic under an identical scoring protocol.

Check out the Sarvam Dub blog for more details: https://www.sarvam.ai/blogs/sarvam-dub

Evaluation-only dataset — not for training.

Supported Languages

This benchmark includes the following languages:

  • —English (en)
  • —Hindi (hi)
  • —Bengali (bn)
  • —Tamil (ta)
  • —Telugu (te)
  • —Kannada (kn)
  • —Malayalam (ml)
  • —Marathi (mr)
  • —Gujarati (gu)
  • —Odia (or)
  • —Punjabi (pa)

Dataset Summary

  • —Speakers: 64
  • —Languages per speaker: 11
  • —Total samples: 704
  • —Setup: One-shot speaker conditioning
  • —Metric: Speaker similarity

Dataset Schema

Each record contains:

  • —reference_audio — speaker prompt audio
  • —target_text — text to dub
  • —target_language — output language code

Speaker Similarity Scoring

Speaker similarity is computed using the SpeechBrain ECAPA speaker embedding model:

Model: https://huggingface.co/speechbrain/spkrec-ecapa-voxceleb

Intended Use

Dubbing benchmarking, voice cloning evaluation, speaker similarity measurement.