CoolFace
Datasetpublic

Harisri/indic-fusion-hi-mr

Indic-FUSION Hindi-Marathi Pseudo-Parallel Speech A research pseudo-parallel speech corpus for Indic-FUSION. Construction Source speech is sampled from ai4bharat/indicvoices_r. For each source utterance: The source dataset transcript is used directly. NLLB generates a direct Indic-to-Indic target transcript. ai4bharat/IndicF5 synthesizes target-language speech using the source utterance as the reference prompt. The target audio is therefore synthetic, not… See the full description on the dataset page: https://huggingface.co/datasets/Harisri/indic-fusion-hi-mr.

sourceHugging Faceupdated 3d agoView on Hugging Face
0likes20downloads
Dataset Card

Indic-FUSION Hindi-Marathi Pseudo-Parallel Speech

A research pseudo-parallel speech corpus for Indic-FUSION.

Construction

Source speech is sampled from ai4bharat/indicvoices_r.

For each source utterance:

  1. 1.The source dataset transcript is used directly.
  2. 2.NLLB generates a direct Indic-to-Indic target transcript.
  3. 3.ai4bharat/IndicF5 synthesizes target-language speech using the source utterance as the reference prompt.

The target audio is therefore synthetic, not naturally recorded parallel speech.

Directions

  • —Hindi → Marathi
  • —Marathi → Hindi

Splits

Train/validation/test are speaker-disjoint.

Intended use

Research on cross-lingual speech-to-speech translation, speaker preservation, and expressive/prosodic preservation.

Provenance

Source: ai4bharat/indicvoicesr translationmodel": "facebook/nllb-200-distilled-1.3B TTS: ai4bharat/IndicF5 Project: Indic-FUSION