CoolFace
Datasetpublic

projectkaira/Pretraining-V1

Indic TTS Unified v1 A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This dataset consolidates 17 distinct source datasets into a single, schema-normalized resource covering Indian / South Asian languages, plus major European, African, MENA, and Central Asian languages, with over 13.7 million utterances and 26,000+ hours of audio. All audio is resampled to 24 kHz mono. Every row follows an identical schema regardless of source… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/Pretraining-V1.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes11kdownloads
Dataset Card

Indic TTS Unified v1

A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This dataset consolidates 17 distinct source datasets into a single, schema-normalized resource covering Indian / South Asian languages, plus major European, African, MENA, and Central Asian languages, with over 13.7 million utterances and 26,000+ hours of audio.

All audio is resampled to 24 kHz mono. Every row follows an identical schema regardless of source, enabling seamless multi-dataset training without per-source preprocessing.


Dataset Summary

StatisticValue
Total utterances13,777,541
Total audio duration26,300+ hours
Languages covered58+
Audio format24 kHz, mono, float32
Configs (subsets)26

Subsets

ConfigRowsHoursSpeakersLanguagesSource Dataset
orpheus_distill_neucodec400~2--1BarryFutureman/orpheus-distill-neucodec
maya_distill_neucodec14,00033.612,8011BarryFutureman/maya-distill-data-neucodec
emodb_neucodec22,04340.351BarryFutureman/EmoDB-neucodec
expresso_neucodec11,59910.941BarryFutureman/expresso-neucodec
nonverbal_tts6,22217.62,2961deepvk/NonverbalTTS
elise1,1942.611MrDragonFox/Elise
elise_hindi1,1472.411ronith09/Elise-Hindi
synthetic_v110,75922.8509kenpath/tts-synthetic-v1
spicor50,46899.421kenpath/tts-SPICOR
indictts294,008527.0151,24714SPRINGLab/IndicTTS (14 datasets)
msft_indian115,392134.9115,3903deepdml/microsoft-speech-corpus-indian
syspin786,6251,706.5189kenpath/tts-SYSPIN
ivr664,2081,656.510,15222ai4bharat/indicvoices_r
rasa582,1951,035.54022ai4bharat/Rasa
kathbath805,7211,475.298512ai4bharat/Kathbath
shrutilipi2,226,7534,665.0--16ai4bharat/Shrutilipi
cv22_sidon3,212,8584,614.2--17sarulab-speech/commonvoice22_sidon
cv22_african725,125~1,500--6sarulab-speech/commonvoice22_sidon (African subset: sw, lg, ha, yo, ig, am)
cv22_central_asian404,288~800--4sarulab-speech/commonvoice22_sidon (Central Asian subset: uz, ka, az, kk)
cv22_mena194,077~370--2sarulab-speech/commonvoice22_sidon (MENA subset: ar, fa)
cv22_de699,462~1,450--1sarulab-speech/commonvoice22_sidon (German)
cv22_fr700,202~1,450--1sarulab-speech/commonvoice22_sidon (French)
cv22_es1,592,537~3,300--1sarulab-speech/commonvoice22_sidon (Spanish)
cv22_european591,663~1,200--6sarulab-speech/commonvoice22_sidon (European subset: it, nl, tr, ru, pt, pl)
arabic_misc49,412122.739,8981Mixed: MohamedRashad, Nourhann, NeoBoy, saleh1312, KejueAI
uq_speech16,18328.016,1831ixxan/mms-tts-uig-script_arabic-UQSpeech
libritts_r358,000585.02,4561parler-tts/libritts_r_filtered
Total14,135,54126,885+58+

Schema

All configs share the same column schema:

ColumnTypeDescription
audioAudio (24 kHz)Audio waveform, resampled to 24 kHz mono
textstringTranscript text. Rasa transcripts may include emotion tags (see below)
speaker_idstringSpeaker identifier (see Speaker ID Policy below)
sourcestringName of the originating dataset (e.g., "kathbath", "rasa")
languagestringFull language name (e.g., "Hindi", "Bengali", "Tamil")
genderstring"Male", "Female", or "Unknown"
durationfloat64Audio duration in seconds

Speaker ID Policy

Speaker identification varies by source dataset:

  • Deterministic 8-character hash: For datasets that provide speaker metadata (kathbath, syspin, ivr, rasa, spicor, synthetic_v1, cv22_sidon, cv22_african, cv22_central_asian, cv22_mena, cv22_de, cv22_es, cv22_fr, cv22_european), the speaker_id is a deterministic hash derived from the original speaker label, ensuring consistency across rows from the same speaker.
  • Random UUID: For datasets without reliable speaker metadata (shrutilipi, msft_indian, indictts), each row receives a unique random UUID. These should not be used for speaker-level grouping.

Duration

Duration values are unfiltered -- no minimum or maximum duration threshold (such as 0.5--60s) has been applied. Downstream consumers should apply their own filtering as needed.

Emotion Tags (Rasa)

The rasa config contains expressive/emotional speech. Transcript text in this subset may include inline emotion tags such as <happy>, <sad>, <angry>, <surprise>, <fear>, <disgust>, and <neutral>. These tags indicate the intended emotion of the utterance and can be used for emotion-conditioned TTS training.


Language Coverage

The dataset spans a broad range of Indian languages. The table below lists languages and the configs in which they appear:

LanguageConfigs
Assameseshrutilipi, ivr, rasa, cv22_sidon
Bengalikathbath, shrutilipi, ivr, rasa, syspin, cv22_sidon
Bodoivr, rasa
Dhivehicv22_sidon
Dogrishrutilipi, ivr, rasa
Dutchcv22_european
English (Indian)spicor, ivr, rasa, indictts
Arabicarabicmisc, cv22mena
Frenchcv22_fr
Germancv22_de
Italiancv22_european
Polishcv22_european
Portuguesecv22_european
Russiancv22_european
Spanishcv22_es
Turkishcv22_european
Uyghuruq_speech
English (Common Voice)cv22_sidon
Gujaratikathbath, shrutilipi, ivr, rasa, syspin, indictts
Hindikathbath, shrutilipi, ivr, rasa, syspin, msftindian, indictts, syntheticv1, cv22_sidon
Kannadakathbath, shrutilipi, ivr, rasa, syspin, indictts
Kashmiriivr
Konkanishrutilipi, ivr, rasa
Maithilishrutilipi, ivr, rasa
Malayalamkathbath, shrutilipi, ivr, rasa, syspin, indictts, cv22_sidon
Manipuriivr, rasa
Marathikathbath, shrutilipi, ivr, rasa, syspin, indictts, cv22_sidon
Nepalishrutilipi, ivr, rasa, cv22_sidon
Odiakathbath, shrutilipi, ivr, rasa, syspin, indictts, cv22_sidon
Pashtocv22_sidon
Punjabikathbath, shrutilipi, ivr, rasa, syspin, indictts, cv22_sidon
Rajasthaniindictts
Sanskritkathbath, shrutilipi, ivr, rasa
Santaliivr, cv22_sidon
Saraikicv22_sidon
Sindhiivr, cv22_sidon
Tamilkathbath, shrutilipi, ivr, rasa, syspin, msftindian, indictts, cv22sidon
Telugukathbath, shrutilipi, ivr, rasa, syspin, msftindian, indictts, cv22sidon
Urdukathbath, shrutilipi, ivr, rasa, cv22_sidon

Detailed Subset Descriptions

orpheusdistillneucodec

Decoded from BarryFutureman/orpheus-distill-neucodec, an Orpheus distillation dataset stored as NeuCodec tokens. Contains 400 English utterances (~2 hours) with emotion conditioning. Text includes emotion wrapper tags (e.g., <happy>...</happy>) and converted vocal expression tags (e.g., <sigh>, <laugh>). Speaker IDs are random 8-character hex values (no speaker metadata in source).

mayadistillneucodec

Decoded from BarryFutureman/maya-distill-data-neucodec, a Maya distillation dataset stored as NeuCodec tokens. Contains 14,000 English utterances (33.6 hours) with emotion conditioning and rich voice metadata. Text includes emotion wrapper tags (e.g., <happy>...</happy>) and vocal expression tags (e.g., <giggle>, <sigh>, <yawn>). Speaker IDs are deterministic 8-character hashes derived from voice_description, yielding 12,801 unique speakers. Gender breakdown: Male 4,602, Female 4,681, Unknown 4,717.

emodb_neucodec

Decoded from BarryFutureman/EmoDB-neucodec, a synthetic emotional speech dataset with GPT-4o-generated English text and NeuCodec-encoded audio. Contains 22,043 utterances (40.3 hours) after deduplication, with 5 speakers and 7 emotion styles (angry, happy, sad, fearful, surprised, disgusted, neutral). Text includes emotion wrapper tags (e.g., <angry>...</angry>). Speaker IDs are deterministic 8-character hashes of the speaker name. All gender values are "Unknown".

expresso_neucodec

Decoded from BarryFutureman/expresso-neucodec, the Expresso corpus encoded as NeuCodec tokens. Contains 11,599 English utterances (10.9 hours) after deduplication, with 4 speakers and multiple expressive styles. Text includes style wrapper tags (e.g., <confused>...</confused>). Speaker IDs are deterministic 8-character hashes of the original speaker ID (e.g., ex01). All gender values are "Unknown".

nonverbal_tts

Sourced from deepvk/NonverbalTTS, a nonverbal-annotated speech dataset combining Expresso and VoxCeleb data. Contains 6,222 English utterances (17.6 hours) with 2,296 unique speakers. Text uses the annotated Result column which includes emoji markers for nonverbal cues (e.g., 🌬️ for breath, 😤 for exhale). Emotion wrapping applied only for happy and sad categories. Gender breakdown: Male 3,872, Female 2,350.

elise

Sourced from MrDragonFox/Elise, a single-speaker English female dataset. Contains 1,194 utterances (2.6 hours). Audio resampled from 22050 Hz to 24 kHz. Text passed through as-is (includes emotion expression tags).

elise_hindi

Sourced from ronith09/Elise-Hindi, a Hindi version of the Elise dataset with the same speaker. Contains 1,147 utterances (2.4 hours). Audio resampled from 22050 Hz to 24 kHz.

synthetic_v1

Synthetic TTS data generated for bootstrapping and augmentation. Covers 9 languages (primarily Hindi) with 50 distinct synthetic voices. 10,759 utterances totaling 22.8 hours.

spicor

The SpiCor corpus of Indian English read speech. Contains 50,468 utterances (99.4 hours) from 2 speakers. Useful for high-quality single-speaker or few-speaker English TTS.

indictts

Derived from the SPRINGLab/IndicTTS collection, which spans 14 individual language datasets. Contains 294,008 utterances (527.0 hours) across 14 Indian languages. Speaker IDs are random UUIDs (no original speaker metadata available).

msft_indian

Sourced from deepdml/microsoft-speech-corpus-indian. Covers 3 languages (Hindi, Tamil, Telugu) with 115,392 utterances (134.9 hours). Speaker IDs are random UUIDs.

syspin

The SYSPIN TTS dataset provides high-quality studio-recorded speech across 9 languages from 18 speakers. With 786,625 utterances and 1,706.5 hours, this is one of the largest single-source contributions. Well-suited for single-speaker and multi-speaker TTS due to consistent recording conditions.

ivr

Derived from ai4bharat/indicvoices_r (IndicVoices-R), a large-scale read speech corpus. Covers 22 languages with 664,208 utterances (1,656.5 hours) from 10,152 speakers. One of the most linguistically diverse configs in this collection.

rasa

The ai4bharat/Rasa dataset of expressive and emotional Indian language speech. Covers 22 languages with 582,195 utterances (1,035.5 hours) from 40 speakers. Transcripts include inline emotion tags (e.g., <happy>, <sad>, <angry>) that indicate the expressed emotion, making this subset uniquely valuable for emotion-conditioned TTS.

kathbath

Sourced from ai4bharat/Kathbath, a read speech dataset covering 12 Indian languages. Contains 805,721 utterances (1,475.2 hours) from 985 speakers.

Language breakdown by hours:

LanguageHours
Tamil176.9
Marathi152.0
Kannada150.3
Telugu146.7
Hindi139.6
Malayalam139.1
Punjabi128.4
Gujarati113.4
Bengali88.0
Odia81.8
Sanskrit80.4
Urdu78.6

Gender breakdown: Female 982.7h, Male 492.5h

cv22_sidon

Sourced from sarulab-speech/commonvoice22_sidon, a SIDON-processed variant of Mozilla Common Voice 22.0. A curated selection of 17 South Asian / Indic language configs is included, covering all splits (train, validation, test, other, invalidated) merged into a single train split per config. Contains 3,212,858 utterances totaling 4,614.2 hours.

Speaker IDs are deterministic 8-character SHA256 hashes of the original Common Voice client_id (preserves speaker grouping across utterances while anonymizing).

Language breakdown:

LanguageCodeRowsHours
Englishen1,687,5622,670.8
Bengalibn957,9371,129.8
Tamilta181,715314.2
Urduur201,883244.8
Pashtops57,16779.3
Odiaor23,32936.5
Dhivehidv23,87533.3
Sindhisd25,01129.1
Hindihi16,25022.7
Marathimr10,83619.1
Malayalamml9,12110.7
Assameseas4,6567.6
Saraikiskr5,8256.7
Punjabipa-IN3,1364.2
Telugute2,2902.6
Nepaline-NP1,4161.6
Santalisat8491.1

Processing pipeline: raw Common Voice audio (typically MP3 at 32--48 kHz) was decoded, downmixed to mono, and resampled to 24 kHz using high-quality resampling. All splits per language were concatenated. Gender values are mapped from the original gender field (male_masculineMale, female_feminineFemale, otherwise Unknown).

cv22_de

German (de) Common Voice 22, sourced from sarulab-speech/commonvoice22_sidon. Contains 699,462 utterances (~1,450 hours). Same processing pipeline as cv22_sidon. Speaker IDs are deterministic 8-character SHA256 hashes of the original Common Voice client_id.

cv22_fr

French (fr) Common Voice 22, sourced from sarulab-speech/commonvoice22_sidon. Contains 700,202 utterances (~1,450 hours). Same processing pipeline as cv22_sidon. Speaker IDs are deterministic 8-character SHA256 hashes of the original Common Voice client_id.

cv22_es

Spanish (es) Common Voice 22, sourced from sarulab-speech/commonvoice22_sidon. Contains 1,592,537 utterances (~3,300 hours) across 320 train shards. Same processing pipeline as cv22_sidon. Speaker IDs are deterministic 8-character SHA256 hashes of the original Common Voice client_id.

cv22_european

A combined config of mid-size European Common Voice 22 languages, sourced from sarulab-speech/commonvoice22_sidon. Contains 591,663 utterances (~1,200 hours) across 6 languages: Italian (it), Dutch (nl), Turkish (tr), Russian (ru), Portuguese (pt), Polish (pl). Same processing pipeline as cv22_sidon. Speaker IDs are deterministic 8-character SHA256 hashes of the original Common Voice client_id.

shrutilipi

The largest config, sourced from ai4bharat/Shrutilipi. Contains 2,226,753 utterances (4,665.0 hours) across 16 languages: Assamese, Bengali, Dogri, Gujarati, Hindi, Kannada, Konkani, Maithili, Malayalam, Marathi, Nepali, Odia, Punjabi, Sanskrit, Tamil, and Telugu. No speaker metadata is available -- each row has a unique UUID as speaker_id, and all gender values are "Unknown".


Usage

Load a specific config

python
from datasets import load_dataset

ds = load_dataset("kenpath/indic-tts-unified-v1", "kathbath", split="train")
print(ds[0])
# {'audio': {'path': ..., 'array': array([...]), 'sampling_rate': 24000},
#  'text': '...', 'speaker_id': 'a1b2c3d4', 'source': 'kathbath',
#  'language': 'Tamil', 'gender': 'Female', 'duration': 5.32}

Streaming mode (recommended for large configs)

python
from datasets import load_dataset

ds = load_dataset(
    "kenpath/indic-tts-unified-v1", "shrutilipi",
    split="train", streaming=True
)

for example in ds:
    audio_array = example["audio"]["array"]
    text = example["text"]
    # Process as needed
    break

Filter by language

python
from datasets import load_dataset

ds = load_dataset(
    "kenpath/indic-tts-unified-v1", "ivr",
    split="train", streaming=True
)

hindi_ds = ds.filter(lambda x: x["language"] == "Hindi")

for example in hindi_ds:
    print(example["text"])
    break

Load multiple configs

python
from datasets import load_dataset, concatenate_datasets

configs = ["kathbath", "syspin", "rasa"]
datasets = []
for config in configs:
    ds = load_dataset(
        "kenpath/indic-tts-unified-v1", config, split="train"
    )
    datasets.append(ds)

combined = concatenate_datasets(datasets)
print(f"Combined: {len(combined)} rows")

Duration filtering

python
from datasets import load_dataset

ds = load_dataset(
    "kenpath/indic-tts-unified-v1", "syspin",
    split="train", streaming=True
)

# Keep only utterances between 1 and 30 seconds
filtered = ds.filter(lambda x: 1.0 <= x["duration"] <= 30.0)

Data Processing

The following normalization steps were applied uniformly across all source datasets during construction:

  1. 1.Audio resampling: All audio resampled to 24 kHz mono using high-quality resampling.
  2. 2.Schema alignment: Every source dataset was mapped to the unified 7-column schema described above.
  3. 3.Speaker hashing: Where speaker labels were available, they were converted to deterministic 8-character hashes for privacy and consistency. Where unavailable, random UUIDs were assigned.
  4. 4.Split merging: Train and test splits from source datasets were combined into a single train split per config.

Intended Use

This dataset is designed for:

  • Text-to-speech (TTS) model training across Indian languages
  • Automatic speech recognition (ASR) pretraining and fine-tuning
  • Speaker verification and speaker embedding research (for configs with reliable speaker IDs)
  • Multilingual and cross-lingual speech research
  • Emotion-conditioned speech synthesis (using the rasa config)

Limitations

  • Speaker IDs for shrutilipi, msft_indian, and indictts are random UUIDs and do not represent actual speaker groupings. Do not use these for speaker-level analysis.
  • Gender metadata is "Unknown" for the entire shrutilipi config and may be incomplete in other configs.
  • Duration is unfiltered. Some utterances may be very short (sub-second) or very long. Apply duration filtering for TTS training.
  • Text quality varies across sources. Some transcripts may contain noise, transliteration inconsistencies, or incomplete sentences.
  • Emotion tags in Rasa are embedded in the transcript text and need to be parsed or stripped depending on the downstream task.

Citation

If you use this dataset, please cite the original source datasets as appropriate:

  • IndicTTS: SPRINGLab/IndicTTS
  • Kathbath: ai4bharat/Kathbath
  • SYSPIN: kenpath/tts-SYSPIN
  • IndicVoices-R: ai4bharat/indicvoices_r
  • Rasa: ai4bharat/Rasa
  • Shrutilipi: ai4bharat/Shrutilipi
  • Microsoft Speech Corpus Indian: deepdml/microsoft-speech-corpus-indian
  • SpiCor: kenpath/tts-SPICOR
  • Common Voice 22 (SIDON): sarulab-speech/commonvoice22_sidon (derived from Mozilla Common Voice Corpus 22.0, CC-0)

License

Please refer to the individual source dataset licenses. This unified collection is provided under CC-BY-4.0 for the aggregation and schema normalization work. The underlying audio and text data retain the licenses of their respective sources.