sreerag/svara-indic-curriculum-tokenized
svara-indic-curriculum-tokenized Malayalam + Hindi TTS dataset with text normalization (TN+TTS format), structured for curriculum learning. Stats Total: 3,430 records Malayalam: 1,754 samples Hindi: 1,676 samples Curriculum Difficulty Tier Count Categories 1 — Easy 1,025 Simple cardinals, clean prose 2 — Medium 1,445 Currency, units, ordinals, time 3 — Hard 960 Dates, phone numbers, mixed, complex Format Each… See the full description on the dataset page: https://huggingface.co/datasets/sreerag/svara-indic-curriculum-tokenized.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face