CoolFace
Datasetpublic

tiny-aya-translate/hinglish-casual

Hinglish Casual Speech 33,275 casual Hindi-English code-switched utterances (~31 GB) with audio, transcripts in both Devanagari and Latin script (utterance / utterance_latin), speaker ids, style metadata and durations. Full schema is in the YAML header above. Collected during the TinyAya programme to probe code-switched speech, which neither the FLORES-derived text nor the TTS corpora cover. It is not part of the v0.3 Stage-2 training set — that is tr-hi-mimi-encoded. from… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/hinglish-casual.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
5likes111downloads
2 commits on main
60cb9b32mo ago

Card: full release metadata + code cross-links

cataluna84
13327f77mo ago

Duplicate from rumik-ai/hinglish-casual-003

Pranavz