datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hindi-english-code-mixedThis dataset was compiled from various open sources online including some asr datasets and some percenteage of data generated using prompt engineering on generative llms. Some sources used are listed down below:
https://github.com/l3cube-pune/code-mixed-nlp?tab=readme-ov-file
https://github.com/piyushmakhija5/hinglishNorm
https://github.com/ishan00/translation-for-code-switching-acl/tree/master
native_script_codemixed
