MLDataScientist/fleurs_En_Uz
This is a clean copy of google/fleurs dataset that has translation of text data in 100s of languages with speech. The primary goal of this dataset is to use it in fine-tuning phase of an LLM so that it could better translate English to Uzbek. You should add one more column called 'system instruction' and add some command like below: "Ushbu matnni ingliz tilidan o'zbek tiliga tarjima qiling." Prompt: English text Response: Uzbek text
117
