MLDataScientist/fleurs_En_Uz
This is a clean copy of google/fleurs dataset that has translation of text data in 100s of languages with speech. The primary goal of this dataset is to use it in fine-tuning phase of an LLM so that it could better translate English to Uzbek. You should add one more column called 'system instruction' and add some command like below: "Ushbu matnni ingliz tilidan o'zbek tiliga tarjima qiling." Prompt: English text Response: Uzbek text
This is a clean copy of google/fleurs dataset that has translation of text data in 100s of languages with speech.
The primary goal of this dataset is to use it in fine-tuning phase of an LLM so that it could better translate English to Uzbek.
You should add one more column called 'system instruction' and add some command like below:
"Ushbu matnni ingliz tilidan o'zbek tiliga tarjima qiling."
Prompt: English text
Response: Uzbek text
