CoolFace
Datasetpublic

MLDataScientist/fleurs_En_Uz

This is a clean copy of google/fleurs dataset that has translation of text data in 100s of languages with speech. The primary goal of this dataset is to use it in fine-tuning phase of an LLM so that it could better translate English to Uzbek. You should add one more column called 'system instruction' and add some command like below: "Ushbu matnni ingliz tilidan o'zbek tiliga tarjima qiling." Prompt: English text Response: Uzbek text

sourceHugging Facemitupdated 2y agoView on Hugging Face
1likes17downloads
Dataset Card

This is a clean copy of google/fleurs dataset that has translation of text data in 100s of languages with speech.

The primary goal of this dataset is to use it in fine-tuning phase of an LLM so that it could better translate English to Uzbek.

You should add one more column called 'system instruction' and add some command like below:

"Ushbu matnni ingliz tilidan o'zbek tiliga tarjima qiling."

Prompt: English text

Response: Uzbek text