datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SlimOrca-Dedup-English-UzbekThis is an Uzbek translated version of https://huggingface.co/datasets/Open-Orca/SlimOrca-Dedup.
It is a single parquet file.
Check here for cleaned Uzbek only slim Orca dataset: https://huggingface.co/datasets/MLDataScientist/SlimOrca-Dedup-Uzbek-cleaned
oasst2_uzbek
Open Assistant Conversations Dataset Release 2 (OASST2) in Uzbek language
This dataset is an Uzbek translated version of OASST2 dataset.
Llama3 chat template + thread formatted dataset based on this translation is also available for model fine-tuning here.
The Uzbek translation was completed in 45 hours using a single T4 GPU and nllb-200-3.3B model.
Based on nllb metrics, you might want to only filter out records that were not originally in English or Russian since… See the full description on the dataset page: https://huggingface.co/datasets/MLDataScientist/oasst2_uzbek.
