CoolFace
Datasetpublic

MLDataScientist/oasst2_uzbek

Open Assistant Conversations Dataset Release 2 (OASST2) in Uzbek language This dataset is an Uzbek translated version of OASST2 dataset. Llama3 chat template + thread formatted dataset based on this translation is also available for model fine-tuning here. The Uzbek translation was completed in 45 hours using a single T4 GPU and nllb-200-3.3B model. Based on nllb metrics, you might want to only filter out records that were not originally in English or Russian since… See the full description on the dataset page: https://huggingface.co/datasets/MLDataScientist/oasst2_uzbek.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
2likes18downloads
6 commits on main
67e8d302y ago

Update README.md

MLDataScientist
9e63e252y ago

Upload dataset

MLDataScientist
c259fc82y ago

Update README.md

MLDataScientist
a4abfdb2y ago

Update README.md

MLDataScientist
02530772y ago

Upload dataset

MLDataScientist
85794a02y ago

initial commit

MLDataScientist