datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KTO-mix-14k-vietnamese-groqOriginal dataset: https://huggingface.co/datasets/trl-lib/kto-mix-14k
This dataset is a KTO-formatted version of argilla/dpo-mix-7k. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using Groq Llama3.3 70B* via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 9 hours for 2k examples.
Usage
from datasets import load_dataset
kto_mix_14k_vi =… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/KTO-mix-14k-vietnamese-groq.kto-gutenberg
kto-gutenberg
This dataset is a merge of jondurbin/gutenberg-dpo-v0.1 and nbeerbower/gutenberg2-dpo.
The dataset is designed for kto training.
KTO-mix-14k-vietnameseCompatible with KTO Trainer of trl library.
Data was filtered to excluded coding examples, so there is no worry of translation errors.
Leave a heart and gud luck, Vietnamese tuners 🤗.
