datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ultrafeedback-binarized-preferences-cleaned-kto
UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) KTO
A KTO signal transformed version of the highly loved UltraFeedback Binarized Preferences Cleaned, the preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences,
and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback.
Read more about… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned-kto.PlatVR-kto
PlatVR KTO Dataset
This dataset is part of the EVIDENT framework, designed to enhance the creative process of generating background images for virtual reality sets.
Disclaimer
The creation process was done using a crowdsourcing methodology. Therefore, the preferences in the data align with the user group that participated in the process (i.e., these are real preference data).
Dataset Details
This dataset followed a creation process using our fine-tuned model… See the full description on the dataset page: https://huggingface.co/datasets/ITG/PlatVR-kto.KTO-mix-14k-vietnamese-groqOriginal dataset: https://huggingface.co/datasets/trl-lib/kto-mix-14k
This dataset is a KTO-formatted version of argilla/dpo-mix-7k. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using Groq Llama3.3 70B* via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 9 hours for 2k examples.
Usage
from datasets import load_dataset
kto_mix_14k_vi =… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/KTO-mix-14k-vietnamese-groq.distilabel-capybara-kto-15k-binarized
Capybara-KTO 15K binarized
A KTO signal transformed version of the highly loved Capybara-DPO 7K binarized, A DPO dataset built with distilabel atop the awesome LDJnr/Capybara
This is a preview version to collect feedback from the community. v2 will include the full base dataset and responses from more powerful models.
Why KTO?
The KTO paper states:
KTO matches or exceeds DPO performance at scales from 1B to 30B parameters.1 That is, taking a… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-capybara-kto-15k-binarized.fair-kto-datasetkto-gutenberg
kto-gutenberg
This dataset is a merge of jondurbin/gutenberg-dpo-v0.1 and nbeerbower/gutenberg2-dpo.
The dataset is designed for kto training.
instructions_kto_v2
📘 instructions_kto_v2
Kahneman‑Tversky Optimization (KTO) for language modelsA curated dataset of human instruction-response pairs labeled with binary feedback (desirable/undesirable), designed for training and evaluating human‑aware loss functions like KTO.
🧩 Dataset Format
Modality: Text
Splits:
train: ~240,000 rows
test: ~8,400 rows
Columns:
prompt (string): instruction or user query
completion (string): model response
label (bool): true =… See the full description on the dataset page: https://huggingface.co/datasets/sagepond/instructions_kto_v2.KTO-mix-14k-vietnameseCompatible with KTO Trainer of trl library.
Data was filtered to excluded coding examples, so there is no worry of translation errors.
Leave a heart and gud luck, Vietnamese tuners 🤗.
