datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orpo-dsorpo-dpo-mix-40k
ORPO-DPO-mix-40k v1.2
This dataset is designed for ORPO or DPO training.
See Fine-tune Llama 3 with ORPO for more information about how to use it.
It is a combination of the following high-quality DPO datasets:
argilla/Capybara-Preferences: highly scored chosen answers >=5 (7,424 samples)argilla/distilabel-intel-orca-dpo-pairs: highly scored chosen answers >=9, not in GSM8K (2,299 samples)
argilla/ultrafeedback-binarized-preferences-cleaned: highly scored chosen answers >=5 (22… See the full description on the dataset page: https://huggingface.co/datasets/mlabonne/orpo-dpo-mix-40k.orpo-dpo-mix-40k-flat
Dataset Card for "orpo-dpo-mix-40k-flat"
More Information needed
med-qa-orpo-dpo
MED QA ORPO-DPO Dataset
This dataset is restructured from several existing datasource on medical literature and research, hosted here on hugging face. The dataset is shaped in question, choosen
and rejected pairs to match the ORPO-DPO trainset requirements.
Features
The dataset consists of the following features:
question: MCQ or yes/no/maybe based questions on medical questions
direct-answer: correct answer to the above question
chosen: the correct answer along with… See the full description on the dataset page: https://huggingface.co/datasets/empirischtech/med-qa-orpo-dpo.LVT-Audio-ORPO-DataGerman-RAG-ORPO-ShareGPT-HESSIAN-AI
German-RAG-ORPO (Odds Ratio Preference Optimization) ShareGPT-Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The ORPO Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities.
The subsets can be for this training step are derived from 3 different sources:
SauerkrautLM Preference Datasets:
SauerkrautLM-Fermented-GER-DPO: is a specialized dataset designed for training… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-ORPO-ShareGPT-HESSIAN-AI.German-RAG-ORPO-Alpaca-HESSIAN-AI
German-RAG-ORPO (Odds Ratio Preference Optimization) Alpaca-Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The ORPO Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities.
The subsets can be for this training step are derived from 2 different sources:
SauerkrautLM Preference Datasets:
SauerkrautLM-Fermented-GER-DPO: is a specialized dataset designed for training… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-ORPO-Alpaca-HESSIAN-AI.Colossal_Translation_Spanish_to_English_AND_English_to_Spanish_ORPO_DPO_Gemma
original dataset
https://huggingface.co/datasets/Iker/Colossal-Instruction-Translation-EN-ES
take dataset
NickyNicky/Iker-Colossal-Instruction-Translation-EN-ES_deduplicated_length_3600
Orpo_Split_Turns_v3_clean_metaalpaca-vs-alpaca-orpo-dpo
Alpaca vs. Alpaca
Dataset Description
The Alpaca vs. Alpaca dataset is a curated blend of the Alpaca dataset and the Alpaca GPT-4 dataset, both available on HuggingFace Datasets. It uses the standard GPT dataset as the 'rejected' answer, steering the model towards the GPT-4 answer, which is considered as the 'chosen' one.
However, it's important to note that the 'correctness' here is not absolute. The premise is based on the assumption that GPT-4 answers are generally… See the full description on the dataset page: https://huggingface.co/datasets/efederici/alpaca-vs-alpaca-orpo-dpo.OpenHermesPreferences-500ksocial-i-qa-orpo-dpo-10kHuggingFaceH4__zephyr-orpo-141b-A35b-v0.1-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-orpo-141b-A35b-v0.1-details.Colossal_Translation_EN_ES_ORPO_DPO_Gemma
original dataset
https://huggingface.co/datasets/Iker/Colossal-Instruction-Translation-EN-ES
take dataset
NickyNicky/Iker-Colossal-Instruction-Translation-EN-ES_deduplicated_length_3600
orpo-dpo-mix-40k_portuguese
Portuguese translation of mlabonne/orpo-dpo-mix-40k
How it was made?
I have used a ctranslate2 version of google/madlad400-3b-mt quantized in int8_f16
Took more than a week to translate all in a single thread using RTX 3060 12GB
Whats next?
Using the same model to translate cognitivecomputations/Wizard-Vicuna-7B-Uncensored
But now I'm using 3 threads, 1 model each.
Plans
I'm focusing in datasets that can be useful for my future finetunes.
The goal… See the full description on the dataset page: https://huggingface.co/datasets/BornSaint/orpo-dpo-mix-40k_portuguese.agillm43-orpo-pairs
AGILLM4.3 ORPO preference pairs
Synthetic {prompt, chosen, rejected} preference pairs for single-stage ORPO
post-training of the AGILLM4.3 base model. Generated with GLM-5.2.
Two preference axes:
Correctness — chosen is right, rejected is fluent-but-wrong. Categories:
capitals, factual QA, arithmetic, worked-step math, science, commonsense,
definitions, short reasoning. Targets the base model's measured weaknesses
(e.g. the "capital of X -> Paris" over-association; word-salad… See the full description on the dataset page: https://huggingface.co/datasets/MarxistLeninist/agillm43-orpo-pairs.PKU-SafeRLHF-orpo-72kWarning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful.
👇original PKU-SafeRLHF datasets (click 🔗 for more details)
what's the advantage of this train dataset over the original one ?
standard chosen/rejected format of preference datasets : make 'chosen' and 'rejected' according to 'better_response_id'
only one file : merge three train datasets(Alpaca-7B、Alpaca2-7B、Alpaca3-8B)… See the full description on the dataset page: https://huggingface.co/datasets/juneup/PKU-SafeRLHF-orpo-72k.orpo-dpo-mix-TR-20k
ORPO-DPO-Mix-TR-20k
This repository contains a Turkish translation of 20k records from the mlabonne/orpo-dpo-mix-40k dataset. The translation was carried out using the gemini-1.5-flash-002 model, chosen for its 1M token context size and overall accurate Turkish responses.
Translation Process
The translation process uses the LLM model with translation prompt and pydantic data validation to ensure accuracy. The complete translation pipeline is available in the GitHub… See the full description on the dataset page: https://huggingface.co/datasets/selimc/orpo-dpo-mix-TR-20k.orpo_test2Multifaceted-Collection-ORPO
Dataset Card for Multifaceted Collection ORPO
Links for Reference
Homepage: https://lklab.kaist.ac.kr/Janus/
Repository: https://github.com/kaistAI/Janus
Paper: https://arxiv.org/abs/2405.17977
Point of Contact: suehyunpark@kaist.ac.kr
TL;DR
Multifaceted Collection is a preference dataset for aligning LLMs to diverse human preferences, where system messages are used to represent individual preferences. The instructions are acquired from five existing… See the full description on the dataset page: https://huggingface.co/datasets/kaist-ai/Multifaceted-Collection-ORPO.medical-o1-reasoning-SFT-orpodfurman__CalmeRys-78B-Orpo-v0.1-details
Dataset Card for Evaluation run of dfurman/CalmeRys-78B-Orpo-v0.1
Dataset automatically created during the evaluation run of model dfurman/CalmeRys-78B-Orpo-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dfurman__CalmeRys-78B-Orpo-v0.1-details.social-i-qa-orpo-dpo-27kGerman-RAG-ORPO-Long-Context-ShareGPT-HESSIAN-AI
German-RAG-ORPO (Odds Ratio Preference Optimization) Long Context ShareGPT-Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The ORPO Long Context Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities.
The subsets are derived from Synthetic generation inspired by Tencent's (“Scaling Synthetic Data Creation with 1,000,000,000 Personas”).
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-ORPO-Long-Context-ShareGPT-HESSIAN-AI.Danielbrdz__Barcenas-Llama3-8b-ORPO-details
Dataset Card for Evaluation run of Danielbrdz/Barcenas-Llama3-8b-ORPO
Dataset automatically created during the evaluation run of model Danielbrdz/Barcenas-Llama3-8b-ORPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Danielbrdz__Barcenas-Llama3-8b-ORPO-details.Nekochu__Llama-3.1-8B-German-ORPO-details
Dataset Card for Evaluation run of Nekochu/Llama-3.1-8B-German-ORPO
Dataset automatically created during the evaluation run of model Nekochu/Llama-3.1-8B-German-ORPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nekochu__Llama-3.1-8B-German-ORPO-details.oasst2_orpo_mix_tokenizer_phi_3_v1
https://huggingface.co/datasets/NickyNicky/orpo-dpo-mix-54k
orpo-dpo-mix-40k-flat-mlx
Dataset Description
This dataset is a split version of orpo-dpo-mix-40k-flat for direct use with mlx-lm-lora, specifically tailored to be compatible with DPO and CPO training.
The dataset has been divided into three parts:
Train Set: 90%
Validation Set: 6%
Test Set: 4%
Example Usage
To train a model using this dataset, you can use the following command:
mlx_lm_lora.train \
--model Qwen/Qwen2.5-3B-Instruct \
--train \
--test \
--num-layers 8 \
--data… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/orpo-dpo-mix-40k-flat-mlx.orpo_mixed_enanakin87__gemma-2b-orpo-details
Dataset Card for Evaluation run of anakin87/gemma-2b-orpo
Dataset automatically created during the evaluation run of model anakin87/gemma-2b-orpo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/anakin87__gemma-2b-orpo-details.
