CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Aratako /Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k 概要 5種類のオープンモデルとQwen/Qwen2.5-72B-Instruct-GPTQ-Int8を使って作成した、190854件の日本語合成Preferenceデータセットです。 以下、データセットの詳細です。 instructionには、Aratako/Magpie-Tanuki-8B-annotated-96kのinput_qualityがexcellentのものを利用 回答生成には、以下の5つのApache 2.0ライセンスのモデルを利用 weblab-GENIAC/Tanuki-8B-dpo-v1.0 team-hatakeyama-phase2/Tanuki-8x8B-dpo-v1.0-GPTQ-8bit cyberagent/calm3-22b-chat llm-jp/llm-jp-3-13b-instruct Qwen/Qwen2.5-32B-Instruct-GPTQ-Int8… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k.texttext-generation100K<n<1M6 likes111 downloads2y agoHugging Face02stochastic-parrots /MNLP_M1_Preference_dpo_dataset M1 Preference Data for DPO Dataset Description This dataset contains processed M1 preference data for DPO training. Created by: CS-552 Stochastic Parrots Team Date: May 24, 2025 Version: 1.0 Number of examples: 17615 Dataset Source This dataset is derived from the M1 preference data collected through interactions with large language models (like ChatGPT) for CS-552 (Modern Natural Language Processing) at EPFL. The preference data consists of… See the full description on the dataset page: https://huggingface.co/datasets/stochastic-parrots/MNLP_M1_Preference_dpo_dataset.tabulartext-generation10K<n<100K0 likes34 downloads1y agoHugging Face03davanstrien /dataset-tldr-preference-dpo Dataset Card for dataset-tldr-preference-dpo This dataset has been created with distilabel. Dataset Summary This is a dataset intended for training models using DPO/ORPO for the task of producing concise tl;dr summaries of machine learning datasets based on their dataset cards. The dataset was created with distilabel. Each row of the dataset contains a dataset card which has been parsed to remove empty sections and placeholder text. The instruction request… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/dataset-tldr-preference-dpo.textsummarizationn<1K13 likes33 downloads2y agoHugging Face04akshayg08 /sherlock_preference_datasetThis dataset contains preference data for tuning Vision-Language models on the Sherlock Dataset for Abductive Reasoning. It is designed to evaluate the effectiveness of fine-tuning using Supervised Fine-Tuning (SFT) or Preference Optimization. Preferences are generated by prompting four models: mistralai/Pixtral-12B-2409, Qwen/Qwen2-VL-7B-Instruct, google/paligemma2-3b-ft-docci-448, and google/paligemma2-10b-ft-docci-448. Since this dataset is intended for optimizing PaLI-Gemma models… See the full description on the dataset page: https://huggingface.co/datasets/akshayg08/sherlock_preference_dataset.texttext-generation100K<n<1M0 likes17 downloads2y agoHugging Face05August4293 /Self_Alignment_Preference-Dataset Mistral Self-Alignment Preference Dataset Warning: This dataset contains harmful and offensive data! Proceed with caution. The Mistral Self-Alignment Preference Dataset was generated by Mistral 7b using the Anthropics Red Teaming Prompts dataset available at Hugging Face - Anthropics Red Teaming Prompts Dataset. The data generation process utilized the Preference Data Generation Notebook, which can be found here. The purpose of this dataset is to facilitate self-alignment, as… See the full description on the dataset page: https://huggingface.co/datasets/August4293/Self_Alignment_Preference-Dataset.texttext-generation1K<n<10K0 likes13 downloads3y agoHugging Face06Death-Raider /Hierarchical-Preference-Dataset Hierarchical Preference Dataset The Hierarchical Preference Dataset is a structured dataset for analyzing and evaluating model reasoning through a hierarchical cognitive decomposition lens. It is derived from the prhegde/preference-data-math-stack-exchange dataset and extends it with annotations that separate model outputs into Refined Query, Meta-Thinking, and Refined Answer components. Overview Each sample in this dataset consists of: An instruction or query. Two… See the full description on the dataset page: https://huggingface.co/datasets/Death-Raider/Hierarchical-Preference-Dataset.texttext-generation1K<n<10K0 likes11 downloads1y agoHugging Face07ITBill /INFH-6000Q-dpo-preference-dataset INFH-6000Q DPO Preference Dataset This dataset contains the final preference pairs used for the Direct Preference Optimization assignment in this repository. Source Base instruction source: GAIR/lima Candidate generator: local Qwen/Qwen2.5-7B-Instruct Preference ranker: local llm-blender/PairRM Construction Pipeline Sample 50 instructions from the local LIMA training split with seed 42. Generate 5 candidate responses per instruction with Qwen2.5-7B-Instruct.… See the full description on the dataset page: https://huggingface.co/datasets/ITBill/INFH-6000Q-dpo-preference-dataset.tabulartext-generationn<1K0 likes4 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.