datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
triple-preference-ultrafeedback-40K
Dataset Card for llama3-ultrafeedback-armorm
This dataset was used to train tpo-alignment/Llama-3-8B-TPO-L-40k, tpo-alignment/Llama-3-8B-TPO-40k, and tpo-alignment/Mistral-7B-TPO-40k.
Dataset Creation
This dataset is built based on the UltraFeedback. We reconstruct UltraFeedback to select three preferences per prompt. First, we rank the responses based on the scores provided in the base dataset. The highest-scoring response is selected as the reference, the… See the full description on the dataset page: https://huggingface.co/datasets/tpo-alignment/triple-preference-ultrafeedback-40K.BeaverTails-single-dimension-preferencepreference_alignment_ultra_cutq-alignment-dynamic-preference-datapreference_alignment_totalq-alignment-preference-data-v5grpo-q-alignment-preference-datafiltered-final-q-alignment-preference-data-th65q-alignment-preference-data-v2food-preference-generalizationverified-q-alignment-dynamic-preference-datapreference_alignment_oasstInterMT-Global-Preferencefinal-q-alignment-preference-dataSelf_Alignment_Preference-Dataset
Mistral Self-Alignment Preference Dataset
Warning: This dataset contains harmful and offensive data! Proceed with caution.
The Mistral Self-Alignment Preference Dataset was generated by Mistral 7b using the Anthropics Red Teaming Prompts dataset available at Hugging Face - Anthropics Red Teaming Prompts Dataset. The data generation process utilized the Preference Data Generation Notebook, which can be found here.
The purpose of this dataset is to facilitate self-alignment, as… See the full description on the dataset page: https://huggingface.co/datasets/August4293/Self_Alignment_Preference-Dataset.banking_alignment_preference_dsfiltered-final-q-alignment-preference-data-th75grpo-q-alignment-preference-data-bon-correct-selectionq-alignment-preference-data-v3filtered-final-q-alignment-preference-dataq-alignment-preference-dataq-alignment-preference-data-v4han-human-preference-alignment-v1
Humanoid Human Preference Alignment Dataset
This dataset captures structured human preference signals
used to align humanoid AI behavior with individual needs.
Use Cases
Personalized interaction
Alignment training
Adaptive response tuning
Fields
human_id
preference_category
preference_value
priority_level
confidence_score
Part of
Humanoid Network (HAN)
License
MIT
verified-q-alignment-dynamic-preference-data-cur-scoresft-inu-preference-alignmentInterMT-Global-Preference0910-tv2t-preferencehello
inu-preference-alignment
