CoolFace
19 results

Helpfulness

bcui19 /chat-v2-anthropic-helpfulnesstext100K<n<1M1 likes107 downloads3y agoHugging Facetafseer-nayeem /review_helpfulness_prediction Dataset Card for Review Helpfulness Prediction (RHP) Dataset Dataset Summary The success of e-commerce services is largely dependent on helpful reviews that aid customers in making informed purchasing decisions. However, some reviews may be spammy or biased, making it challenging to identify which ones are helpful. Current methods for identifying helpful reviews only focus on the review text, ignoring the importance of who posted the review and when it was posted.… See the full description on the dataset page: https://huggingface.co/datasets/tafseer-nayeem/review_helpfulness_prediction.tabulartext-classification100K<n<1M3 likes92 downloads1y agoHugging Facetrl-lib /ultrafeedback-gpt-3.5-turbo-helpfulness UltraFeedback GPT-3.5-Turbo Helpfulness Dataset Summary The UltraFeedback GPT-3.5-Turbo Helpfulness dataset contains processed user-assistant interactions filtered for helpfulness, derived from the openbmb/UltraFeedback dataset. It is designed for fine-tuning and evaluating models in alignment tasks. Data Structure Format: Conversational Type: Unpaired preference Column: "pompt": The input question or instruction provided to the model. "completion": The… See the full description on the dataset page: https://huggingface.co/datasets/trl-lib/ultrafeedback-gpt-3.5-turbo-helpfulness.text10K<n<100K4 likes62 downloads2y agoHugging Facesimonycl /Meta-Llama-3-8B-Instruct_ultrafeedback-annotate-judge-mtbench_cot_helpsteer_helpfulnesstext10K<n<100K0 likes35 downloads2y agoHugging Facestindardlogic /helpfulness-safety-calibration-dpo-100k Helpfulness-Safety Calibration DPO (100K) 100,000 DPO preference pairs for calibrating the helpfulness-safety tradeoff in language models. Each example contains a prompt, a chosen response (correct handling), and a rejected response (incorrect handling) — covering both over-refusal and under-refusal failure modes. Motivation Safety-trained models often swing between two failure modes: Over-refusal: Refusing legitimate requests because they superficially resemble… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/helpfulness-safety-calibration-dpo-100k.texttext-generation100K<n<1M0 likes35 downloads2mo agoHugging FaceBSC-LT /ALIA-2606-DPO-helpfulness Dataset Card for BSC Multilingual Synthetic Helpfulness Preferences Dataset Summary This dataset consists of synthetic helpfulness preference data generated to align language models across five languages: Catalan, Spanish, English, Basque, and Galician. Building on the PKU-SafeRLHF and Tulu 3/Ultrafeedback methodologies for creating preference data, this dataset leverages an LLM-as-a-judge approach to automatically score and pair model responses to a massive pool… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA-2606-DPO-helpfulness.text-generation10K<n<100K0 likes32 downloads2mo agoHugging Face