models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
rlcd-modernbert-151mArmoRM-Llama3-8B-v0.1llama-3.1-8b-oracle-rm-hh-rlhf-harmlessnessllama-3.1-8b-oracle-rm-hh-rlhf-helpfulnessdeberta-v3-large-tasksource-rlhf-reward-modelhh_rlhf_rm_open_llama_3bWorldPM-72B-RLHFLowQwen2.5-7B-SafeRLHF-RMLlama-3.1-Tulu-3-8B-RL-RM-RB2tulu-v2.5-13b-hh-rlhf-60k-rmQwen2.5-7B-SafeRLHF-CMdistilbert-base-uncased-finetuned-clincDecision-Tree-Reward-Llama-3.1-8BRewardModel-Mistral-7B-for-DPA-v1hh-rlhf-reward-model-anyscaleDecision-Tree-Reward-Gemma-2-27BTLDR-Mistral-7B-SmallSFT-RMhh_rlhf_reward_modelprojekt_ga_tf_rltgpt2-rlhf-rewardrlhflow-llama-3-sft-8b-v2-segment-rm-700kdistilbert-base-uncased-finetuned-emotionQwen2.5-7B-hh-rlhf-RewardTLDR-Llama-3.2-1B-SmallSFT-RMrlt_2409_1450rlhf-reward-modelrlhf_dxtoicd_rewardRM_1B_iter0rl-grp-prj-per-clsrlhf_reward_model
