models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
ArmoRM-Llama3-8B-v0.1rlcd-modernbert-151mllama-3.1-8b-oracle-rm-hh-rlhf-harmlessnessllama-3.1-8b-oracle-rm-hh-rlhf-helpfulnessHugston_code-rl-Qwen3-4B-Instruct-2507-SFT-30bdeberta-v3-large-tasksource-rlhf-reward-modelhh_rlhf_rm_open_llama_3bqwen3-0.6b-rlcd-decisionWorldPM-72B-RLHFLowQwen2.5-7B-SafeRLHF-RMLlama-3.1-Tulu-3-8B-RL-RM-RB2tulu-v2.5-13b-hh-rlhf-60k-rmQwen2.5-7B-SafeRLHF-CMdistilbert-base-uncased-finetuned-clincRewardModel-Mistral-7B-for-DPA-v1Decision-Tree-Reward-Llama-3.1-8Bhh-rlhf-reward-model-anyscaleDecision-Tree-Reward-Gemma-2-27BTLDR-Mistral-7B-SmallSFT-RMdecision-head-qwen3.5-4b-rlcd-32khh_rlhf_reward_modelprojekt_ga_tf_rltgpt2-rlhf-rewardrlhflow-llama-3-sft-8b-v2-segment-rm-700kdistilbert-base-uncased-finetuned-emotionQwen2.5-7B-hh-rlhf-Rewardyelp-review-quality-v2rlhf-reward-modelrlhf_dxtoicd_rewardTLDR-Llama-3.2-1B-SmallSFT-RM
