models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
GPT2-124-poetry-RLHF-GGUFLinkbricks-Horizon-AI-Korean-llama3.1-sft-rlhf-dpo-8BRLHF-VRLHF-V-SFTtulu-v2.5-dpo-13b-hh-rlhf-60ktulu-v2.5-ppo-13b-hh-rlhf-60ktulu-v2.5-dpo-13b-hh-rlhfgpt2-open-instruct-v1-Anthropic-hh-rlhfbloomz-rlhfchat-opt-1.3b-rlhf-critic-deepspeedQwen3-1.7B-DPO-hh-rlhfchat-opt-1.3b-rlhf-actor-deepspeedhh-rlhf-sftchinese-alpaca-2-1.3b-rlhfNV-Llama2-13B-RLHF-RMpythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-modelchat-opt-1.3b-rlhf-actor-ema-deepspeeddeepspeed-chat-step3-rlhf-actor-model-opt1.3bMetaAligner-HH-RLHF-7BDDeduPModelv7-RLHF-Merged-16bitDDeduPModelv7-RLHFv2Phi4-rlhf-trained-specializedOpenBezoar-HH-RLHF-SFTMetaAligner-HH-RLHF-1.1BOPT-1.3B-RLHF-DSChatLoRAOpenBezoar-HH-RLHF-DPOPythia-2.8B-HH-RLHF-Iterative-SamPOllama-3-8b-Instruct_ftjob-2581e9f8d338arco-rlhf-test-2MetaAligner-HH-RLHF-13B
