models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
CON-PPO-TEST2besstie-sarcasm-deberta-v3critic_600_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-1000-f553c1779bRLHF-PPO-RewardModel-LLama3-3B-v2RLHF-PPO-RewardModel-LLama3-1B-v1PPOCR_v5ppo-reward-model-bertPPOCR_v6RLHF-PPO-RewardModel-LLama3-1B-v1.1pythia_tldr_ppo_1b_valueppo-reward-modelfinetuned-amazon_reviews_multihuggingface_trainllm-course-hw2-ppocritic_250_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-500-6394f91930critic_250_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-500-517c4b730fQwen-1.5B-Instruct-ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-sub-8cd26db347critic_16_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-5000-0f12cee30ecritic_200_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-1000-c08ed8b533llama-3.2-3b_ppo_lr5e-07_rm_data-mix_no_sys_msg_unfilteredcritic_800_Qwen-1.5B-Instruct-ppo-run-math-training-prompt-len-800-response-len-4096-0c4e7a4810critic_450_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-500-26d437d2d8critic_400_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-1000-6314c2edc2critic_50_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-500-96479447bbcritic_16_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-500-a-9a44e3cd58critic_450_ppo-run-math-training-prompt-len-800-response-len-4096-bce-loss-temperatur-9fe16df365critic_1200_ppo-run-math-training-prompt-len-800-response-len-4096-bce-loss-temperatu-6aa1e360d1critic_50_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-5000-ccb6db349dcritic_600_deepseek-r1-distil-1.5b-ppo-run-math-training-prompt-len-800-response-len-07fa1b4078besstie-sentiment-deberta-v3
