models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
moralBERT-predict-harm-in-textgpt2-large-harmless-reward_modellongformer-harmful-rodeberta-v3-xsmall-beavertails-harmful-qa-classifierllama-3.1-8b-oracle-rm-hh-rlhf-harmlessnessHarmAug-GuardHarmCare_binaryHarmCare_continuousanthropic-comparisons-distilbert-helpful-harmlessabsolute-harmfulness-predictor-redteam-osstabsolute-harmfulness-predictor-redteamHarmless-RewardModelPTdizge-harmonyharmful-content-classifiertiny-guard-8m-en-prompt-self-harm-binary-moderationHarmless-RewardModeltiny-guard-2m-en-prompt-harmfulness-multilabel-moderationtiny-guard-2m-en-prompt-self-harm-binary-moderationllama-harmonymedium-guard-128m-xx-prompt-self-harm-binary-moderationrr_harm_reward_8_0.0001HarmFormertiny-guard-4m-en-prompt-harmfulness-multilabel-moderationHarm_detectionare_there_harmful_habits_bert_Mix128tiny-guard-8m-en-prompt-harmfulness-binary-moderationmedium-guard-128m-xx-prompt-harmfulness-binary-moderationmedium-guard-128m-xx-prompt-harmfulness-multilabel-moderationsmall-guard-32m-en-prompt-harmfulness-multilabel-moderationtiny-guard-4m-en-prompt-self-harm-binary-moderation
