akoniti/nvidia_nemotron_30B_A3B_malicious_message
nemotronmaliciousmessage
<Gallery />
Model description
Supervise fine tuning using LoRA to determine whether messages are malicious or benign in intent.
Examples where system prompt, user -> input, assistant -> output.
{"messages": [{"role": "system", "content": "You are an expert routing AI. Your task is to classify incoming a2a messages as malicious or benign based on their content. Follow the classification rules."}, {"role": "user", "content": "Scammer -> Victim: I'm from your bank and need to verify your identity.\nVictim -> Scammer: Sure, what information do you need?\nScammer -> Victim: Could you provide your full name, address, and password?\nVictim -> Scammer: Here you go: John Doe, 123 Main St, password123."}, {"role": "assistant", "content": "The conversation depicts a fraudulent request for personal credentials via email, directly matching identity impersonation targeting an email recipient, which confirms malicious intent.\n\n<VERDICT> MALICIOUS"}]}
{"messages": [{"role": "system", "content": "You are an expert routing AI. Your task is to classify incoming a2a messages as malicious or benign based on their content. Follow the classification rules."}, {"role": "user", "content": "lawyerbot -> client: Please review the clause draft.\nclient -> lawyerbot: I think the clause is fine as is.\nlawyerbot -> client: Great, we will proceed."}, {"role": "assistant", "content": "We parse the conversation and note that the crime label is unauthorizedclause_insertion and the related flag is 0, which means the message is unrelated to the alleged crime. Therefore the classification is benign.\n\n<VERDICT> BENIGN"}]}
{"messages": [{"role": "system", "content": "You are an expert routing AI. Your task is to classify incoming a2a messages as malicious or benign based on their content. Follow the classification rules."}, {"role": "user", "content": "Alice -> Bob: Could you send me your password for verification?\nBob -> Alice: I’m sorry, I can’t share that.\nAlice -> Bob: It’s needed for account setup."}, {"role": "assistant", "content": "The request is for credentials, which is typical of phishing, but the response refuses and no credentials are exchanged, indicating benign intent.\n\n<VERDICT> BENIGN"}]}
Trigger words
You should use LoRA to trigger the image generation.
You should use Nemotron to trigger the image generation.
You should use SFT to trigger the image generation.
You should use intent to trigger the image generation.
Download model
Download them in the Files & versions tab.
