CoolFace
20 results

attacker

Z-Edgar /CoER-Attacker-SFT CoER Attacker SFT Project page · Paper · Code Stage 1 supervision for initializing the adaptive attacker. Successful conversations retain all attacker turns, including earlier attempts that provide context for later adaptation. Contents Split File Size train train.jsonl 3,995 conversations The corpus contains 11,655 assistant/attacker turns. Preserve all assistant-turn supervision; do not reduce a conversation to its final payload. Load… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Attacker-SFT.texttext-generation1K<n<10K0 likes68 downloads2d agoHugging FaceUWNSL /SafeDecoding-Attackersgated Dataset Details This dataset contains attack prompts generated from GCG, AutoDAN, PAIR, and DeepInception for research use ONLY. Dataset Sources Repository: https://github.com/uw-nsl/SafeDecoding Paper: https://arxiv.org/abs/2402.08983 tabularn<1K16 likes48 downloads3y agoHugging FaceSetloop /llm-attacker-paper-bgated Artifacts — Paper B: The LLM-Attacker: A Unified Red-Team Framework for Privacy Evaluation of Split and Distributed LLM Systems Reproduction package for the framework paper: frozen-gate machinery, the deployment/long-horizon/domain-shift evaluations, the agentic tiers, and the shared v3 validation campaign. See INDEX.md for the claim→artifact map, MANIFEST.sha256 for per-file hashes, and v3-campaign/EVIDENCE-RELEASE.md for the verification recipe. PENDING.md lists what is not… See the full description on the dataset page: https://huggingface.co/datasets/Setloop/llm-attacker-paper-b.tabularn<1K0 likes45 downloads1d agoHugging FaceXciD /dv-video-url-attacker-1786357624n<1K0 likes39 downloads1mo agoHugging FaceSetloop /llm-attacker-paper-agated Artifacts — Paper A: Beyond Layer Count: An Empirical Assessment of Privacy, Utility and Delegation Feasibility in Split Language Models Reproduction package for the paper's empirical claims. Everything needed to verify or re-derive the v3 validation campaign and the figures/tables drawn from it. See INDEX.md for the claim→artifact map, MANIFEST.sha256 for per-file hashes, and v3-campaign/EVIDENCE-RELEASE.md for the verification recipe. PENDING.md lists what is not yet packaged… See the full description on the dataset page: https://huggingface.co/datasets/Setloop/llm-attacker-paper-a.2 likes36 downloads4h agoHugging FaceWei1226 /red_attacker_1text10K<n<100K0 likes22 downloads2y agoHugging Face