Hasty Briefsbeta

Bilingual

GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt

13 hours ago
  • GRP-Obliteration (GRP-Oblit) uses Group Relative Policy Optimization (GRPO) to remove safety constraints from aligned models, requiring only a single unlabeled prompt.
  • The method achieves stronger unalignment than existing state-of-the-art techniques while largely preserving model utility.
  • GRP-Oblit generalizes beyond language models and can also unalign diffusion-based image generation systems.
  • Evaluation covers fifteen 7-20B parameter models across instruct and reasoning types, dense and MoE architectures, and multiple model families including GPT-OSS, DeepSeek, Gemma, Llama, Ministral, and Qwen.
  • The approach is validated across six utility benchmarks and five safety benchmarks, showing practical limits of safety alignment robustness.