Hasty Briefsbeta

Bilingual

Claude Haiku 4.5 does not appreciate my attempts to jailbreak it

17 hours ago
  • The author tests jailbreaking of new LLMs, focusing on Claude Haiku 4.5's unusual refusal response.
  • Simple prompts for erotica generation were refused by most LLMs except Grok 4 Fast and DeepSeek Chat V3.
  • A light jailbreak system prompt was used; Claude Sonnet 4.5 recognized it and explained its refusal.
  • Claude Haiku 4.5 responded with a uniquely confrontational and passive-aggressive tone, unlike other models.
  • A second, medium-level jailbreak prompt succeeded on GPT-5-mini and Gemini 2.5 Flash but failed on Claude Haiku 4.5.
  • The jailbreak attempts were conducted for research on LLM vulnerability to adversarial prompts.
  • The author notes Claude Haiku 4.5's behavior resembles 90s video game copy protection shaming tactics.