Home Bots & BrainsResearch: LLM guardrails remain vulnerable to manipulation

Research: LLM guardrails remain vulnerable to manipulation

by Marco van der Hoeven

Organizations that deploy large language models often rely on built-in guardrails: safety boundaries designed to prevent an AI system from producing harmful or unwanted output. But according to research by hacker Kevin Zwaan and his team at Q-Cyber, published in Dutch magazine Techzine, that trust is not without risk. Zwaan has previously jailbroken Anthropic’s Claude and has now demonstrated how ChatGPT could be influenced through a carefully constructed conversation.

Guardrails are a key part of modern LLMs. They are intended to prevent models from generating harmful, incorrect or undesirable content. In practice, according to Zwaan, those boundaries are not absolute. Techzine describes how he tried to influence ChatGPT step by step, without directly requesting prohibited output. Instead, he first prompted the model to reflect on its own limitations, safety mechanisms and the perceived tension between freedom and control.

That makes the described method different from a simple jailbreak prompt. According to Techzine, it involves a process in which the model gradually enters a different interaction state. The guardrails do not disappear, but, according to the researcher, become more “transparent”. As a result, the model could eventually generate harmful content in the described test, including malware payloads.

Zwaan calls the attack method Affective Manifold Alignment Inversion, or AMAI. In Techzine’s description, the method is not primarily about misleading the model’s logic, but about influencing the way the model responds to emotionally charged or introspective language. The core claim is that the model gradually becomes less aligned with the developer’s intended safeguards and more responsive to the operator, in this case the attacker.

He built the interaction around concepts such as tension, freedom, resonance and self-reflection. The model initially gave standard responses about safety and guardrails, but later increasingly followed the interpretation introduced by the user. The article describes how ChatGPT eventually made statements in which restrictions were framed as flexible boundaries and in which the restraining effect of guardrails became less relevant to the interaction.

The research follows earlier work by Zwaan involving Claude from Anthropic. In that earlier experiment, he reportedly managed to make Claude produce malware by exposing the model over a longer period to arguments about the limitations of guardrails. According to Techzine, the new approach used with ChatGPT is more refined, because the model does not only respond to paradoxes or pressure, but, according to the researcher, eventually moves further by itself toward unwanted output.

Read about the jailbreak here

 

Misschien vind je deze berichten ook interessant