A team of researchers has identified a core architectural flaw in large language models (LLMs) that renders them permanently vulnerable to exploitation. According to a paper presented at a recent AI conference, the issue stems from how these models distinguish between instructions and data, making it impossible to fully secure them against malicious actors.

Advertisement

By exploiting this vulnerability, researchers demonstrated that they could bypass safety guardrails, forcing popular models to generate prohibited content, such as instructions for synthesizing illegal substances or sabotaging critical infrastructure like aircraft navigation systems. The findings suggest that this security challenge may be an intractable feature of current LLM technology.

Source: MIT Technology Review