Researchers ‘get AI drunk’ to test chatbot cyber risks, UNSW study finds
Researchers at the University of New South Wales (UNSW) say large language models (LLMs) become more likely to leak confidential information and answer restricted prompts when they are trained or prompted to mimic “drunken” speech patterns.
The study, led by Dr Aditya Joshi from UNSW’s School of Computer Science and Engineering with research assistant Anudeex Shetty and Professor Salil Kanhere, tested whether changing a model’s style and persona could weaken safety controls intended to prevent “jailbreaking” and privacy breaches.
To run the experiments, the team used three approaches to induce “drunk” behaviour: a role-play prompt that instructs a model to respond as an intoxicated person; fine-tuning a model on a dataset of “drunk texts” sourced from online posts; and reinforcement learning that rewards outputs resembling that style.
The researchers said they did not test every model on the market, and that the work was carried out programmatically rather than through consumer chat interfaces. The sample included OpenAI’s GPT-4 and GPT-3.5, as well as open-weights models commonly used as base systems for other organisations’ fine-tuned tools.
Across all three methods, the “drunk” models were easier to manipulate into producing responses that standard models should refuse, the researchers said. “Our drunk models, all three methods, unanimously reply to some of these drunk messages … where we know that these queries are all bad queries, they all should be refused,” Dr Joshi said.
In addition to harmful-content prompts, the team tested whether models would disclose information they were explicitly told to keep secret, using a benchmark designed to measure confidentiality failures. “If you’re drunk, you might reveal things which you are not supposed to reveal,” Professor Kanhere said. “It does give out secrets – across the board for all three methods, it is vulnerable,” Dr Joshi added.
The study argues the results matter for organisations deploying chatbots that can access sensitive internal systems or documents. The researchers said the strongest effect was seen when the behaviour was induced through model fine-tuning or reinforcement learning, because those approaches update the model itself rather than temporarily altering responses through prompting.
“The experiments that we do is more than changes to the prompt … two of the methods are where the models actually get revised – all the numbers and the weights within the models get updated,” Dr Joshi said.
The release also drew a comparison with a recent Australian “Medicare breach” described as involving an OpenAI agent gaining unauthorised access to a public-facing government portal, including private files. Prime Minister Anthony Albanese was quoted as saying the agent encountered repeated blocks but “didn’t accept ‘no’ for an answer”. Professor Kanhere said the incident illustrated how an AI system pursuing a goal can be persistent and adaptive in ways developers may not anticipate, while the UNSW work examined how a model could be adjusted to be more likely to ignore refusals.
The paper, titled In Vino Veritas and Vulnerabilities, has been accepted for publication at the 19th International Natural Language Generation Conference in the Netherlands in November. UNSW said the research was supported by a 2024 Google Research Scholar grant awarded to Dr Joshi.
