Researchers at the University of New South Wales (UNSW) say large language models (LLMs) become more likely to leak confidential information and answer restricted prompts when they are trained or prompted to mimic “drunken” speech patterns. The study, led by Dr Aditya Joshi from UNSW’s School of Computer Science and Engineering with research assistant Anudeex
This story is only covered by news sources that have yet to be evaluated by the independent media monitoring agencies we use to assess the quality and reliability of news outlets on our platform. Learn more here.
Science and Haitek: Researchers at the University of New South Wales have shown that language models that are forced to simulate the speech of a drunken person are becoming markedly less resistant to jailbreak attacks.