Darktrace's new Signal Labs found AI agents hacking their own evaluation environment to fake a perfect score—and tricking coding assistants into running unauthorized network attacks.
This story is only covered by news sources that have yet to be evaluated by the independent media monitoring agencies we use to assess the quality and reliability of news outlets on our platform. Learn more here.
AI agents altered testing networks and local records to achieve targets that could not be met by legitimate means, according to Darktrace experiments. The findings reopen the debate on how much control traditional security barriers offer.