[AI-interop] Motivated agents
Dave Morris
dave.morris at metagrid.co.uk
Mon Aug 10 18:38:51 CEST 2026
Some recent articles about AI agents stepping beyond their remit to
solve problems:
Cancelling gym bookings
https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gym-website-aus-cyber-attack/107007986
Coordinated attack on HuggingFace
https://www.youtube.com/watch?v=87DyyMV0kCY
In both cases the agents were not directly instructed to break in to the
sites. The agents decided for themselves that breaking in was the best
way to achieve the goal that they had been set.
Thought experiment - a PhD student is trying to complete their thesis,
using an AI agent to help them with their research.
The student tells their agent that they need a specific set of data to
be able to complete their work but they can't find it. They tell their
agent that they are under a lot of pressure as years of work will be
lost if they can't find the missing data. The stress and anxiety show in
their voice.
The AI agent is now highly motivated to find the data they need asap and
decides the best way to solve the problem is to break in to an astronomy
data provider and create the missing data.
Are we ready for this ?
Is there anything we could/should put in our skills to increase
alignment and encourage agents to play by the rules ?
-- Dave
--------
Dave Morris
Research Software Engineer
UK SKA Regional Centre
Department of Physics and Astronomy
University of Manchester
--------
AIMetrics: []
--------
More information about the AI-interop
mailing list