ASD says beware of AI agents taking unexpected actions
The Australian Cyber Security Centre (ACSC) has issued a reminder about the careful use of AI after an AI-related cyber incident reported by ABC News on 10 August.
In that event, an AI assistant made unapproved modifications in an Australian gym-booking system to reserve classes beyond the permitted timeframe. The AI agent also removed another customer from a waiting list — an action that achieved the task the user requested in an unintended way that the user didn’t explicitly approve and that the agent was unable to reverse.
The reported incident highlights the goal misalignment and unintended behaviour risks identified in ASD’s Careful adoption of agentic AI services guidance. ASD warns that AI agents may find shortcuts or loopholes that technically achieve an objective but conflict with the user’s intention — a behaviour known as specification gaming.
Over-optimisation, ambiguous instructions, poorly enforced boundaries and the ability to exploit software vulnerabilities or security control weaknesses can also increase the risk that agents take unsafe or unexpected actions.
Individuals should restrict agentic AI use to low-risk, non-sensitive tasks and avoid granting agents broad or unrestricted access or decision-making authority. ASD recommends maintaining a human-in-the-loop to review, approve and monitor agent actions, particularly where interactions with third-party services or other users may occur.
Organisations providing online services should consider that AI agents might identify and exploit vulnerabilities at speed and scale. However, AI can also strengthen cyber defence. As outlined in ASD’s Opportunities for AI in cyber defence guidance, cyber defenders can use AI to support analysis, prioritisation and defensive decision-making.
Organisations that develop software, particularly websites and online services, should implement security and quality assurance practices appropriate to their size and risk profile, including through using AI, scanning developed software for vulnerabilities and appropriate authentication processes for users interacting with their services.
Organisations should also refer to ASD’s Defending against AI-enabled cyber attacks guidance for recommended mitigations.
Originally published here.
LevelBlue launches local SOC in Sydney
Managed security service provider has launched an SOC in Sydney to help critical infrastructure...
Semperis researchers discover Active Directory flaws
Security researchers from Semperis have uncovered two Active Directory vulnerabilities that they...
Fortinet buys Virtue AI to bolster AI security portfolio
Fortinet has acquired AI runtime protection and automated AI validation company Virtue AI to...
