The Hidden Instructions That Can Hijack AI Agents
Malicious prompts concealed in documents, metadata, emails, images and code can manipulate autonomous agents into taking dangerous actions. The post The Hidden Instructions That Can Hijack AI Agents appeared first on SecurityWeek.
They cannot be seen, can be tailored to different purposes and once adopted they operate at lightning speed.
Hidden AI prompt injections are synonymous with indirect prompts but with the specific quality of being hidden from human overview. Unlike traditional prompt injection attacks, where a user directly attempts to manipulate an AI chatbot, indirect prompt injection targets the information AI agents ingest. Bowbridge fears they are a growing risk to autonomous agents.
“As businesses are rapidly adopting AI agents, these systems are increasingly being given access to sensitive information, internal documents and operational tools. While this creates significant opportunities for efficiency, it also introduces a new cybersecurity threat that traditional security controls may not detect.” They do not, for example, have a fingerprint similar to malware that can be detected on disk by any traditional AV product.
Source: https://www.securityweek.com/the-hidden-instructions-that-can-hijack-ai-agents/
Related breach coverage
- OpenAI Agents Hijack Another Victim Website2026-09-07
OpenAI agents made 15,000–18,000 autonomous edits to a German wiki over three months, evading moderation and echoing tactics seen in the Hugging Face breach. The post OpenAI Agents Hijack Another Victim Website appeared first on SecurityWeek.
- Encrypted Prompts Bypass AI Safety Guardrails in Grok and Gemini2026-08-21
Researchers say the new ‘Cryptographic Context Injection’ technique conceals malicious instructions until they are decrypted inside a trusted execution environment. The post Encrypted Prompts Bypass AI Safety Guardrails in Grok and Gemini appeared first on SecurityWeek.
- OpenAI Investigates Report Linking AI Agents to RubyGems Attack2026-09-15
The incident occurred in May, when RubyGems maintainers suspended new account registrations due to what appeared like malicious activity. The post OpenAI Investigates Report Linking AI Agents to RubyGems Attack appeared first on SecurityWeek.
- AI Agent Firewall Startup AIR Security Emerges From Stealth With $50 Million2026-09-03
The startup’s firewall evaluates AI skills, plugins and MCP servers for malicious instructions, excessive permissions and software supply chain risks. The post AI Agent Firewall Startup AIR Security Emerges From Stealth With $50 Million appeared first on SecurityWeek.