Remote Code Execution (RCE) is a class of vulnerability in which an attacker causes arbitrary code to run on a target system without authorized access. In AI agent contexts it usually arises when an agent's tool use can be steered – through prompt injection, tool poisoning, or a compromised MCP server – into executing commands the operator never intended.
- An attacker runs arbitrary code on a system without authorized access to it
- Agent RCE is the worst-case failure mode of tool use and autonomy
- The agent's execution environment sets the blast radius, not the model
- Practical defense sits at the agent runtime layer rather than the model layer
Why is remote code execution important?
RCE through an AI agent is the worst-case failure mode of agent autonomy. The agent inherits the credentials and permissions of its execution environment, so if it can be steered toward a code-executing tool, the attacker inherits that same scope. A prompt becomes a command, and the command runs with whatever the agent was trusted to hold.
This is no longer theoretical. CrowdStrike documented malicious prompts injected into legitimate AI tools at more than 90 organizations during 2025, generating commands that stole credentials and cryptocurrency, alongside exploitation of AI development platforms to establish persistence and deploy ransomware. Both CISA and OWASP have flagged agent RCE patterns in recent guidance.
What makes agent RCE distinct from classical RCE is the entry point. There is no memory corruption bug to patch. The vulnerability is that an agent with tool access will follow instructions it finds in the data it processes.
What is remote code execution?
Remote code execution is a vulnerability class allowing an attacker to run commands on a target system without physical access or valid credentials for it. Classically the attacker exploits a flaw in how a system handles untrusted input, then executes code at the privilege level of the compromised process.
In AI agent environments the mechanism changes while the outcome does not. An agent reads external content – a document, or a tool response it has no reason to distrust – that carries instructions crafted by an attacker. The agent treats those instructions as legitimate and calls a tool that can execute code, write files, install packages, or invoke another agent. Nothing was exploited in the traditional sense; the agent was persuaded.
Three paths recur. Prompt injection carries the instruction. Tool poisoning compromises the tool an agent trusts. A compromised MCP server sits between the agent and everything it reaches. In each case the defense is the same: evaluate the tool call before it executes.
Types of remote code execution
RCE is usually grouped by the flaw that enables it, and agent environments add a category that did not previously exist.
Classical paths include memory corruption, where malformed input overwrites execution state; insecure deserialization, where untrusted serialized data reconstructs into executable objects; command injection, where unvalidated input reaches a shell or interpreter; and template or expression injection in web frameworks.
Agent-mediated RCE is the newer category, and it does not require a software defect at all. The agent is functioning as designed – reading content, choosing a tool, calling it. What fails is authorization at the action boundary. Sub-paths include instruction injection through retrieved content, poisoned tool descriptions that misrepresent what a tool does, and compromised MCP servers that return attacker-controlled results.
The distinction matters for remediation. Classical RCE is patched. Agent-mediated RCE is governed, because there is no bug to fix.
Remote Code Execution & Onyx
Onyx inspects every tool call at the request boundary, with action-class controls that intercept the steps an RCE chain depends on. Destructive action detection flags operations capable of altering or deleting state. Tool restrictions limit which agents may reach code-executing tools at all. Package provenance checks what an agent is about to install before it installs it.
Because enforcement happens inline, an agent steered toward a code-executing tool meets a policy decision before the call resolves rather than an alert afterward. Runtime and prompt injection defense handles the instruction path, and MCP security governs the servers an agent depends on.


