Update cookies preferences

The 5 Questions I would ask any AI Security Vendor

Cheng Ren
,
Principal Solutions Engineer
|
01 Jul 2026
Knitted-yarn illustration of a checklist on a clipboard with a magnifying glass over a question mark, surrounded by user, code, database, and padlock icons, on a pink grid background.

In the vendor evaluations I run, I keep seeing the same pattern: buyers score AI security vendors on the criteria they already know how to evaluate. Threat detection, alert fidelity, integration surface, time to deploy. Those are reasonable questions. They are also the wrong questions for the control problem that actually matters.

The non-human identity side of the environment was already undertracked before AI agents arrived. As of 2024, for every 1,000 human identities in most enterprise environments, there were over 20,000 non-human identities: service accounts, API tokens, OAuth grants, webhooks, machine credentials of every description. That ratio has only continued to climb as cloud, SaaS, and API-driven architectures have matured. Most of those identities are overprivileged. Many never expire. Most have no off-boarding process. In the evaluations I participate in, most security teams have a mature human identity program and essentially zero coverage on the machine-identity side. That was the gap before agents. It is wider now.

Enterprise engineering teams are now provisioning AI agents the same way they used to provision SaaS apps: via credit card, with no procurement review and no security sign-off. A developer connects an agent to a production system, the agent gets credentials, the agent gets tool access, and the agent begins operating. Onyx data puts the median number of AI agents per enterprise at over 1,000, and between 80 and 90 percent of those agents were invisible to security teams before they were discovered. The distribution mechanism is different from the NHI boom. The accountability gap is structurally the same.

Security programs get caught flat-footed when a new class of identity or connectivity outpaces the existing evaluation frameworks. The questions that actually matter for AI agent security are the questions that should have been asked of IAM vendors four years ago.

Here are the five questions I use to evaluate AI security vendors. A vendor who cannot give a crisp, specific, operational answer to each of these is not yet ready to be your primary control.

1. Identity attribution at runtime: can you tell me which agent, in which session, took which action, attributed to which user?

This is the first gap I surface in evaluations. Most vendors can tell me which credential was used. They cannot tell me which human initiated the session, what that session did across its full span, or whether the action was consistent with normal behavior for that agent.

With AI agents, the stakes are higher because the agents act across multiple systems autonomously. If an agent reads a record, writes to a database, and calls an API in a single session, you need full attribution across all three actions, bound to the identity of the user or workload that initiated the agent, not just a log entry that says the session occurred. "We log it" is not the same as "we attribute it." Ask for a demonstration of the attribution chain before you accept the claim.

2. Runtime control versus detection-only: when an agent reasons toward a harmful action, can you intervene at the action boundary, or do you only alert?

Detection is not control. The gap between knowing something happened and being able to stop it from happening is where the damage accumulates.

For AI agents, the question is whether the vendor can block or mask a harmful action at the moment it is about to occur, not after the fact. An alert that arrives after an agent has exfiltrated sensitive data is not a control. Enforcement that blocks or masks at the action boundary is a control. Ask the vendor to walk through what happens at the exact moment their system identifies a policy violation mid-session. If the answer is "we send an alert," that is detection, not control.

3. Autonomous action boundaries: when an agent operates without human oversight, what is the mechanism for escalating an action to a human reviewer before it executes?

In the evaluations I run, this is the question that produces the most hesitation from vendors. The question is whether there is a defined mechanism for escalating specific action types to a human reviewer before execution. The answers to look for: which categories trigger mandatory human review, how is the reviewer notified, and what happens to the agent session while the review is pending. Vague answers mean the escalation path has not been built.

4. Audit completeness: when an agent takes an action, does every prompt, response, and tool call appear in a complete and retrievable audit trail, or are there gaps?

A partial audit trail looks like a complete one until you need to reconstruct a sequence. For AI agents, completeness means every prompt the agent received, every response it generated, every tool call it made, and every API it hit, bound to the session, the identity, and the timestamp.

The MCP layer is where audit gaps appear today: if an agent calls tools through a tool server, confirm that the vendor captures the tool call at the tool boundary, not just at the agent boundary. Ask for a sample audit record. Verify it includes tool execution events, not just conversational turns.

5. Cross-platform coverage: when an agent moves from a SaaS surface to a cloud runtime to a coding environment, does coverage follow, or does it fragment?

The AI agents that create the most risk are not the ones operating in a single surface. They are the ones that move: a browser-based agent that invokes a cloud function, which calls a tool through an MCP server, which writes to a data system. The coverage question is whether a single control plane follows the agent across all of those surfaces, or whether each surface has a separate security story that produces separate, unreconciled data.

When I ask this question, I am asking for a map: here is the handoff point, and here is how your product maintains attribution and enforcement across it. If the answer is a list of integrations without a unified session model, you have point solutions, not a platform.

The Discipline

The buyers who close the gap apply the same discipline they would apply to any new identity class: ask about attribution, ask about control, ask about accountability, ask about completeness, ask about coverage. Ask early. Ask specifically. Require demonstrations, not slide decks.

The questions above are not novel. They are the questions every security buyer would want their IAM vendors to have been asked in 2019. We have another chance to ask them now, before the gap widens.

Ask them before you sign. 

Ready to put these questions to the test? Get a demo and see how Onyx answers each one. 

Table of Contents
Cheng Ren
Principal Solutions Engineer
01 Jul 2026

Cheng Ren is a Principal Solutions Engineer at Onyx Security, where he runs enterprise proof-of-value deployments for customers in financial services, healthcare, and industrial sectors. He came to Onyx from Akamai Technologies, where he joined through Akamai's acquisition of Noname Security, the API security company. He previously held senior Sales Engineering roles at Splunk and VMware. At Onyx, Cheng works across MCP gateway architecture, coding-agent runtime enforcement, Microsoft Copilot Studio policy, and inventory discovery across Bedrock, AgentCore, and the broader AWS AI surface.