Ask a RevOps leader what they want from an agentic tool this year and you will hear a list that would have sounded strange in 2024: a permissioned data layer, budget and usage controls, audit trails, bounded autonomy. Apollo's write-up of how revenue operations leaders are deploying agents says it plainly: the question has moved from whether to deploy agents to how to "deploy them safely, govern them rigorously, and measure their direct impact," and vendors are being evaluated "on control plane capabilities first, not feature breadth" (Apollo, 2026).
One practitioner survey Apollo cites puts 73% of organizations past the experimentation phase, running agents inside core GTM workflows (Revenue Wizards, single survey, worth holding loosely). The exact number matters less than the consequence: when agents run inside the workflows that touch customers, the buying question flips. Capability is table stakes. Boundedness is the spec.
The failure stories are unboundedness stories
Nobody pulls an agent pilot because the drafts were mediocre. Mediocre drafts get edited.
Pilots get pulled when something ran that should not have: the sequence that went to current customers, the tool that kept sending while someone hunted for the off switch, the outreach nobody could explain to the account team afterward. Every one of those is the same failure. Not a capability failure, a bounds failure. The tool could do more than anyone had decided it should.
Which is why the procurement conversation now starts where the incident reviews ended. The buyers writing this year's requirements documents are, in large part, the people who cleaned up last year's pilots.
The five questions
Strip the vocabulary and the spec is five questions. They take less time than the demo's opening slide.
Can I name the channels it writes to? Not "it integrates with your stack." Which systems can it touch, listed, with everything else off by default. An agent that can write anywhere is not flexible. It is unscoped.
Can I cap it? A usage limit that the operator sets, so one noisy Tuesday cannot spend a month of budget, or of channel goodwill, before lunch.
Can I stop all of it, immediately? One control, global, now. A suspend that finishes the current batch first is not a suspend. It is a confession with a delay.
Can I read what each run saw and did? Per-run records: what was seen, why it qualified, what happened. Event-level and exportable, not a monthly activity summary.
Can I see why it scored what it scored? The axes, the weights, the threshold. If the answer is "the model decides," you are being asked to delegate a judgment nobody can inspect.
- Can I name the channels it writes to?
- Can I cap it?
- Can I stop all of it, immediately?
- Can I read what each run saw and did?
- Can I see why it scored what it scored?
Ask them in the scheduling email; count what gets shown versus narrated.
0 of 5 shown
Check off what the vendor actually showed you, not what the deck claimed.
Score what you were shown, not what was claimed. A vendor either shows you the control or narrates around it.
The useful property of these questions is that they are falsifiable in a demo. A vendor either shows you the channel list, the cap, the suspend, the record, and the score, or narrates around them. There is no third behavior. You will know within ten minutes which one you are watching.
