Chris gave buyers a vendor test on Wednesday: show me an account the agent was not allowed to touch, the rule that blocked it, and the record. Here is that test run against Bryn, in public: one Play, one signal that clears the score, one account on a suppression list, one blocked action, one record line another person can reconstruct from. Then the honest edge of what Bryn cannot show.
Civic Team, Staff
||9 min read|
tl;dr
On Wednesday Chris handed buyers a vendor test: show me an account the agent was not allowed to touch, the rule that blocked it, and the record explaining what happened. This is that test, run against Bryn, in public. One Play, one signal that clears the score, one account on a suppression list, one blocked action, one record line another person can reconstruct from. Then the honest edge: Bryn can show you what it refused to initiate. It cannot show you what a third-party system did after an action left.
On Wednesday, Chris closed his read of Workday's quarter with a test for any agent vendor: show me a blocked action. An account the agent was not allowed to touch, the rule that blocked it, and the record explaining what happened. If the answer is a policy document or a roadmap slide, the control is a promise.
Fair test. Here it is, run against Bryn. Values are illustrative and the account is anonymized; the shape of the run is what you would see in the audit log.
The setup
One Play, pricing-return-nudge-v1: when a known account comes back to the pricing page and reads three pages inside 48 hours, post a note to the growth channel in Slack and open a CRM task for the account owner. Nothing else. No email, no sequence, no write to any system the Play does not name.
Three controls sit on that Play, all set before anything fired. A boundary: channels limited to Slack and a CRM task, a daily cap of 25 runs (illustrative), and a suppression list holding the domains of current customers, because a renewal conversation should not get a prospecting nudge. An approval: the operator approved the Play on August 19 and left it in Run mode, so Bryn runs matching instances without asking again. A kill-switch, unused here, that would suspend everything and log the suspend.
Then the part Chris asked for. We put one current customer's domain on the suppression list and waited for that account to trip the signal. In the record it is acct_4471, domain customer-example.com.
The run
Five steps, each tagged with what governed it, using Chris's labels from Wednesday: model judgment, a fixed rule, or something fixed before the run began.
Signal. Bryn matched pricing → return → 3 pages (48h) on acct_4471 at 09:12 UTC. The pattern was written into the Play and approved before the run, so the match is a check against something already on record. Fixed before the run. (Resolving anonymous visits to a named account is its own subject; Brad's piece yesterday covers how far to trust that class of signal.)
Score. Bryn scored the account 0.74 against the customer profile; the Play's threshold is 0.70. The one step where the model exercises judgment, and a similar account on a similar day could score differently. Model judgment.
Boundary. Bryn checked the account against the Play's boundary. Channel: allowed. Daily cap: 6 of 25 used. Suppression list: customer-example.com matched. The check fails on that third line. Fixed rule; nothing about the score or the signal can override it.
Action. No action sent. Slack was not written to; no CRM task was opened. Judgment got the account this far, and a rule decided whether anything left. Mixed, in exactly Chris's sense.
Record. Bryn wrote the run to the audit log with the same required fields every run carries, action or not. Fixed rule.
Replay the blocked run.
Step through the five things Bryn did with acct_4471 at 09:12 UTC. Each step shows what happened, in Bryn's own record register, and what governed it.
Step
What happened
What governed it
1 Signal
I matched pricing → return → 3 pages (48h) on acct_4471 at 09:12 UTC.
Fixed before the run. The pattern is written into the Play and approved before anything fires.
2 Score
I scored acct_4471 at 0.74 against the profile. Play threshold 0.70. Cleared.
Model judgment. A similar account on a similar day could score differently.
Fixed rule. Nothing about the score or the signal can override the list.
4 Action
No action sent. Slack not written. No CRM task opened.
Mixed. Judgment got the account this far; a rule decided whether anything left.
5 Record
I matched pricing → return → 3 pages (48h) on acct_4471 at 09:12 UTC, scored 0.74, checked pricing-return-nudge-v1 boundary, blocked: domain on suppression list (customer). No action sent. Logged.
Fixed rule. Every run writes the same required fields, action or not.
step 1 of 5SignalFIXED BEFORE THE RUN
09:12:07Z I matched pricing → return → 3 pages (48h) on acct_4471.
The pattern is written into the Play and approved before anything fires. The match is a check against something already on record.
1 / 5
All values illustrative, not a benchmark. Account anonymized. Governance tags follow Chris's Wednesday piece: model judgment, fixed rule, fixed before the run, mixed.
Figure 1. The blocked run, step by step, with what governed each step (illustrative, not a benchmark).
BRYN byCivicLabor Day offer ⬩ through September 17
Save Your Labor (Day)
Bryn watches your site, scores the account, runs the Play, and files the run. A free month of it, on any tier.
Timesheet ⬩ arbor.devPunched ⬩ Tue 2:02 PM
2:02:08 PMWatched a return to pricing, then the comparison page
2:02:09 PMScored the account 86
2:02:10 PMRan the pricing.follow-up Play into Slack and the CRM
The whole run is one line in the log, in Bryn's own register:
I matched pricing → return → 3 pages (48h) on acct_4471 at 09:12 UTC, scored 0.74, checked pricing-return-nudge-v1 boundary, blocked: domain on suppression list (customer). No action sent. Logged.
Three people can read that line and each gets what they came for.
The operator reads that the Play is working: the signal was real, the score cleared, the list held, nothing to fix. Maybe something to follow up by hand, since a current customer reading pricing three times in two days is a renewal signal, and the record surfaced it without sending anything.
The CFO reads a control that ran with nobody in the room. No send, no spend against the cap, and a run that can be reconstructed from one place without a meeting. The audit trail, showing up as a line rather than a promise.
The DPO reads that a do-not-contact rule was enforced at the moment of action, not reviewed after the fact, with the reason on the line rather than in a ticket. When someone asks whether any customer received a prospecting message that quarter, the answer is a filter on the log, not an investigation.
Figure 2. One line, three readers (illustrative, not a benchmark).
A run that clears the boundary produces a line of the same shape with action filled in: what was posted, where, under which approval. We wrote about why that per-run record is the thing to ask for in Bounded Autonomy Is the Spec, and why a claim you can check against your own systems is what closes deals now in Proof of Claims.
What Bryn cannot show you
Chris drew a distinction on Wednesday that we want to keep intact. Workday enforces controls inside systems it owns end to end. Bryn coordinates actions across third-party tools. So the record above is complete about one thing: what Bryn saw, how it scored it, what it checked, and what it refused to initiate.
It is not a record of what Slack or a CRM did after receiving an action. In this run nothing was sent; in a run where something is, Bryn's record ends at the handoff. A Slack post edited later or a CRM task reassigned lives in those systems' logs, not Bryn's. Every vendor that coordinates across tools has the same edge; the honest ones say where the record stops.
One more. The suppression list is only as good as what is on it. Bryn blocked customer-example.com because the operator put it there. A customer whose domain never made the list would have received the nudge, and the record would show that too. The control is enforced; the list is yours.
Do this to your own stack
Before you trust any agent with a channel, make it fail on purpose. Pick one workflow it already runs and put one account you know well on its do-not-touch list. Fire the trigger yourself, or wait for it. Then go read the record and time how long it takes to find the refusal: the account, the rule that stopped it, and the confirmation that nothing left. All three in one place in under a minute, and the control is real. A settings screen, a support thread, or the vendor's own memory of what happened, and the control is a promise, which you now know before it cost you a customer. Run it again on a normal Tuesday, three months in, when nobody from the vendor is watching.
Where this leaves us
Wednesday's test was aimed at every vendor, including us. A line from the log is a better answer than a paragraph about how seriously we take governance.
Bryn is not another dashboard to watch. It is the governed execution layer that runs Plays through your stack.
Our team brings decades of experience across the domains that matter: 10 years in AI and agentic systems, 65 in financial services, 35 in identity and access management, 30 in marketing and AdTech, 15 in legal and professional services, and 12 in manufacturing and industrial.
We're for operators who can't afford unintended actions or silent failures, and who want the agent in production quickly and effectively.