Skip to main content
Field Notes/bryngtmdeliverabilityoutboundemailproof

Deliverability Is a Commons

Email deliverability is a commons: every over-sender degrades the channel for everyone, including themselves. Google and Yahoo enforce a 0.3% spam-rate ceiling, most fully autonomous AI SDR pilots get pulled inside 90 days, and warmup tricks do not fix a grazing problem. A usage cap protects the budget. An audience cap protects the channel.

Civic Team
Civic Team, Staff
8 min read
Hero: Deliverability Is a Commons. Five sender agents graze one shared band of inbox trust. Four draw on it lightly; one coral over-sender draws three lines into the same pasture, and the band reddens beneath it. A fence-line chip marks enforcement at 0.3 percent spam rate. Caption: every over-sender degrades the channel for everyone, including themselves.
tl;dr

Email deliverability is a commons: every over-sender degrades the channel for everyone, including themselves. Google and Yahoo now enforce a 0.3% spam-complaint ceiling, most fully autonomous AI SDR pilots get pulled inside 90 days, and warmup tricks do not fix a grazing problem. The fix is an audience cap: an explicit ceiling on how much of your market any sequence may touch. A usage cap protects the budget. An audience cap protects the channel. They are not the same cap.

Your domain reputation is grazing land.

Every mailbox provider maintains a shared ledger of trust between senders and inboxes, and every message you send draws on it. So does every message your competitors send, and every message their agents send. In 2026 the herd got very large very fast: agents that write, sequence, and send without a human touching the workflow, each one grazing the same pasture your deals have to cross.

Economists have a name for what happens next. In a commons, each grazer captures the whole benefit of one more animal and pays only a fraction of the cost, so everyone adds one more animal until the field is dirt. Nobody intended the dirt. The incentives did it.

That is not a metaphor stretched over email. It is a literal description of how inbox trust works: a finite, shared resource, drawn down fastest by whoever sends most, with the damage billed to everyone on the channel.

The 90-day pattern

On Tuesday, Chris wrote about Ramp: a company that shut down the AI SDR program behind a meaningful share of its pipeline and spent four years rebuilding outbound around selection, bounded action, and a record. Ramp is the famous case. The pattern underneath it is bigger, and it runs on a schedule.

Most fully autonomous AI SDR pilots get pulled inside 90 days. Digital Applied's 2026 aggregation of AI SDR benchmarks puts a number on the leading cause: 47% of attempted deployments hit a domain-reputation wall inside their first 90 days, and roughly a fifth never recover the inbox placement they started with. The 47% is that one aggregation's figure, drawn from sender-platform data, so hold it loosely. The 90-day shape it describes, volume up fast, reputation down faster, pilot pulled, repeats across the independent 2026 summaries of the same benchmark set.

The cycle is almost boring in its regularity. A pilot launches. Volume multiplies, because volume is the one thing the new tooling makes free. Complaint rates creep. The provider throttles. Reply rates fall, which the dashboard misreads as a copy problem, so the pilot sends more. And somewhere around day 90, someone with a title pulls the plug on a program that was killed by its own throughput.

Warmup tricks, domain pools, and clever openers move the wall a few weeks. They do not remove it, because the wall is not a filter to outsmart. It is the pasture running out.

The fence line

The fence exists, and it is public.

Since February 2024, Google's sender guidelines require senders to keep spam-complaint rates reported in Postmaster Tools below 0.3%, with bulk senders held to the same 0.30% line. The same page tells you where Google actually wants you: keep spam rates below 0.10%, and avoid ever reaching 0.30% at all. Yahoo's sender requirements hold the identical line: keep your spam rate below 0.3%.

Read that number the way it is written. Three complaints per thousand delivered messages is the enforcement threshold. One per thousand is the comfort zone. The distance between a healthy channel and a burned one is two annoyed recipients per thousand, and a fully autonomous sender can close that distance in an afternoon.

The asymmetry is the part most teams learn the expensive way. Damage is fast and recovery is not. Google's guidance says plainly that it can take time for improvements in spam rate to reflect positively on spam classification. The same aggregation reports post-incident reputation recovery measured in weeks: roughly three at Google, closer to seven at Microsoft 365. The over-grazing took days. The regrowth takes a quarter.

Channel health

Sends this month against the published spam-rate lines. Only the two dashed thresholds are sourced; the curve is the shape of the story, not your data.

Sends vs addressable audienceIllustrative complaint rateZone
100% (1x)0.05%Inside Google's recommended 0.10% ceiling
200% (2x)0.13%Above the recommendation, drifting toward the fence
350% (3.5x)0.31%At the 0.30% enforcement threshold Google and Yahoo publish
500% (5x)0.57%Past the threshold; recovery measured in weeks

Curve values illustrative, not a benchmark. Thresholds sourced to Google's sender guidelines and Yahoo's sender requirements.

Drag the slider, or use arrow keys. Only the two dashed lines are sourced: Google enforces at a 0.30% spam rate and recommends staying below 0.10%; Yahoo requires below 0.3%. The curve shape and all rate values along it are illustrative, not a benchmark. Damage is faster than recovery: post-incident reputation rebuilds are reported in weeks, not days.

BRYNbyCivic Running now

What would this essay do if it could act? It just did.

Essay, alone

Someone reads it. Maybe they fit your ICP. The minute passes and nobody downstream ever knows.

Your chance to reach your engaged, identified prospect: Gone

Every run lands on the record.

Two caps, two jobs

Part of what Ramp spent four years building, per Tuesday's piece, was volume discipline enforced in the system rather than promised in a slide. That discipline has a name worth separating from its look-alike.

A usage cap limits how much work a system may do for you: actions run, spend incurred. It protects your budget. It is priced into every plan tier on the internet, ours included, and it knows nothing about your audience.

An audience cap limits how much of your addressable market any sequence may touch in a given window. It protects the channel. It knows your audience and how often that audience has already been grazed, and no plan tier can set it for you, because it is a policy about your pasture, not a feature of anyone's product.

A usage cap is not an audience cap Two caps, two jobs. Vendors blur them. Your channel cannot afford to. USAGE CAP Protects the budget. WHAT IT LIMITS How much work the system may do: actions run, spend incurred. WHO IT PROTECTS You. It keeps the invoice inside the plan. WHAT IT KNOWS Your plan tier. Nothing about your audience. WHAT IT CANNOT DO Stop a sequence from touching the same market twice in a week. AUDIENCE CAP Protects the channel. WHAT IT LIMITS How much of your addressable market any sequence may touch this month. WHO IT PROTECTS Everyone sharing the channel, including you next quarter. WHAT IT KNOWS Your audience, and how often it has already been touched. WHAT IT IS A policy you set and enforce against the record. Not a plan tier. Do not let anyone sell you the first as the second. text detail ⬩ policy design, not product UI
Two caps, two jobs. Policy design, not product UI.

The failure mode is buying the first and believing you bought the second. A usage cap with no audience cap will happily spend its entire monthly allowance on the same three thousand people, twice a week, until the complaint rate crosses the fence line. The budget was protected the whole way down.

Do not let anyone, including us, sell you the first as the second.

Where Bryn stands

Bryn is the governed execution layer that runs Plays through your stack, and its position on this is deliberately narrow. Plays write only to the channels you name. Usage limits constrain what a month of execution can cost. Suspend stops execution. And every run writes a per-run record: what was sent, where, under which Play.

What Bryn does not do is set your volume policy. The audience ceiling, how much of your market a sequence may touch and how often, stays yours to set. That is not a missing feature. A commons survives when the grazers own their fences.

What the record changes is enforceability. An audience cap you can audit against actual sends, channel by channel and run by run, is a policy. One you cannot audit is a hope. The difference between the two is what the 90-day pattern keeps collecting on.

The Monday checks

Two checks, neither of which requires buying anything.

First, open Google Postmaster Tools and read your spam-complaint rate against the published lines: below 0.10% you have room, between 0.10% and 0.30% you are drifting toward the fence, and at 0.30% you are on the line Google enforces. If you have never opened Postmaster Tools, that is the finding.

Second, set an audience cap: an explicit ceiling on the share of your addressable market any sequence may touch this month. Write it down, tell the team, and check actual sends against it at the end of the month. If nothing in your stack can produce the actual-sends number, that is the second finding.

The pasture recovers. That is the good news buried in every commons story: fences work, and the grazers who set them are the ones still selling into the channel next year.

Bryn starts at $49 a month with a 7-day trial. Pricing is public at civic.com/bryn/pricing.


Further reading

Sources

Civic Team

Civic Team

Staff

More essays by Civic

Our team brings decades of experience across the domains that matter: 10 years in AI and agentic systems, 65 in financial services, 35 in identity and access management, 30 in marketing and AdTech, 15 in legal and professional services, and 12 in manufacturing and industrial.

We're for operators who can't afford unintended actions or silent failures, and who want the agent in production quickly and effectively.