We asked forty-one teams to run Meridian as their primary paging path for ninety days and to let us look at the numbers afterwards. All forty-one finished. This is what the data says, including the parts that are less flattering than the headline.
The headline
Median alert volume per team per week fell from 217 to 2. That figure has appeared on our home page for some months and it is accurate, but a median compresses a lot. The distribution is more interesting than the midpoint.
- Eleven teams ended below one page per week. Most were running fewer than a dozen services.
- Twenty-four teams landed between one and four pages per week.
- Six teams stayed above ten. Every one of them had a dependency they did not control.
That last group is the honest part. If a payment processor or a regional cloud provider is failing weekly, no alerting model will make that quiet, and it should not. Those teams got fewer pages than before, but they were still being woken, because something genuinely was breaking their objectives on a weekly basis. Our system correctly declined to hide it.
Where the reduction came from
It is tempting to attribute the drop to the signal model alone. The breakdown is less romantic. Roughly half the reduction came from deduplication — one failure producing one page instead of fourteen, as each service downstream of it reported its own symptom.
About a third came from budget-aware severity: incidents that previously paged now opened a room and waited for business hours. The remainder came from teams deleting rules they had been carrying for years and had stopped believing in, which required no technology at all — only an occasion to look.
We deleted about sixty alert rules in the first week. Half of them nobody could explain the origin of.
PLATFORM LEAD · TEAM 19
What went badly
Three teams missed an incident in the first fortnight. In two cases they had set an objective far looser than their users’ actual expectations — a 99.5% target on a checkout flow — so the budget absorbed a real problem without complaint. The system did what it was told. It was told the wrong thing.
The third case was ours. A regional ingest path was dropping a fraction of spans under load, which understated the error rate for about four hours. We have written that up separately and it changed how we compute burn during ingest degradation.
What we would tell you before you start
Set objectives you would defend in front of the people who depend on the service, not ones that are comfortable to hit. Run the new path in parallel with the old one for two weeks and compare what each would have woken you for. And expect the first fortnight to be noisier than the steady state, because you will be discovering which of your objectives were fiction.
None of the forty-one teams asked to switch back. We would rather that were because the model is right than because switching twice is tiring, so we will run the study again at a year.