Tideline streams live sport to about two million subscribers, which means their traffic is not a curve but a series of cliffs. Fifteen services, twenty-four engineers, and a load profile where the difference between a quiet Tuesday and a Saturday fixture is two orders of magnitude.
Their alerting had been tuned for the cliffs, which meant it was permanently wrong on the flat. A threshold that made sense during a match produced a steady drip of noise the rest of the week, and nobody wanted to be the person who loosened it before a fixture.
The first month went badly
Tideline is the team we point to when someone asks what goes wrong. They set a 99.5% objective on playback start, which felt conservative and was in fact extremely loose — about three and a half hours of failure a month. For four weeks their error budget absorbed real problems without complaint, including a regional CDN fault that degraded playback for a segment of their audience for most of an evening.
Reyes tightened playback to 99.95% in week five and split it by region, which is when the numbers in this study start to mean anything. Everything before that is a measurement of the wrong thing.
What the corrected objectives produced
From week five to the end of the window, Tideline averaged two alerts a week against a hundred and eighty before. That is the largest reduction of any team in the ninety-day study, and it is largely a function of how bad the starting position was: alerting tuned for peak load, applied to a system that is off-peak most of the time.
The regional split mattered more than the tightening. A fault affecting one CDN edge now burns that region's budget rather than being averaged into a global number that stays comfortably green while a segment of the audience cannot watch.
Match days
The behaviour Reyes wanted and did not expect to get is that a burn during a fixture is treated as more urgent than the same burn on a Tuesday, without anyone configuring that. It falls out of the model: at match-day volume a given error rate consumes the month's budget in minutes, so the burn rate is steep and it pages. The same error rate at Tuesday volume takes days, so it opens a room and waits.
- Match days — fast burn, paged immediately, twice in the window.
- Weekdays — slow burn, incident room opened, reviewed in hours.
- Neither required a separate rule, a schedule, or a maintenance window.
Tideline's written-review completion also moved, from a median of nine days after an incident to under two. Reyes is blunt about the cause: the draft is already correct about the facts, so the work is the analysis rather than the transcription.
