TL;DR - The platform team is the one team in engineering that runs a production service for internal customers and usually has no pager, no objective and no incident process for it. The reason is that platform incidents do not look like outages: they block change rather than break traffic, so nothing on a dashboard turns red while forty teams sit unable to deploy. Fix it in this order. Page on the journey, which means a synthetic service that builds and deploys through the real road every half hour and wakes someone when it cannot. Tier your own services honestly, because a CI outage at 3am on a Saturday is a Monday ticket and a broken secret store is not. Separate the pager from the interrupt shift, because answering questions and restoring a road are different promises. Publish what your users can expect and then hold it, including the part where the promise is business hours. And declare an incident when the road breaks, even though nothing is technically down, because the alternative is that your users quietly build a way around you.
I build and run platforms for a living, and I have spent years telling other teams to own their services in production. Set an objective. Wire a pager to it. Run a review when it breaks. Then I have gone back to my own team, where the shared pipeline, the base images, the secret store, the deploy controller and the templates that generate half the estate all run on the reliability model of “someone will notice”.
This is the least defensible thing platform teams do, and it is close to universal. The whole discipline rests on the claim that the platform is an internal product with real customers. A product with real customers in production has an on-call rotation. Ours usually has a Slack channel and a person who happens to know how the thing works.
The tell
There is one question that settles whether your platform has an operational practice or a habit. How did you find out about your last platform outage?
If the answer is a developer asking in a channel whether CI is weird for anyone else, you do not have on-call. You have a distributed group of unpaid monitoring agents who are also your customers, and they are reporting the incident to you an unknown number of minutes after it started, in a channel that nobody reads at weekends.
The rest of the picture is usually consistent with that. There is no severity for platform problems, so every one of them is handled at the urgency of whoever picked it up. There is no comms, so eleven people each ask separately and each get a slightly different answer. There is no restoration target, so nobody can say whether ninety minutes was good or bad. And there is no review afterwards, because the thing that broke was a template, and templates do not feel like production even when they generate production.
Meanwhile the team is running an incident practice for the whole organisation and telling stream-aligned teams that reliability is a property you design in.
Platform incidents have a different shape
Before staffing anything, it is worth being precise about why the standard playbook does not transfer cleanly. Platform failures differ from application failures in four ways, and each one has an operational consequence.
- They block change rather than break traffic. The most common platform incident does not take a single user-facing service down. It stops anyone from shipping: builds fail, the deploy hangs, the registry rejects pushes, the scaffold command errors. Every availability number you own stays green through the entire event, which is why this class of outage is systematically under-reported and systematically under-staffed. If you only alert on things being down, you have instrumented the failure mode you almost never have.
- They fan out instead of fanning in. An application incident affects one service and its callers. A bad base image affects every service that rebuilds this week. Your blast radius is not measured in requests, it is measured in teams, and it grows every quarter you succeed at adoption.
- They are frequently latent. A template change ships a defect into every repository created after it. The incident started three weeks ago and will keep starting for as long as the template is live. There is no spike on any chart, just a slowly growing population of services with a problem you introduced.
- The failure is often yours, running inside somebody else’s repository. Your pipeline definition, your library, your policy check, executing in a codebase you do not own and whose logs you may not be able to read. The debugging loop crosses an ownership boundary in the middle, which is exactly where most incident processes assume it does not.
That last one is worth sitting with, because it changes the severity conversation. The honest unit of impact for a platform incident is how many teams cannot ship right now, and that number is knowable within a minute if you have thought about it in advance, and unknowable for an hour if you have not.
Page on the road, not on the components
The highest-leverage thing a platform team can build for its own reliability is not a dashboard. It is a robot that uses the product.
Run a synthetic service through the real golden path, on a schedule. Every thirty minutes, a job creates or updates a trivial service from the actual template, pushes it through the actual pipeline, deploys it to a real environment, hits its health endpoint, and tears it down. When that fails twice in a row, something pages a human. That is roughly two days of work and it is worth more than any amount of component monitoring, because it asserts the only property your users care about: can a change get to production right now?
This is the same principle as measuring an objective on the user journey rather than on host uptime, which I have argued at length in the post about error budgets nobody spends, applied to a platform where the user journey happens to be a deploy. It is also the difference between watching the components and knowing what the system did to the people using it, which is the whole argument in dashboards versus real observability. Your control plane can be perfectly healthy while nobody can ship, and only the robot will tell you.
Around that canary, tier what you actually run, because “the platform” is not one service:
- Runtime primitives. The ingress, the secret store, the service mesh if you have one, the cluster control plane, the thing that holds credentials. When these fail, customer-facing services fail with them. This tier is production in the ordinary sense and needs ordinary production treatment.
- The change path. CI, the registry, the deploy mechanism, the templates. Failure here blocks work rather than breaking traffic. It is urgent during working hours and, for most organisations, genuinely not urgent at 4am on a Sunday.
- The conveniences. The portal, the catalogue, the dashboards, the internal CLI’s nice-to-have subcommands. These can be broken for a day. Say so out loud, because a portal that looks like the platform invites people to treat its outage as a platform outage.
Writing this tiering down is most of the work. It converts every future 2am judgement call from an argument into a lookup.
An SLO for a road, and the awkward part
You can put an objective on a platform, and you should, but there is one measurement problem you have to solve first or the number will be dishonest.
The obvious indicator is pipeline success rate. The problem is that most pipeline failures are supposed to happen: the tests failed, the linter caught something, the developer pushed a typo. Those are the road working. If you count them, your objective measures the quality of other people’s code and you will never be able to spend or defend it.
So the engineering task is classification. Every failure in your pipeline needs to be attributable to one of two buckets: the road broke or the change was bad. That means your steps exit with codes you own, your platform-owned stages are distinguishable from user-owned stages, and infrastructure faults are tagged as such at the point of failure. If you cannot do this, do not set an objective yet. Go and make the failures classifiable, then come back.
Once you can, a small set of honest indicators covers most of it:
- Change path availability. Share of platform-attributable pipeline runs that complete successfully. The canary above is the cleanest possible version of this signal because it removes user code from the equation entirely.
- Deploy latency at the high percentile. Time from merge to running, measured end to end. A road that works but takes forty minutes is a road people will start avoiding, and the avoidance will look like teams routing around a good product rather than like the reliability problem it is.
- Primitive availability, on the user’s journey. Not “the secret store was up” but “workloads could fetch credentials”.
Then the uncomfortable arithmetic. You are a hard dependency for every team you serve, so their objective cannot be better than yours. If you ask forty teams to hold three nines and the road holds two, you have handed them a target they cannot reach with a straight face, and the first team that works this out will stop taking your reliability advice seriously. Publish your own numbers first, and publish them where your users are.
Six people and a rota
Here is where most of this advice usually dies, and it deserves a straight answer rather than an aspiration. A platform team of six cannot run a follow-the-sun 24/7 rotation. Any post that tells you otherwise is describing a company with three platform teams in three timezones.
What you can do:
- Match the promise to the tier, and write the promise down. Runtime primitives get out-of-hours paging because their failure is a customer-facing outage. The change path gets a business-hours pager with a stated response window. “The road is a business-hours service, and here is what to do if you need an emergency deploy at midnight” is a completely legitimate contract. An unwritten expectation of 24/7 heroics is not, and it is the version that burns your senior engineer out in eight months.
- Give the emergency path a real escape hatch, and make it loud. There will be a night when someone must ship a fix and the road is down. The break-glass procedure should exist, be documented, require a human to consciously choose it, and generate a ticket automatically. An undocumented escape hatch gets invented under pressure and then becomes a permanent second way to deploy, which is sprawl arriving by the back door.
- Separate the pager from the interrupt shift. These get conflated constantly and they are different jobs with different promises. The interrupt shift absorbs questions, review requests and “how do I” so that the rest of the team is genuinely untouchable, exactly as the first ninety days should have set up. The pager restores a broken road. The same person can hold both in a small team, but the response times, the escalation and the definition of done are not the same, and treating a question as an incident is how the whole thing becomes noise.
- Price the rota into the plan. On-call is not free capacity that appears between projects. It is part of the run cost that year two quietly fills up with, and if it is not in the quarterly allocation as a number, it will be paid out of the evenings of whoever is most conscientious.
- If you cannot staff the promise, shrink the promise. Not the documentation of it. A team that publishes “best effort out of hours” and means it is more trustworthy than one that publishes an SLA it silently misses.
Say something before they ask
Your users are engineers, which makes incident communication easier than in most products, and platform teams still do it badly.
- Announce before you are asked. One channel, one status page, updated by the person on the pager, first message within minutes of detection even when it only says “we know, investigating”. The single largest cost of a platform incident is not the downtime, it is forty engineers each independently spending twenty minutes debugging your problem inside their own repository.
- Say what is broken and what still works. “Deploys to production are failing, running services are unaffected, builds still pass” is the sentence that prevents an afternoon of duplicated panic. Impact statements written in terms of what your users can and cannot do beat statements about which component is unhealthy.
- Publish the workaround, and then take it back. Give people the manual path if one exists, and remove it explicitly when the road is fixed. A workaround left standing becomes an unofficial road, and unofficial roads are how you end up with two of everything.
- Own the incidents you did not cause. When your CI provider is down, you are still the organisation’s status page for it. “Not our system” is technically true and operationally useless to the people you serve, and being the calm source of truth during someone else’s outage buys more credibility than most of what you will ship this quarter.
Road breakage is an incident, even when nothing is down
I have made this claim in passing before, when arguing that a paved road should be changed and closed the way you would change a public API. Here is the operational version of it.
When a template change breaks builds across the estate, declare an incident. Use the same severity language, the same channel and the same review as you would for a customer-facing failure. The test I use: would we have declared this if a customer-facing service had done this to a customer for the same duration? If yes, declare it, because the only difference is that your customer sits three desks away and is too polite to churn immediately.
Then run the review properly, which means a review that changes something rather than a document that gets filed, the whole argument in the postmortem that changed nothing. Platform postmortems have a specific and repeated conclusion, which is that the change was shipped to the entire estate at once because shipping to the entire estate at once is the default in most template repositories.
Which leads to the actual fix, and it is not a process fix. A change to a golden path is a production deploy to every service on it, so ship it like one. Version the path. Roll new template and pipeline versions to a canary cohort of friendly services before the rest. Keep the ability to pin, and the ability to roll back, and test the rollback at least once, because a rollback path you have never exercised is a guess, which is the same reason an environment you have never broken on purpose is lying to you. Almost none of this is possible with a fat scaffold that copies logic into every repository, and all of it is possible with thin templates over shared primitives you can update centrally.
When this is the wrong answer
The honest off-ramps, because plenty of teams should not build any of this yet.
- You are three people running one road for six services. There is no rota here, there is a group chat and reasonable goodwill. Building a severity ladder and a tiering document at this size is the same category error as running Kubernetes for three services. Add the deploy canary anyway, because it is two days of work and it pays back immediately. Skip the rest.
- The platform is not load-bearing yet. If nobody is blocked when your road breaks, that is not resilience, it is a lack of adoption, and your problem is winning the first users, not staffing a pager. Reliability apparatus ahead of dependency is ceremony.
- The organisation already has an incident function. Join it. A parallel platform-only incident process with its own severities, its own channel and its own review template is a second system for the same job, and it will diverge within two quarters. Adopt the existing language and add the platform’s services to the existing tiering.
- What breaks overnight is not something you can fix overnight. If the realistic 3am failure is your cloud provider or your SaaS CI vendor, an engineering pager gives you a woken engineer refreshing a status page. What you need there is a comms rota and a documented degraded mode, not an escalation policy.
- Your users are asleep when your service is. A platform whose load is entirely in office hours does not need out-of-hours engineering cover for the change path. Say that plainly and spend the money on making the working day better instead.
One guardrail for the other direction, and it is the one I would be most careful about. A platform team with a good pager and a reputation for calm competence becomes the place every unclaimed production problem drifts toward. It starts as helpfulness and it ends with the platform team as the escalation path for the whole estate, which is precisely the central-ops bottleneck the discipline was invented to dissolve. Page yourself for your own services. Provide the tools, the runbook template and the enablement for everyone else to run theirs, in the X-as-a-Service shape rather than a permanent rescue mode. The moment your rota is carrying somebody else’s service, you are not a platform team any more.
The bottom line
Every argument platform engineering makes about internal products applies to the platform itself, and on-call is where that gets tested. The awkward truth is that we exempt ourselves quietly, not deliberately: the failures do not look like outages, the users tell us politely in a channel, and there is always something more visible to build than a rota nobody will thank you for.
So do the small version first. Build the robot that deploys a hello-world service through the real road every half hour and pages when it cannot. Tier your own services and write down which ones justify waking someone. Separate questions from breakage. Publish what people can expect, including the part where the answer is business hours. Declare an incident when the road breaks, review it like you mean it, and ship template changes the way you would ship anything else into production.
And ask the question that this whole post turns on: if the road broke at 2am tonight, how long until someone knew, and would that person be a member of your team or a customer? If the honest answer is a customer, you are not running a platform. You are hosting one and hoping.
If you are standing up a platform team’s operational practice and cannot work out how to page six people fairly, or you have a road that everybody depends on and nobody watches, that is exactly the kind of problem I help teams work through. Let’s talk.