You Build It, You Run It. With What?

TL;DR - “You build it, you run it” is usually announced rather than equipped. Teams get a pager, a link to a wiki and a vague expectation, and the first time something hard breaks at 2am they call the platform team, because the platform team is the only group that understands the substrate. Help once and you are the escalation path forever. The fix is to treat operability the way you already treat deployment: as a paved-road default. Every service created on the road should arrive with a rota, alerts on the user journey, a dashboard, production access that works and a generated runbook that is true on day one. Split the runbook into the generic half the platform owns and generates, and a short specific half the team owns and drills. Offer enablement as a time-boxed engagement with an exit date, not as a standing rescue service. And make every rescue leave something behind in the kit, so the second call for the same problem never has to happen.

I build and run platforms for a living, and in nearly every organisation I have worked with, the phrase “you build it, you run it” is on a slide somewhere. It is a good principle. The team that changes a service is the team best placed to understand why it broke, and the feedback from being woken up is the most honest code review there is.

It is also, in most places, a principle that was declared and never delivered. The slide happened. The pager rotation was created. And then a product team with six engineers, none of whom have ever debugged a node running out of file descriptors, was told that production is theirs now.

Last week I wrote about who pages the platform team, and I ended it with a guardrail I deliberately did not develop: a platform team with a good pager becomes the place every unclaimed production problem drifts toward. This post is about that drift, and about the other half of the closing argument, the part that says “provide the tools, the runbook template and the enablement for everyone else”. That sentence is doing a lot of work, and it deserves more than a clause.

Ownership by announcement versus operable by default

The tell

One question settles whether a stream-aligned team actually runs its service. When their service broke badly the last time, who got called in the first thirty minutes?

If the answer is someone on the platform team, “because they know Kubernetes”, then you have not moved operations to the teams. You have moved the pager to the teams and kept the operations. The platform team is still the de facto operations function for the estate, it is just reached through a side door, unscheduled, and without any of the capacity or authority a real operations function would have.

That is the DevOps-engineers-as-a-human-bottleneck outcome the whole discipline was supposed to dissolve, rebuilt from goodwill. Nobody designed it. It grew one kind favour at a time.

Why the gravity always pulls toward you

It is worth being precise about the mechanism, because it is not laziness on either side.

  • The hardest incidents live in the substrate. A service’s own bugs are usually diagnosable by the team that wrote it. The incidents that genuinely frighten a product team are the ones where the answer is in the layer underneath: DNS, a noisy neighbour, a certificate, a connection pool against a managed database, a sidecar. That layer is yours. They are not wrong to call you.
  • The first incident is on unfamiliar ground. A team that has been on call for three weeks has not built the muscle yet. The first real outage is the one where they are least equipped, most anxious, and most likely to pull in whoever sounds confident.
  • Helping is rewarded and fixing the cause is invisible. The engineer who joins a bridge at 2am and saves the night gets thanked in public. The engineer who spends two days making sure that class of problem is self-diagnosable next time gets nothing, because nothing happens.
  • Every rescue is a precedent. Once you have been called once and it worked, you are the fastest path to resolution, and the next team hears about it. Within a year there is an unwritten second tier of on-call and it is your team, unpaid and unplanned.

That last point is where it compounds. It is the same interrupt load that quietly eats a platform team’s second year, except this version arrives at night and does not show up in any queue.

Ownership without kit is a tax

Here is the claim. Handing a team the pager without handing them the capability to answer it is not ownership. It is the same mistake as shipping a security rule without a road that satisfies it: a responsibility with no delivery mechanism, which people will route around, and the route they find leads to you.

What a team actually needs to run a service is fairly well understood, and almost none of it is specific to their service:

  • A rota that exists in the paging tool, with escalation and overrides, before the first alert fires.
  • Alerts on the user journey, with burn-rate paging rather than a raw threshold, for the reasons I went through in the error budget nobody spends.
  • A dashboard that answers “is it us?” in under a minute: the service’s own signals next to the health of the things it depends on.
  • Production access that works at 2am. Read access to logs, traces and the orchestrator, provisioned in advance, not requested through a form during the incident.
  • A way to declare an incident that creates the channel, sets the severity, and posts to the status page, in one command.
  • A runbook that is true.

Look at that list and notice that it is a template. It is the same six things for nearly every service on a given road. Which means the platform team can generate most of it, and should, the same way it already generates the pipeline and the Dockerfile. If a service created from your first golden path arrives with a deploy pipeline and without alerting, you have paved half the road and left the other half as an exercise.

So the practical rule is: operable by default. A service that comes off the scaffold command comes with its rota stub, its alert rules, its dashboard, its access groups and its runbook skeleton, all wired to the same primitives as everything else. Teams can change any of it. Nobody has to invent it.

Runbooks are mostly fiction, so split them

Most runbooks I have read fall into two categories. The ones written in a burst of enthusiasm the week the service launched, and never updated since. And the ones that say “check the logs and escalate to the platform team if needed”, which is not a runbook, it is an escalation policy with extra steps.

The fix is not a better runbook template. It is noticing that a runbook contains two very different kinds of knowledge, owned by two different teams, and decaying at two different rates.

  • The generic half belongs to the platform, and it is generated. How to deploy. How to roll back. Where the logs, traces and metrics live. How to scale out. How to restart safely. How to see what this service depends on and what depends on it. How to get production access. How to declare an incident. None of this is specific to the service, all of it changes when the platform changes, and all of it should be rendered from the platform’s own source of truth so that it is correct by construction. If the rollback command changes, every runbook changes with it, the same day.
  • The specific half belongs to the team, and it is short. What this service does, in two sentences. What it is allowed to be down for, and who notices first when it is. The three or four failure modes the team has actually seen, with what they looked like and what fixed them. Which dependencies are hard and which it can degrade without. That is one page, and a page is maintainable.

Two tests keep the specific half honest. First, can someone who has never seen this service get through the first fifteen minutes of an incident using only the runbook? If they need to phone the author, it is not a runbook yet. Second, has anyone used it on purpose? A runbook that has never been exercised is a guess, for the same reason an environment you have never broken on purpose is lying to you. Put a last-exercised date at the top and let it look embarrassing.

One warning while we are here. Do not solve the runbook problem by building a portal page for it. A runbook lives next to the code, in the repository, where it gets reviewed with the change that makes it wrong. The developer portal trap is especially seductive here, because a grid of services with a green “runbook: yes” badge looks exactly like operational maturity and measures nothing.

Enablement is a mode with an exit date

Generated kit gets a team most of the way. The rest is skill, and skill has to be taught, which is where platform teams tend to make the opposite mistake: they either refuse to teach anything (“it’s in the docs”) or they never stop.

The vocabulary from Team Topologies in practice is useful here, because it names what a healthy version looks like. Enablement is a temporary interaction mode. It has a start, a purpose and an end, and then the relationship returns to X-as-a-Service. What goes wrong is that the temporary mode never ends, and “enabling” turns into a permanent rescue arrangement that nobody signed up for.

What a bounded version looks like in practice:

  • Shadow the first on-call week. When a team takes the pager for a service for the first time, someone from the platform team is a named secondary for that week only. Not the primary, not the fixer. The person who sits next to the primary and asks “what would you check next?”.
  • Run one game day on their service. Break something realistic in a controlled window: kill a dependency, fill a disk, expire a certificate. Let the team diagnose it with their own runbook. You learn more about the gaps in your kit in two hours than in a quarter of support tickets.
  • Write the exit criteria before you start. “The team has handled two real pages without us, the runbook has been exercised, and the dashboard answered ‘is it us?’ both times.” When they are met, you leave, and you say that you are leaving.
  • Cap it. One team at a time, or two. Enablement that is offered to everyone at once is support by another name.

The second-call rule

This is the one habit that, more than any tooling, stops the drift.

The platform team helps with a class of problem once. The second time, the help arrives as a change to the kit. A new panel on the generated dashboard. A new section in the generic runbook. A better error message from the thing that failed. A new alert. A diagnostic command that answers the question the team had to phone you to ask.

It is the operational version of the argument that the fifth identical question is a product bug rather than a documentation bug. Every rescue is a free requirements document for the kit. If you rescue and ship nothing, you have spent your night teaching one engineer something that the next team will have to call you about again.

This pairs naturally with the reviews. When a stream-aligned team’s incident review says “we called the platform team and they found it”, that is a finding about the platform, not about the team, and the action belongs to you. Reviews that let that sentence pass without an owner are the reviews in the postmortem that changed nothing.

Draw the escalation boundary before the first night

The last piece is a contract, and it should be written down while everyone is calm.

The platform team is escalated to when something the platform owns is implicated: the road, the runtime primitives, the shared networking, the secret store, the base images. That is a real promise with a response time, and it is backed by the platform’s own pager. The platform team is not escalated to because the service is behaving strangely and nobody on the owning team knows why. That is the owning team’s incident, and the platform’s contribution is the kit that lets them find out.

The honest middle is triage. Sometimes the team genuinely cannot tell whether it is them or the substrate, and making them prove it before you will look is hostile. So offer a bounded triage promise instead: within a stated number of minutes, someone will tell you whether the platform is involved. If it is, it is now your incident. If it is not, you hand it back with whatever you saw, and the “is it us?” dashboard gets a new panel the next day.

Then measure the boundary, because it will erode quietly otherwise. Two numbers are enough: the share of platform escalations where the platform was not at fault, and rescues per team per quarter. If the first is high, your triage kit is not good enough. If the second is not falling for a team you enabled, the enablement did not work, and it is worth finding out why rather than just continuing to answer.

When this is the wrong answer

The off-ramps, because “every team runs its own service” is not a universal law.

  • You have twenty engineers. One shared on-call rotation across the whole engineering team is correct at this size, and splitting it by service gives you four rotas of five people each, all exhausted. The same instinct as not running Kubernetes for three services applies: generate the kit anyway, because it is cheap and it will pay back later, but do not fragment the pager.
  • The teams are two people each. A rota needs about five or six people to be humane. Tiny stream-aligned teams should pool into shared rotations across related services, with the generated runbooks doing the work of making someone else’s service operable. That is still “you run it”, at the level of a group rather than a squad.
  • Regulation says otherwise. Some environments genuinely require a separate operations function with production access that developers do not have. Then the principle becomes “you build it, you are on the bridge when it breaks”, and the kit is aimed at the operations team instead. Do not fight the regulator with a slogan.
  • The platform itself is not reliable yet. If most real incidents are in the substrate, the teams are calling you because it really is you. Fix the road first, and give yourself a pager of your own before you ask anyone else to hold one.
  • There is no road to generate from. If services are still hand-built, “operable by default” has nowhere to live. That is a path-one problem, not an enablement problem, and teaching teams to operate snowflakes one at a time is how you end up as their permanent operations team.

One guardrail for the other direction, because I have also seen this go wrong the other way. “You build it, you run it” is sometimes used as a polite way of dumping operational load onto product teams with no support at all, and then blaming them for their incident rate. Refusing every escalation is not independence, it is abandonment, and teams that are abandoned will build their own shadow platform out of necessity, which is sprawl by a different route. The goal is not that nobody ever calls you. It is that nobody calls you twice for the same thing.

The bottom line

Operational ownership is a capability, not a declaration, and it only moves to the teams if the platform ships it the same way it ships everything else. Most of what a team needs to run a service is identical across services, which makes it a template, which makes it the platform team’s job to generate and keep true. What remains is short, specific, and owned by the team, and it only stays true if someone uses it on purpose.

So do the boring version. Make every service operable on the day it is created: rota, journey alerts, a dashboard that answers “is it us?”, working access and a generated runbook. Keep the team-owned half to a page and drill it. Offer enablement one team at a time with an exit date. Help once, and ship the second answer as a change to the kit. Write the escalation boundary down, promise fast triage rather than open-ended rescue, and watch the two numbers that tell you whether the boundary is holding.

And ask the question this whole post turns on: if a product team’s service broke tonight, would they solve it with what we gave them, or with our phone numbers? If it is the phone numbers, you have not handed over operations. You have handed over a pager and kept the job.

If you are moving operational ownership to stream-aligned teams and your platform team has quietly become everyone’s second-tier on-call, that is exactly the kind of problem I help teams work through. Let’s talk.