<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>On-Call on F. Latini - IT Engineer</title><link>https://latini.dev/tags/on-call/</link><description>Recent content in On-Call on F. Latini - IT Engineer</description><generator>Hugo</generator><language>en</language><lastBuildDate>Thu, 24 Sep 2026 08:00:00 +0000</lastBuildDate><atom:link href="https://latini.dev/tags/on-call/index.xml" rel="self" type="application/rss+xml"/><item><title>You Build It, You Run It. With What?</title><link>https://latini.dev/posts/you-build-it-you-run-it-with-what/</link><pubDate>Thu, 24 Sep 2026 08:00:00 +0000</pubDate><guid>https://latini.dev/posts/you-build-it-you-run-it-with-what/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; - &amp;ldquo;You build it, you run it&amp;rdquo; is usually announced rather than equipped. Teams get a pager, a link to a wiki and a vague expectation, and the first time something hard breaks at 2am they call the platform team, because the platform team is the only group that understands the substrate. Help once and you are the escalation path forever. The fix is to treat operability the way you already treat deployment: as a paved-road default. Every service created on the road should arrive with a rota, alerts on the user journey, a dashboard, production access that works and a generated runbook that is true on day one. Split the runbook into the generic half the platform owns and generates, and a short specific half the team owns and drills. Offer enablement as a time-boxed engagement with an exit date, not as a standing rescue service. And make every rescue leave something behind in the kit, so the second call for the same problem never has to happen.&lt;/p&gt;</description></item><item><title>Who Pages the Platform Team?</title><link>https://latini.dev/posts/who-pages-the-platform-team/</link><pubDate>Thu, 10 Sep 2026 08:00:00 +0000</pubDate><guid>https://latini.dev/posts/who-pages-the-platform-team/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; - The platform team is the one team in engineering that runs a production service for internal customers and usually has no pager, no objective and no incident process for it. The reason is that platform incidents do not look like outages: they block change rather than break traffic, so nothing on a dashboard turns red while forty teams sit unable to deploy. Fix it in this order. Page on the journey, which means a synthetic service that builds and deploys through the real road every half hour and wakes someone when it cannot. Tier your own services honestly, because a CI outage at 3am on a Saturday is a Monday ticket and a broken secret store is not. Separate the pager from the interrupt shift, because answering questions and restoring a road are different promises. Publish what your users can expect and then hold it, including the part where the promise is business hours. And declare an incident when the road breaks, even though nothing is technically down, because the alternative is that your users quietly build a way around you.&lt;/p&gt;</description></item></channel></rss>