I’m a Platform and Site Reliability engineer. I design, build, and run the cloud infrastructure and developer platforms that distributed, event-driven systems sit on - mostly on GCP, some AWS, managed as code with Terraform.
Day job: Site Reliability Engineer at AxonIQ. The rest of the time I consult for engineering teams that have outgrown the “one DevOps person heroically holding it together” stage, and I build my own products - NutriFinder, InPedana, and OSDb - which is where most of the opinions below come from.
What I write about here
Platform engineering, mostly, and the parts of it people skip. Golden paths and who actually walks them. Why a developer portal is usually premature. Error budgets nobody spends and DORA dashboards nobody acts on. Where the boring stack beats the impressive one, and the specific cases where it does not. Every post tries to earn its thesis and name the situation where it is the wrong answer.
I run Kubernetes for a living and ship my side projects on Flask, HTMX, and a single VM. Both of those are the same opinion, not a contradiction. The databases underneath started on SQLite and moved to Postgres once the product gave me a concrete reason, which is that same opinion a third time.
If any of this is a problem you’re currently living with, here’s how to work with me.
> Articles by Francesco Latini
- · ~13 min · 2,736 words LATEST
# Your Platform Team Is Hiring for Year One →
Most platform teams hire against an infrastructure job description long after the job stopped being infrastructure. Why the interview loop selects for builders when year two needs product, support, teaching and programme instincts, how to rewrite the job description from the team's actual calendar, what to test for instead, and why the career ladder quietly undoes all of it.
- #platform-engineering
- #hiring
- #engineering-management
- #org-design
- #developer-experience
- · ~14 min · 2,912 words
# You Build It, You Run It. With What? →
Most organisations hand stream-aligned teams the pager and call it ownership, then watch every hard incident drift back to the platform team because they are the ones who understand the substrate. Why operational ownership has to ship as a paved road rather than a slogan, how to split runbooks into the part the platform generates and the part the team owns, and how to help on the first incident without becoming the escalation path for the whole estate.
- #platform-engineering
- #site-reliability-engineering
- #on-call
- #runbooks
- #developer-experience
- · ~17 min · 3,553 words
# The Migration With a Date on It →
A runtime end of life, a datacentre exit, a cloud service retirement: the migration you did not choose, with a deadline you cannot move. Why these fail as governance rather than as engineering, how to find the real internal date, why the inventory is the critical path, and how a platform team turns sixty bespoke migrations into one migration applied sixty times.
- #platform-engineering
- #migration
- #engineering-management
- #internal-developer-platform
- #devops
- · ~15 min · 3,186 words
# Who Pages the Platform Team? →
Platform teams sell on-call, SLOs and incident practice to everyone else, then run their own service on hope and Slack. Platform incidents have a different shape: they block change rather than take things down, they fan out across every team at once, and the users have a workaround that costs you more than the outage. How to page on the road rather than the components, what a platform SLO can honestly measure, and how to staff a rota with six people.
- #platform-engineering
- #site-reliability-engineering
- #on-call
- #incident-response
- #developer-experience
- · ~16 min · 3,327 words
# The Platform Team's Second Year →
Year one of a platform team is a build. Year two is a run, and almost nobody re-plans for it. The maintenance ratio nobody budgets, why success is what generates the load, how to put the run cost in the plan in public, and the year-two failure mode nobody names: inventing a project to prove the team is still needed.
- #platform-engineering
- #engineering-management
- #developer-experience
- #internal-developer-platform
- #org-design
- · ~18 min · 3,759 words
# How to Retire a Golden Path →
Every paved road eventually has to be closed, and almost nobody plans for it. Deprecating a golden path is a product launch in reverse: it needs a destination that already works, a migration you wrote yourself, a funded owner for the last twenty percent, and a real turn-off date. The half-finished deprecation is the most expensive outcome available to a platform team.
- #platform-engineering
- #golden-paths
- #developer-experience
- #internal-developer-platform
- #migration
- · ~18 min · 3,652 words
# Platform Adoption Without a Mandate →
You built a good paved road and half the teams still went around it. Adoption is not a persuasion problem or a compliance problem - it is a product problem, and a team routing around your platform is telling you something a mandate would have hidden. The four reasons teams say no, the switching-cost arithmetic nobody does, and how to win the moment of creation instead of the argument.
- #platform-engineering
- #developer-experience
- #golden-paths
- #internal-developer-platform
- #engineering-management
- · ~16 min · 3,237 words
# The Platform Team's First 90 Days →
A new platform team's first ninety days decide whether it gets a second year, and almost every team spends them building something to prove it exists. What the first quarter should actually produce: a written mandate including the no, baseline numbers you can only measure now, the inherited ticket queue read as a requirements document, and one small painful thing fixed by week six.
- #platform-engineering
- #engineering-management
- #org-design
- #developer-experience
- #internal-developer-platform
- · ~16 min · 3,344 words
# Your Cloud Bill Is Not a Finance Problem →
Cloud cost is a feedback problem, not an accounting one. The invoice arrives monthly, aggregated, weeks late, addressed to someone who cannot change a line of it - while the people who chose the instance size never see a number at all. How to make cost a platform capability: attribution as a paved-road default, unit cost instead of total spend, cheap defaults, and showback before chargeback.
- #platform-engineering
- #finops
- #cloud-cost
- #developer-experience
- #internal-developer-platform
- · ~16 min · 3,317 words
# Your Security Policy Needs a Paved Road →
Secrets management and policy-as-code fail in a specific way: the platform team ships the rule without shipping the road that satisfies it. A rule without a road is a tax on developers, and taxes get evaded. How to turn security controls into paved capabilities - workload identity instead of distributed secrets, observe-warn-enforce rollouts, and exceptions that expire.
- #platform-engineering
- #secrets-management
- #policy-as-code
- #devsecops
- #internal-developer-platform
- · ~13 min · 2,570 words
# Build vs Buy Is the Wrong Question →
The platform build-vs-buy spreadsheet almost always lies, because it compares the sticker price of a license against the sticker price of construction and forgets that a thing you build is a thing you operate forever. The real decision isn't build or buy - it's where you draw the line. Buy the undifferentiated substrate, build only the thin layer of opinion that is actually yours, and fund the result like a product, not a project.
- #platform-engineering
- #build-vs-buy
- #internal-developer-platform
- #platform-funding
- #engineering-leadership
- · ~13 min · 2,680 words
# The Error Budget Nobody Spends →
An SLO is only a reliability practice if the error budget under it is a decision you actually make. Most teams set a target from ambition, measure the wrong thing, and never change a single plan when the budget runs out. How to pick a target with slack, measure it on the user's journey, and write the policy that makes a blown budget mean something.
- #site-reliability-engineering
- #slo
- #error-budget
- #platform-engineering
- #observability
- · ~5 min · 1,058 words
# The Real Project Was Never the App →
NutriFinder is the friendly face, but the real project sits one layer down: OSDb, the Open Supplements Database - an open, normalised reference for sports-nutrition and supplement data, refreshed by an LLM-powered Playwright pipeline and heading toward a public API anyone can build on.
- #side-projects
- #sports-nutrition
- #open-data
- #data-pipeline
- #product
- #api
- · ~12 min · 2,357 words
# Golden Paths, Part 2: The Second and Third Path →
The first golden path is an adoption project. The second and third are a portfolio problem, and the dangers invert: sprawl, forked templates, and a platform team that says yes to everything. How to pick the next road, reuse the first one's guts, and keep the paved-road count honest.
- #platform-engineering
- #golden-paths
- #developer-experience
- #internal-developer-platform
- #platform-as-product
- · ~7 min · 1,367 words
# The Product I Wish Existed When I Started Racing →
I built NutriFinder, a neutral place to browse and compare endurance-sports nutrition on the numbers that actually matter - carbs, sodium, caffeine, price per portion - instead of on brand marketing. What it is, why it exists, and the open-data idea underneath it.
- #side-projects
- #sports-nutrition
- #endurance-sports
- #product
- #data-pipeline
- #flask
- · ~11 min · 2,153 words
# Team Topologies in Practice →
Everyone adopts the four team types from Team Topologies and ignores the three interaction modes, where the whole value lives. How to organise a platform team around cognitive load instead of the org chart.
- #team-topologies
- #platform-engineering
- #org-design
- #developer-experience
- #engineering-management
- · ~7 min · 1,363 words
# The Comparison That Refuses to Compare →
The hard part of building a 'compare two athletes' feature was not the chart. It was the honesty: a rank only means something inside its own cohort, so the feature refuses to crown a winner unless the two actually competed in the same field. A field note on software that says 'I don't know' out loud, from inpedana.com.
- #side-projects
- #data-visualization
- #honest-metrics
- #sqlite
- #flask
- · ~7 min · 1,409 words
# DORA Metrics Without the Dashboard Theatre →
The four DORA metrics are a thermometer, not a treatment. Most teams chart them, watch the lines go up, and change nothing about how software actually ships. How to read each metric honestly, the four ways teams game them, and the one question that tells a real DORA practice apart from a vanity dashboard.
- #platform-engineering
- #dora-metrics
- #developer-experience
- #engineering-metrics
- #site-reliability-engineering
- · ~8 min · 1,690 words
# Your Developer Portal Is a Trap (Until It Isn't) →
Backstage is the most popular way to make platform engineering look real before it is. A portal is the index of your paved roads - and writing the index before the book is the classic first-six-months mistake. When a developer portal earns its keep, what to put in it, and the cheaper version you almost certainly want first.
- #platform-engineering
- #developer-portal
- #backstage
- #internal-developer-platform
- #developer-experience
- · ~11 min · 2,193 words
# Learn Terraform Fluently →
Knowing terraform apply is not knowing Terraform. A pragmatic guide to the mental model that makes infrastructure-as-code fluent: state, the plan loop, modules, drift, and how to organise it all - plus when to put the tool down.
- #terraform
- #infrastructure-as-code
- #platform-engineering
- #devops
- #opentofu
- · ~11 min · 2,274 words
# Building Your First Golden Path →
A platform is a slide deck until the first golden path ships. How to pick, build, and get adoption for the one paved road that makes platform engineering real.
- #platform-engineering
- #golden-paths
- #developer-experience
- #ci-cd
- #internal-developer-platform
- · ~10 min · 1,953 words
# Your LLM Platform Is Repeating Your ML Platform's Mistakes →
An SRE and platform-engineering take on why most LLM platforms are quietly rebuilding the anti-patterns that made ML platforms painful in 2018-2022, and the small set of decisions that break the loop.
- #platform-engineering
- #sre
- #llm
- #ai-infrastructure
- #mlops
- #reliability
- · ~9 min · 1,861 words
# The Postmortem That Changed Nothing →
Most postmortems produce a document, not a change. A pragmatic SRE take on why incident reviews fail to fix anything, and the small set of habits that turn them into real learning.
- #sre
- #platform-engineering
- #incidents
- #postmortem
- #reliability
- · ~9 min · 1,834 words
# Your Staging Environment Is Lying To You →
Most staging environments are expensive theatre. A pragmatic SRE take on what staging is actually for, why it always drifts from production, and the production-safety habits that pay off more than fidelity ever will.
- #sre
- #platform-engineering
- #staging
- #production
- #feature-flags
- #canary
- · ~8 min · 1,660 words
# The Boring Stack Manifesto →
The SRE who runs Kubernetes for a living ships side projects on Flask, SQLite, and HTMX. A manifesto for the boring stack, with inpedana.com as the live case study.
- #flask
- #sqlite
- #htmx
- #platform-engineering
- #side-projects
- #simplicity
- · ~6 min · 1,133 words
# Are You Really Monitoring Your Infrastructure? →
Most teams confuse 'we have dashboards' with 'we have observability'. A practical guide to monitoring that actually catches problems before your customers do.
- #observability
- #monitoring
- #sre
- #platform-engineering
- #prometheus
- · ~5 min · 919 words
# Is Kubernetes the Right Tool for You? →
Kubernetes is brilliant when you need it and a tax when you don't. A pragmatic decision framework - with honest alternatives - from someone who runs Kubernetes for a living.
- #kubernetes
- #platform-engineering
- #cloud
- #architecture
- #sre
- · ~6 min · 1,122 words
# DDDD - Domain Driven Design for Dummies →
A practical introduction to Domain Driven Design: bounded contexts, ubiquitous language, and the tactical patterns that actually matter - without the 500-page book.
- #domain-driven-design
- #architecture
- #software-design
- #microservices
- · ~7 min · 1,320 words
# What Is Platform Engineering? A Beginner's Guide →
Platform Engineering explained: what it is, how it differs from DevOps and SRE, and why it matters for building scalable developer platforms.
- #platform-engineering
- #devops
- #sre
- #kubernetes
- #developer-experience
- · ~3 min · 427 words
# gnome-control-center keeps crashing on Fedora 35 →
Debugging a gnome-control-center crash on Fedora 35: tracing it back to corrupted shared libraries via journalctl, and fixing it with dnf reinstall.
- #fedora
- #gnome
- #linux
- #debugging
- · ~1 min · 201 words
# The new website is alive →
First post on latini.dev - kicking off a writing habit alongside freelance Platform Engineering work. Hugo on a DigitalOcean droplet in Germany.
- #meta
- #hugo
- #writing
fl@latini.dev:~/authors/francesco-latini/$
