The Migration With a Date on It

TL;DR - Sooner or later something external sets you a deadline: a runtime goes end of life, a datacentre lease ends, a managed service is retired, a certificate authority stops being trusted. This is not a deprecation you are choosing, and almost every instinct that works for a paved road you decided to close works badly here. The date is not negotiable, so scope is the only variable you have. The inventory, not the migration, is the critical path, and it is the thing teams start last. The real internal date is one to two quarters before the vendor’s, and if you publish the vendor’s date you will finish after it. Buying time is a legitimate engineering strategy and gets dismissed for aesthetic reasons. The single most common way these projects die is turning into a rewrite, because everything is being touched anyway. And a forced migration ends with the estate exactly where it started, on purpose, which is why it needs a named senior owner rather than a platform team squeezing it in between tickets.

I build and run platforms for a living, and I have now been through enough of these to recognise the opening scene. Somebody forwards an email. A version you depend on has a support end date roughly eleven months out, or the lease on a facility ends in eighteen, or a cloud provider has decided that a service two hundred of your workloads sit on will be switched off. Everybody agrees it is important. Somebody says “we should get ahead of this”. The thread goes quiet, and nine months later it is an emergency with a hard stop.

When I wrote about retiring a golden path on purpose, I deliberately drew a line around this case and pushed it into an off-ramp: an externally dated forcing function is not a deprecation, it is a migration project, and it needs project governance. That off-ramp has been bothering me ever since, because “it needs project governance” is the kind of sentence that sounds like an answer and is actually a shrug. This post is the part I left out.

A forced migration run as a platform side quest versus run as a dated programme

The date is the only thing you do not control

Every other migration you will run has three adjustable dials: scope, quality and date. In a normal deprecation the date is the softest of the three, which is why so many of them quietly slip into permanence.

A forced migration removes the dial you are used to reaching for. The vendor does not care about your quarter. The lease does not care that Q3 was heavy. Which means the entire character of the project changes:

  • Scope becomes the only real variable. Not “what do we move”, because everything on the old thing has to leave. What you actually control is what leaves as a migration, what leaves as a deletion, and what leaves as something you consciously buy an extension for.
  • Urgency is free and value is absent. A paved-road retirement you chose needs you to manufacture urgency, and it has a story about why the destination is better. This one has urgency handed to you and no story at all. At the end of it, the estate runs exactly as it did before, on a supported version, and nobody’s day got better. That absence shapes everything: staffing, morale, and how the work gets read in a performance cycle.
  • The failure mode is a cliff, not a drift. A half-finished deprecation leaves you running two roads forever, which is expensive and survivable. A half-finished forced migration leaves you on the wrong side of a date with a production dependency that is unsupported, unpatched, or switched off. Those are different categories of bad.

That third point is worth being blunt about internally, because it is the only leverage the project has. “We will be running an unpatched runtime in production, and the next CVE in it has no fix” is a sentence that moves a quarter. “We should really get on with the upgrade” is not.

It is not a platform project, even when the platform team does all the work

Here is the governance claim, stated properly rather than as a shrug.

A forced migration touches services owned by teams that do not report to you, in quarters that were planned without it, to produce an outcome none of those teams asked for. The platform team can build the mechanism, own the destination, write most of the code, and answer every question. What the platform team cannot do is take a fortnight away from a product team that has a launch. That is an authority the platform team does not have, and asking for a mandate to compensate for a product you have not made easy is the expensive mistake in normal times.

This is the exception I named in that post, and it is the only one I would defend: a dated external constraint is legitimately mandate-shaped. But the mandate has to be held by somebody who can actually reallocate people, and it comes with obligations in both directions.

What the project needs, minimally:

  • One named senior owner of the date. Not a working group. A person, senior enough to say “this team is not doing that feature this quarter”, whose name is on the outcome. If you cannot get one, you do not have a project, you have a recurring meeting.
  • A platform owner of the mechanism. That is you: the destination, the migration tooling, the compatibility layer, the docs, the escalations. Distinct role, distinct accountability.
  • A pre-agreed escalation path with a stated latency. “A team that has not booked a slot by week N gets escalated to their director within three working days.” Agreed before it is needed, while everyone is still calm and generous.
  • A standing trade, in writing. The platform team is about to spend a large fraction of two or three quarters on this. Say what is not happening because of it. This is the run, evolve and new split of a second-year platform team being spent in one direction, and if you do not name the cost, the year still ends with an unexplained shortfall and the org concludes you got slower.

Work backwards from a date that is not the vendor’s

The vendor’s date is a cliff edge, not a target. Targeting it guarantees you arrive late, because every single thing between you and it consumes time you have not modelled.

Work backwards instead. From the external date, subtract:

  • The change freeze. Most organisations have one, sometimes two. A December deadline in a retail company is really a mid-November deadline.
  • The soak. The last migrated service needs to run in production long enough for anything odd to appear. Two weeks minimum, a month if the workload is seasonal.
  • The tail. In every migration I have seen, the last ten to twenty percent of the estate takes as long as the first eighty. Not because anyone is lazy, but because the things left at the end are left for reasons.
  • The thing you find in month three. There is always one: a library with no upgrade path, a binary dependency nobody can rebuild, a vendor integration certified against the old version. Budget a month of unallocated time for it by name, so that when it appears it consumes budget rather than credibility.

What comes out is your internal date, and it will usually sit one to two quarters before the external one. Publish the internal date and never mention the other one again. The moment both dates are in circulation, every team plans against the later one, and you have given away your entire buffer to the teams least likely to need it.

The other half of that discipline: block creation on the old thing on day one. Remove it from the templates, the scaffold command, the module defaults, the docs. It costs almost nothing and it is the difference between a shrinking estate and a treadmill. Every service created on the old platform after the announcement is a service you will personally migrate later, usually in the worst week.

The inventory is the critical path

The most common structural mistake is treating discovery as a preliminary. It is not preliminary. It is the longest-lead item in the project and it should start the day the email arrives, before anybody has decided on an approach.

“Everything running Python 3.9” sounds like a query. In practice it is: the things in the service catalogue, plus the things not in the catalogue, plus the base images two layers down, plus the build-time dependency in a pipeline nobody owns, plus a Lambda deployed from a laptop in 2023, plus the vendor appliance that embeds it, plus the data job that only runs at quarter end and will therefore be discovered in month seven.

Generate the list from what the machines see, never from what people remember. Registry image layers, package manifests at build time, runtime inventory from the orchestrator, IAM principals still authenticating, network flows into the old facility, billing line items. A service catalogue is a hint, not an inventory, because a catalogue records what somebody registered.

Then sort what you find, because these buckets get completely different treatment:

  • Alive and actively developed. The easy majority. These move with tooling and a booked slot.
  • Alive and frozen. Nobody has touched it in a year, it serves real traffic, the team that wrote it has reorganised twice. Highest risk per service, because there are no tests and no living knowledge, which is exactly the situation where an environment nobody has broken on purpose turns out to be lying.
  • Running but dead. It is up, it costs money, nothing calls it. Do not migrate this. Delete it. In a large estate this bucket is routinely ten percent of the work, and switching it off is the cheapest progress available in the entire programme.
  • Genuinely unowned. Not a migration problem. An ownership escalation, and it needs resolving in week three rather than month eight, because the resolution takes longer than the migration does.
  • Not yours to move. A vendor appliance, a managed thing, an acquired company’s stack. This is a contract conversation and a long lead time, and the person who owns the date needs to know about it immediately.

A number worth tracking from week one: services with no booked migration date. Not percent complete. Percent complete flatters you early, moves fastest when the work is easiest, and says nothing about whether the end is reachable. The count of unbooked services tells you what is actually in front of you, and it is the number I would put on the weekly slide.

Make it a diff, not a memo

The platform team’s real contribution to a forced migration is a conversion: turn sixty bespoke migrations into one migration applied sixty times. Everything else is logistics.

Concretely, in roughly this order of value:

  • A codemod or an automated change, even a rough one. A script that gets a typical service ninety percent of the way there, leaving a human to review, is worth more than a perfect runbook. It also makes the work reviewable, which means it can be checked by somebody other than its author.
  • Open the pull requests yourself. The arithmetic is the same as in any adoption problem: the cost is theirs and now, the benefit is yours and later. A migration that arrives as a reviewable diff with a green pipeline is a ten-minute decision. A migration that arrives as a wiki page is a project on somebody else’s backlog.
  • A compatibility layer, if one is honestly available. A shim, a dual-write, a proxy that speaks both protocols. This decouples “move off the old thing” from “adopt the new thing properly” and buys you the ability to sequence.
  • A canary cohort and a staged rollout. The same discipline as shipping a golden-path change like a production deploy. Pick three friendly services, migrate them yourself end to end, and only then write the guide. The guide written before the first real migration is fiction.
  • One place that shows the state of the estate. Generated, not maintained by hand, and visible to the teams as well as the steering meeting. A list that disagrees with reality after week two is worse than no list.

If the estate was built on thin templates over shared primitives with the version stamped into every generated repository, the first three of those are days of work rather than months. If every service got a full copy of the build logic at creation time, you are now paying for that decision with interest, and it is worth saying so out loud while the lesson is expensive and vivid.

Buying time is a strategy, not a defeat

The option engineers routinely refuse to price: pay someone to move the date.

Extended support contracts, a vendor’s paid long-term-support tier, a third-party maintainer backporting security fixes, a negotiated lease extension, a period of running the old thing air-gapped and unchanged behind a proxy. These get dismissed quickly and usually on aesthetic grounds, because they feel like failure and because the invoice is visible while engineering time is not.

Run the actual numbers. Two quarters of senior engineering time across several teams is a large amount of money next to most support contracts, and the comparison is exactly the asymmetry I have complained about before: the invoice is visible and payroll is not, so the option with a price tag loses an argument it should win.

Two conditions make it honest rather than an evasion. Name what you are buying the time for, as a plan with dates, not as breathing room. And make it genuinely once. A second extension is not a strategy, it is a subscription to a problem that gets more expensive every year while the population of people who understand the old system shrinks.

Do not let it become a rewrite

This is how these projects die, and it kills good teams rather than bad ones.

Everything is being touched anyway. While we are in there, we may as well move to the new framework. And since we are changing the framework, this is the natural moment to split that service. And if we are splitting it, we should finally move it onto the new orchestrator. Each step is individually reasonable and the aggregate is a rewrite with somebody else’s deadline attached.

The rule I hold hard: a forced migration changes one variable. Same architecture, same language version policy, same deployment model, same boundaries, new substrate. If the old design is genuinely bad, that is a separate project with its own justification, run after this one, at a moment you choose. A migration and an improvement have different risk profiles, and merging them means the improvement inherits a date it cannot negotiate while the migration inherits a risk it did not need.

There is a softer version of the same failure, and it is the one I would watch for in a mature team: the migration becomes the cover story for building the thing the team wanted to build anyway. That is the invented project of a second-year platform team wearing a hard hat, and it is harder to challenge because the deadline is real even though the scope is not. The test is the same one: which number that we already track does this part move, and what happens on the vendor’s date if we skip it?

What you should get to keep

A forced migration produces no user-visible value, so the only durable return is capability, and you have to deliberately keep it. There will be another one. Runtimes go end of life every two to four years, cloud providers retire services continuously, and hardware leases keep ending.

What is worth extracting before the programme disbands:

  • The inventory pipeline. You just built the thing that answers “what is actually running, and on what version”. Keep it running and wire it into the platform rather than letting the spreadsheet rot.
  • Provenance in every generated repository. Path name and version stamped in at creation, so the next question of this shape is a query rather than an archaeology project.
  • The migration muscle. Codemods, bulk pull-request tooling, a canary cohort of friendly teams, a rehearsed comms pattern. This is the machinery that makes ordinary version bumps cheap for the rest of the platform’s life.
  • A written ending. What the estate looks like now, what was deleted, what was granted an exception and when that expires, what surprised you. Run the review like a postmortem that actually changes something rather than a completion email, because the next one of these lands on people who were not here.

The honest framing for the org is that you spent two quarters converting an unbounded liability into a bounded one and came out with tooling that makes the next occurrence cheaper. That is the whole story, and it is a good one. It is not a product launch and pretending it is will not help.

When this is the wrong answer

The off-ramps, because plenty of dated migrations should not become programmes.

  • Twelve services and one team. There is no governance layer here. There is a two-week spike and a shared calendar. Standing up a steering meeting, a burn-down and an escalation matrix for this is the same category error as running Kubernetes for three services.
  • The estate is converged and the paths are versioned. If everything runs through three paved roads with thin templates, a runtime bump is a version change in the template, a rebuild and a canary. That is a fortnight of the ordinary drift treadmill, not a project, and if that is your situation you should notice you are being repaid for earlier discipline.
  • The date is further away than the tooling debt. Sometimes the correct first move is not to start migrating, but to spend a quarter making the estate migratable: inventory, ownership, provenance, a way to open bulk pull requests. That quarter is not a delay. It changes the slope of everything after it.
  • “Forced” is being used to win an argument. Check the actual consequence on the actual date. Community support ending is not the same as a service being switched off, and a version falling off a compliance matrix is not the same as a security cliff. Some of these buy you a year and a different conversation. A migration justified by a date that turns out to be soft spends credibility you will want for the one that is real.
  • You do not have a platform yet. Then this is a cross-team engineering project and should be run as one, by whoever runs projects. Do not build a platform under cover of the migration. Path one deserves a design partner and a chosen problem, not a hostage estate and a deadline, and a platform whose origin story is a forced migration tends to encode that emergency in its defaults for years.

One guardrail for the other direction, since I have just argued for governance. The apparatus is not the work. I have watched a fourteen-month migration acquire a weekly steering meeting, a RAG status, a risk register and a dedicated programme manager, for an engineering effort that was three people and a codemod. Every hour in that room was an hour not spent opening pull requests. Scale the ceremony to the number of teams you have to coordinate, not to the scariness of the date.

The bottom line

The migrations you choose are product decisions and you control the pace. The migrations you do not choose are governance problems wearing an engineering costume, and they fail in governance ways: no owner of the date, an inventory that is still being argued about in month six, a buffer given away by publishing the vendor’s deadline, a tail nobody funded, and a scope that quietly grew a rewrite inside it.

So take the boring version. Start the inventory the day the email arrives and generate it from what the machines see. Set an internal date a quarter early and never say the real one out loud. Get one senior name on that date and write down what the platform team is not doing in exchange. Block creation on the old thing immediately. Turn the migration into a diff you open yourself, price the option of buying time honestly, and refuse every improvement that wants to ride along. Then keep the tooling, because this will happen again.

The question I would ask at the start of one of these: if the date were tomorrow, what would still be running on the old thing, and who would we call? If nobody in the room can answer that, the project’s first deliverable is not a migration plan. It is a list. And the whole premise of platform engineering as an internal product is that you should already know what your product is carrying, which is why the second one of these is so much cheaper than the first.

If you are staring at a hard external date with an estate you cannot fully enumerate, or you have a migration that has been at eighty percent for two quarters, that is exactly the kind of problem I help teams work through. Let’s talk.