Your Platform Team Is Hiring for Year One

TL;DR - Almost every platform team I meet is hiring against the job description it wrote when the team was founded: Kubernetes, Terraform, a cloud certification, “built X from scratch”. That loop selects for one thing, the ability to build infrastructure, which is exactly what year one needed and exactly what a mature platform team already has too much of. By year two the work is mostly product judgement, support that shrinks itself, teaching, migration programmes and running a production service for internal customers, and none of that is in the interview. The result is a team of excellent builders doing a job that is mostly not building, and quietly inventing builds to stay sane. The fix is not a unicorn hire. Write the job description from last quarter’s calendar instead of last year’s ambition, test for the skills the calendar actually needs, hire some of your platform engineers from among your users, and change the career ladder so the unglamorous work is promotable. Otherwise the ladder will undo every hire you make.

I build and run platforms for a living, and over the last few months I have written about the shape of the job after the founding: why the second year is a run rather than a build, why the platform team needs its own pager, how to run a migration somebody else dated, and what it takes to actually equip teams to run their own services. Each of those posts ends up, sooner or later, in the same uncomfortable place. The work it describes is real, the organisation needs it done, and the people on most platform teams were not hired to do it.

That is not a criticism of the people. It is a criticism of the loop that selected them.

What the platform job ad asks for versus what the platform job actually is

The tell

Pull up your current platform engineer job description and put it next to an honest account of where the team’s time went last quarter.

The job description almost certainly says some version of: deep Kubernetes, Terraform modules at scale, CI/CD design, a major cloud provider, observability tooling, “a track record of building platforms from the ground up”. It might mention developer experience in a bullet near the bottom.

The calendar, if the team is past its first year, says something different. Answering questions in Slack. Rewriting an error message for the third time. Chasing eleven teams through a base-image bump. Running a game day for a team that just took the pager. Arguing, politely, about whether the data team’s request is a new road or a variant. Being on call for the pipeline. Writing the upgrade guide. Opening forty pull requests in repositories you do not own.

If the two documents describe different jobs, you are hiring for the year you have already finished.

Why the loop drifts toward builders

Nobody designs this. It happens for reasons that are each individually sensible.

  • The job description was written at the founding. When the team was created, the job genuinely was “build a platform”. The first job ad was accurate. Then it was copied for the second hire, and the fifth, and nobody re-derived it because the title did not change.
  • Builders are easy to interview. You can test Kubernetes networking on a whiteboard. You cannot easily test whether someone will notice that the fifth identical question is a product bug rather than a documentation gap. So the loop measures what it can measure, and over time it comes to believe that what it measures is the job.
  • Builders interview builders. A team assembled for its infrastructure depth will, with complete sincerity, recognise infrastructure depth as the mark of a strong candidate. The person with an unusual mix - three years on a product team, a habit of writing things down, a history of running a migration across a dozen teams - reads as “not quite senior enough technically”.
  • The market rewards the keyword. Candidates optimise their CVs for the job ads that exist, and the job ads that exist are infrastructure job ads. The pool you search is already filtered before you open it.

The cost shows up a year later, in a team that is very good at the twenty percent of the work it enjoys and visibly restless about the other eighty.

What the job actually needs

When I map the year-two calendar back to skills, five groups fall out, and only one of them is what most loops test for.

  • Substrate depth. The thing the job ad already asks for: Kubernetes, networking, IaC, the cloud provider, how the pieces fail. It is still essential. The road has to work, and somebody has to be the person who knows why the connection pool behaves like that under load. It is just no longer most of the job.
  • Product judgement. Deciding what not to build. Telling a variant from a new workload class. Running a user interview without leading the witness. Noticing that the clever abstraction is clever for you and expensive for them. This is the skill that keeps a platform from becoming an internal product nobody asked for, and it is almost never in the loop.
  • Support that shrinks itself. Not patience with tickets. The instinct to answer a question once and then remove the reason it was asked: the better error message, the default that makes the mistake impossible, the diagnostic command. In year two this is the highest-leverage skill on the team, because support per user is the only number you can actually drive down.
  • Teaching. Shadowing a first on-call week, running a game day, explaining why the road works the way it does to a sceptical senior engineer on another team. Enablement is a bounded mode with an exit date, and leaving on time is itself a skill. Most infrastructure engineers have never been asked to do it, and some are excellent at it, and you would never find out from the current interview.
  • Programme work. The migration across sixty services, the retirement with a tail, the version bump with forty pull requests. This is coordination, written communication, tracking and stubbornness, and it is the work most platform engineers describe, privately, as the part they were never hired for. They are right. They were not.

Look at that list and notice that a single person who is strong at all five is a person you will not find, and should not try to. The point is the mix across the team, not the profile of each hire.

Write the job description from the calendar

The fix starts with a document, and it is a cheap one.

Take the last two quarters. For each person on the team, estimate roughly where the time went, in the five groups above. Do not overthink the precision; you want the shape. Then add up the team and compare it with the skills the current members are actually strong in.

You will usually find two things. The team is heavily over-indexed on substrate depth, which is fine, because that is what it was built from. And there is one category, often programme work or teaching, that is being done reluctantly by whoever lost the argument, which shows up as the thing that is always late.

That gap is your next hire. Not “another senior platform engineer”, but a platform engineer whose strongest card is the thing the team is weakest at, and who is good enough at the substrate to be credible. Write the job description so that the gap is the headline and the substrate is the floor. “Has led a migration across many teams to completion” in the first three bullets will change who applies more than any amount of employer branding.

Two practical notes. Keep the substrate floor honest: a platform engineer who cannot debug the road will not be trusted by the people who can, and that matters for the team’s internal health. And say in the ad what the job actually is, including the support and the on-call. The candidate who withdraws because the ad mentions answering questions was going to be unhappy in month four anyway.

Test for it

If the job description changes and the loop does not, nothing changes. Some exercises I have found more predictive than another system-design round:

  • Write the error message. Give the candidate a real failure from your road, with the stack trace and the context, and ask them to write the message a developer should see. You will learn more about how they think about users from that than from an hour on scheduling.
  • Work a support ticket. A real, anonymised question from your queue. Watch whether they answer it, or answer it and then ask why it was possible to ask.
  • Plan the migration. “Sixty services are on a runtime that goes end of life in ten months. What do you do in the first two weeks?” The strong answer starts with the inventory and the owner of the date, not with the codemod.
  • Review a template change from the user’s side. Hand them a pull request to your golden path and ask what will break for the teams already on it. This is where you find the people who think about the road as a public API.
  • Explain something to a sceptic. Ten minutes, an interviewer playing a senior engineer on a product team who does not want to move to the road. You are not testing persuasion. You are testing whether they listen.

None of these replace a substrate round. They sit beside it, and the scoring rubric should say out loud that a candidate can be the strongest in the loop without being the strongest on Kubernetes.

Hire some of your users

The most reliable source of product judgement on a platform team is people who used to be its customers.

An engineer from a stream-aligned team who has shipped services on your road, been paged for them, and complained about the template in the open hour knows things about your platform that nobody on your team can see from the inside. They know which docs page everyone actually reads, which error message makes people give up, and which part of the road gets routed around and why. That is the same signal adoption without a mandate asks you to go and collect through interviews, already installed in a person.

It also works in the other direction. The Team Topologies interaction modes assume people move between team types over a career, and a platform engineer who spends a year embedded in a product team comes back with a far better sense of what “self-service” means than any survey will give you. In the second-year post I argued for moving people out rather than losing them when a team should shrink. Rotation in both directions is the cheap, permanent version of that.

The career ladder will undo all of this

Here is the part that makes most of the above pointless if you skip it.

Most engineering ladders reward scope, technical complexity and things you can point to. “Designed and built the new deployment platform” is a promotion case. “Reduced platform support questions per service by forty percent over three quarters” sounds, to a promotion committee calibrated on product engineering, like customer service. “Led the runtime migration to completion with no missed dates” sounds like project management. “Ran enablement for six teams who now handle their own incidents” sounds like nothing at all, because the evidence is an absence.

So you hire someone for their programme instincts, they do excellent programme work for a year, they go to promotion, and they are told to come back with something more technically ambitious. The lesson lands across the whole team within a week. Next quarter, somebody proposes a rewrite.

This is the mechanism behind the year-two failure where a platform team invents a project to prove it is still needed. It is not only a capacity problem. It is an incentive problem, and the ladder is where the incentive lives.

What to change, concretely:

  • Write platform-specific evidence into the ladder. Ratios moved, support deflected, migrations completed, teams enabled with an exit, incidents that stopped recurring. If the ladder does not name these, a committee will not recognise them.
  • Treat a finished migration like a launch. Same write-up, same visibility, same weight in a promotion case. The retirement post argued for announcing an ending like a launch for the users’ sake. It matters as much for the people who did it.
  • Ask the “what did this make cheaper?” question in calibration. It is the year-two review question for the team. It should also be the year-two review question for the individual.

If you cannot change the ladder, say so to candidates, and adjust what you are hiring for. Hiring people for work the organisation will not reward is a slow way to lose good people.

When this is the wrong answer

The off-ramps, because plenty of platform teams should keep hiring builders.

  • You are genuinely in year one. If the road does not exist yet, the job really is building it, and the founding job description is accurate. Hire substrate depth, plus one person with the product instinct to keep the team honest about what the first ninety days are actually for. Revisit when the first road has users.
  • You are three people. A three-person platform team cannot have a skills mix; it has three people who all do everything. Hire for range rather than any one strength, and do not formalise a five-way skills matrix for a team that fits around one desk.
  • The substrate is genuinely hard. If you run your own datacentres, a bespoke scheduler, or anything regulated down to the kernel, deep specialists are the job and will stay the job. Most organisations are not that, and most that think they are have chosen a substrate harder than the problem needed.
  • The company is doubling every year. New workload classes keep arriving, the build agenda is real for several years running, and under-building is the bigger risk. Hire builders, and keep one eye on the support curve.
  • You want to fix it by hiring a platform product manager. A good PM helps. A PM as a substitute for product instinct on the team produces a team that builds what it is told and a PM who gets the blame. Hire the PM if the team is large enough, but not instead of changing the loop.

One guardrail for the other direction, because I have watched it happen. A platform team that over-corrects into product, support and coordination, and stops hiring substrate depth, becomes a ticket concierge in front of infrastructure nobody on the team can debug. The road degrades slowly, the platform’s own pager gets louder, and the first serious incident gets escalated to a vendor. The goal is not to stop hiring engineers. It is to stop hiring only one kind.

The bottom line

A platform team is a product team that happens to build infrastructure, and the hiring loop of most platform teams still describes an infrastructure team that happens to have users. That mismatch is invisible in year one, when the job really is building, and expensive every year afterwards, when the job is mostly product judgement, support that shrinks itself, teaching, migration programmes and running a service.

So do the boring version. Put last quarter’s calendar next to the job description and see whether they describe the same job. Make the team’s weakest skill the headline of the next hire, with substrate depth as the floor rather than the whole ad. Add exercises that test what the calendar needs: an error message, a support ticket, a migration plan, a template review, a sceptic. Hire some of your users, and lend some of your engineers out. And change the ladder, because a team will become whatever its promotions reward, regardless of who you hired.

The question the whole thing turns on is simple: if your best year-two platform engineer applied for their own job today, would your loop hire them? If the honest answer is “probably not, they’re not deep enough on Kubernetes”, the loop is screening out the people the platform most needs.

If you are hiring for a platform team and suspect the job description stopped matching the job somewhere around last spring, that is exactly the kind of problem I help teams think through. Let’s talk.