The Estate You Cannot Enumerate

TL;DR - Ask most engineering organisations “what is running on the vulnerable version, and who owns each one?” and the honest answer is “give us two weeks”. Not because the question is hard, but because nobody owns the answer. The list gets rebuilt from scratch for every CVE, every forced migration, every audit, every cost sprint and every reorg, and thrown away afterwards. That is a platform capability being paid for as a recurring emergency. The fix is not a better catalogue, because a catalogue records what people declared, and the estate is whatever is actually running. Build the inventory from what the machines observe, treat the difference between declared and observed as the product, attach ownership to something that survives a reorg, make “unowned” a first-class state with a queue, and measure the whole thing by how long it takes to answer a question. Start with a nightly job and a table, not a portal.

I build and run platforms for a living, and the moment I trust least in any engagement is the first time I ask for a list. Not a diagram, not a strategy, just a list: every deployable thing, what it runs on, and the team that would pick up the phone. I have asked that question in organisations with forty engineers and in organisations with four thousand, and the answer has a remarkably consistent shape. A senior person says “we have that in the catalogue”, somebody else says “the catalogue is a bit out of date”, and a third person quietly opens a spreadsheet from the last time this came up.

When I wrote about running a migration somebody else dated, I argued that the inventory is the critical path of any forced migration, and that the pipeline you build to produce it is the single most valuable thing to keep once the programme disbands. I then left it there, as a bullet in a list. This post is that bullet, developed: what it looks like when the inventory stops being something you assemble in a hurry and becomes something the platform simply has.

Inventory as a recurring project versus inventory as a standing platform capability

The tell

Think back to the last time somebody needed the list urgently. A critical CVE in a library everyone uses, landing at four in the afternoon on a Friday. An auditor asking for every system that touches personal data. A finance review asking what the forty untagged accounts actually contain. A reorg, after which nobody is sure which of the old team’s services moved where.

Now answer three questions about that moment:

  • How long did it take to produce the first complete list? Not a partial one with a “we’re still checking” footnote. Complete enough that you would have bet on it.
  • What was it built from? If the answer includes a Slack thread asking teams to reply with what they run, it was built from memory.
  • Where is it now? If the answer is “a spreadsheet in somebody’s drive” or “I think it was attached to the ticket”, you paid for it in full and then threw it away.

If the first answer is measured in days and the third is a shrug, you do not have an inventory. You have a recurring emergency that produces one, briefly.

The same list, rebuilt five times a year

What makes this expensive is not any single occasion. It is that the occasions are frequent, they come from different directions, and each one looks unique to the people handling it.

  • The CVE. Security wants every deployable that contains a given package at a given version. The question is about the contents of images and build artefacts.
  • The forced migration. A runtime goes end of life, a managed service is retired. The question is about what runs on what substrate.
  • The audit. Compliance wants every system in scope for a regulation, with a named owner. The question is about classification and accountability.
  • The cost review. Finance wants to know what is behind the spend. The question is about attribution, and it is the reason the untagged resource is a platform bug rather than a team failure.
  • The orphan. Something is paging and nobody recognises it. The question is about ownership, and it arrives at the worst possible time.

Five different requesters, five different columns, and one underlying question: what exists, what is it made of, and who answers for it? Each time, a different team builds a fresh spreadsheet by hand, with a different join key and a different definition of “service”, and each time it decays within weeks of being finished. Nobody connects the five, because nobody owns the question. That is exactly the gap a platform team exists to close: if the same request keeps arriving, it should be a product, not a person.

Declared is not observed

The first instinct, when an organisation notices this, is to fix the catalogue. Make it mandatory. Run a campaign. Add a field. I understand the instinct and I have never seen it work, for a structural reason.

A catalogue is a declaration. It records what somebody chose to register, in the form they chose to register it, at the moment they did it. The estate is an observation: whatever is actually running, right now, regardless of whether anyone wrote it down. Those are different kinds of fact, and they drift apart for boring reasons. The Lambda deployed from a laptop was never registered. The service that was decommissioned last spring was never unregistered. The thing registered as one service is actually three deployables sharing a repository.

I made the argument in the developer-portal post that a catalogue nobody trusts is worse than no catalogue, and that you should not build a portal until someone will keep the ownership data true. This is the other half of that sentence: the way you keep it true is to stop relying on people to keep it true.

So build both, and treat the gap as the point:

  • Observed but not declared. Running, costing money, possibly exposed, unknown to the catalogue. This is your shadow estate, and it is where the worst surprises live.
  • Declared but not observed. In the catalogue, not running anywhere. Dead entries. Each one teaches a reader that the catalogue lies, so deleting them is worth more than it looks.
  • Both, but disagreeing. The catalogue says team A, the deploy history says team B has shipped every change for a year. The catalogue says Python 3.11, the image says 3.9. Every disagreement is a finding, and most of them are small, cheap and embarrassing.

That diff is the product. A list of what exists is useful. A list of where your organisation is wrong about what exists is what actually changes behaviour.

What the machines can tell you

Generate the observed side from systems that cannot forget and cannot be lazy. The sources are the same ones a migration reaches for in a panic, just wired in permanently:

  • The orchestrator and the cloud accounts. Every workload, function, database, bucket and queue that exists, from the provider’s API rather than from anyone’s description of it. Enumerate the accounts too; the account nobody knew about is the classic finding.
  • The registry and the build. What every image and artefact contains, from a software bill of materials generated at build time. This is the source that turns the CVE question from an investigation into a lookup.
  • Provenance stamped at creation. Which golden path, which version, which template generated this repository. Stamping that in when a road creates something is the cheapest join key you will ever add, and the one most estates are missing.
  • Identity. Which workload identities and service principals are still authenticating, and to what. A principal that authenticates every hour belongs to something alive, whatever the catalogue says.
  • The edges. DNS records, certificates, load balancer targets, network flows. The things that are reachable, which is the set an attacker cares about.
  • The bill. Every line item is an observation that something exists and somebody is paying for it.

None of this needs to be sophisticated. Each source is a scheduled job that writes rows. The work is in the join: deciding what a “deployable” is, giving it a stable identity, and matching rows from six systems to it. Provenance makes that join easy for everything the platform created. Everything else gets matched by heuristics and flagged for a human, which is fine, because that is a shrinking set if the road is doing its job.

Ownership is the hard half

Enumerating what runs is a solved engineering problem. Knowing who answers for it is not, because ownership is a social fact, and social facts decay at the speed of reorgs.

The common pattern is an owner: field with a person’s name or a team slug in a YAML file in the repository. It is accurate on the day it is written and starts rotting the following morning. People leave. Teams merge. A reorg renames four teams in an afternoon and nobody edits two hundred YAML files.

What I would do instead:

  • Owner is a team, never a person. A person is a contact. A team is an owner. If the only owner of something is an individual, it is one resignation away from being unowned, and the inventory should say so now rather than after the leaving drinks.
  • Point at a team identity that the reorg itself updates. The group in your identity provider, the team in the org system, whatever HR and IT already maintain because payroll depends on it. When the reorg happens, the source of truth changes as a side effect, and the inventory sees it the next night.
  • Corroborate with signals. Who deploys it, who approves its pull requests, whose rota gets paged, whose budget pays for it. When those agree with the declared owner, you have high confidence. When they disagree, you have a finding, and usually a real one: the team that deploys every week is the owner, whatever the file says.
  • Make “unowned” a state, not an absence. Every deployable is owned, contested, or unowned, and the last two have a queue, an age and a named arbiter who settles them. Unowned for a week is normal after a reorg. Unowned for a quarter is a decision somebody is avoiding.

That last point connects directly to operations. Asking teams to run what they build only works if “what they build” is a list both sides agree on, and the first question in any incident on an unfamiliar system is the one the platform’s own on-call has to answer under pressure: who do we call? If the inventory cannot answer that at two in the morning, it is not finished.

Keep it true by derivation, not by asking

Every inventory I have watched die was killed by the same thing: it depended on humans updating it as a separate activity. Forms. Quarterly attestations. A “please review your services” email that everyone archives.

The rule I hold to is simple: derive everything that can be derived, and ask humans only for what only humans know.

  • Derive what runs, where, on which version, built from which path, deployed by whom, costing what, reachable from where. None of that should ever be typed by a person.
  • Ask for intent: is this business critical, what data class does it handle, is it meant to exist. Those are judgements, and they cannot be observed.
  • Ask at a moment they are already there. At scaffold time, when the road creates the service. In the pull request that changes the thing. Never in a separate form on a separate day, because a separate form is a chore, and chores lose to deadlines every time.
  • Expire what nobody confirms. A judgement nobody has reaffirmed in a year is stale. Route it back to the owning team through the channel they already watch, and if it still goes unanswered, mark it unknown rather than letting an old answer pretend to be current.

This is the same move a security control makes when it ships with a paved road: stop relying on compliance and make the correct state the default one. An inventory that teams have to maintain is a rule. An inventory that maintains itself from the road is a road.

Make it a query, not a page

The temptation, once the data exists, is to build a beautiful interface for browsing it. Resist that until you know which questions people actually ask. The value of an inventory is not that you can look at it. It is that a question which used to take two weeks now takes one query.

So start from the questions. The ones I would expect to see within a month:

  • What contains package X at versions below Y, and who owns each? The CVE question.
  • What runs on substrate Z? The migration question.
  • What has no confirmed owner, and for how long? The hygiene question.
  • What has not been deployed in twelve months but is still running and costing money? The deletion question, and in most estates the cheapest progress available.
  • Who owns this IP address, bucket or hostname? The incident question.

Then measure the capability the way you would measure any product, by outcome rather than by completeness:

  • Time to answer. From “we need the list” to a list you would bet on. Weeks becoming minutes is the whole return.
  • Owned coverage. The share of observed deployables with a confirmed owning team, and the age distribution of the unowned queue.
  • Drift. The number of open declared-versus-observed disagreements, and whether it trends down.

Notice that “percentage of services registered in the catalogue” is not on the list. It is the number that looks like progress and measures the wrong side of the diff.

The cheapest version first

None of this needs a new product on day one. The first version I would build is almost embarrassingly small: a nightly job per source writing rows into one database, a join keyed on provenance where it exists, a generated report of the three diff buckets, and a handful of saved queries for the questions above. A single SQLite file and a cron job is a respectable start, in the spirit of the boring stack: pick the dull tool that is easy to keep running.

Run it for a month before you show it to anyone except the people who will fix what it finds. The first report will be humbling, and it is better to work through the embarrassing part privately. Then put the report somewhere teams see it, attach the unowned queue to the arbiter, and only after people are asking it questions every week should you think about whether it needs a front end. If it does, it can feed the catalogue or portal you already have, which finally gets the accurate content your golden paths were always supposed to index.

On staffing: this is a small, permanent slice of the run budget, not a project. It belongs with the drift treadmill and the interrupt load in the second-year accounting, and the honest pitch for it is that it converts five unplanned emergencies a year into five queries.

When this is the wrong answer

The off-ramps, because plenty of estates do not need any of this.

  • Fifteen services and one team. Everybody knows what runs and who owns it, and they are right. A markdown file in a repository is the correct inventory. Build the pipeline when the markdown file is wrong and nobody notices.
  • The estate is genuinely converged. If everything is created by the road, deployed by one pipeline into one orchestrator with ownership labels enforced at admission, the orchestrator is already your inventory. Write the saved queries and stop.
  • A CMDB already exists and is actually true. It happens. Test it before you believe it: pick ten random things from the cloud bill and see whether the CMDB knows about them. If it gets nine, feed it rather than competing with it.
  • Your industry mandates a specific system of record. Then the observed pipeline feeds the mandated system and reports its diff. Do not fork the source of truth to win an argument with an auditor.
  • You have no road yet. Without provenance, the join is mostly heuristics. Spend the effort on the first ninety days and on stamping provenance into the first path, and come back to this once the road is creating things.

One guardrail for the other direction, because inventories attract ambition. The project that sets out to model every attribute of every system ends up with a four-hundred-field schema, a governance board, and a tool that answers no question faster than the old spreadsheet did. Every field needs a question somebody actually asked in the last year. And be honest about what it is for: an inventory used to find owners and fix things gets trusted, and one used mainly to produce compliance slides and blame gets gamed, which brings you straight back to declared-but-not-true.

The bottom line

Every organisation already pays for an inventory. It just pays for it as a series of emergencies: a security fire drill, a migration discovery phase, an audit scramble, a cost sprint, an orphaned alert at two in the morning. Each produces a list, each list is thrown away, and the next request starts from zero with a different spreadsheet.

So build the boring version once. Observe the estate from the systems that cannot forget, rather than from what people remember. Keep the catalogue as the declared side and treat the difference as the product. Attach ownership to a team identity that the reorg updates for you, corroborate it with who actually deploys and gets paged, and give “unowned” a queue and an arbiter. Derive everything you can, ask humans only for intent and only at moments they are already there, and measure the thing by how long it takes to answer.

The question I would ask any platform team: if a critical CVE landed at four this afternoon, how long until you had a complete list of affected deployables, each with an owning team you could page? If the answer is a number of days, the first thing to build is not a feature. It is the list, built so you never have to build it again.

If you are staring at an estate nobody can fully enumerate, or you have rebuilt the same ownership spreadsheet three times this year, that is exactly the kind of problem I help teams work through. Let’s talk.