Skip to main content

The Most Important Engineer You Haven't Hired Yet

Arlo Gilbert · September 30, 2026

The Most Important Engineer You Haven't Hired Yet

For most of my career, at most of my companies, software ran on a calendar.

Every quarter we reviewed the platform. We looked at libraries that had fallen behind, dependencies with security advisories, and services that needed attention before they turned into a liability. Once a year the operating system vendors shipped a major release, and every engineering team spent a few weeks making sure nothing broke. Microsoft made the rhythm official in 2003 with Patch Tuesday. The rest of the industry settled into roughly the same heartbeat.

Change arrived on a schedule, so we staffed around the schedule. Engineers owned features and ops owned uptime. Security owned patches, and somebody owned the upgrade plan. Every moving part in the system had a name next to it.

What changed

The server side of most software has gotten pretty stable. Cloud infrastructure is mature, frameworks are mature, and coding agents now handle a growing share of routine upkeep. The code that holds a product together changes less often than it did five years ago.

The most important component in a lot of products now changes constantly. The frontier labs ship a major model every few months. New snapshots, price changes, and deprecation notices land every few weeks. Open weight models are improving so quickly that the quality you paid flagship prices for last year now costs a fraction of that.

Each of those events changes how your product behaves. Some make it better, some change your costs, and a few quietly break things. None of them show up in a pull request, and none of them wait for your quarterly review.

Who owns the models?

If you ask most engineering leaders who is responsible for the models their product runs on, you'll get a pause. The engineer who built the integration picked a model a year and a half ago and moved on to other work. The platform team treats the model API like any other vendor dependency. Finance sees the invoice. Everyone assumes someone else is watching.

That assumption made sense when a dependency was a payments API or a database. Those are versioned and documented, and they behave next month the same way they behave today. A model is closer to a contractor on your team. Their skills improve, their rates change, and every so often they come back from a long weekend doing the work a little differently. Nobody would leave a contractor unmanaged for eighteen months.

The job

I think this deserves a named role. Call it model stewardship, or model ops. Pick whatever title fits your org. What matters is that one person starts every week responsible for how the models inside your product are behaving and what they cost.

At a small company this is part of a senior engineer's job. At scale it becomes a small team. You want someone with an engineer's instincts and an analyst's patience. They can write an eval, read a model card, and hold their own in a budget conversation with the CFO.

What the work covers

Evals. Everything else depends on this. An eval suite is a set of real tasks from your product with known good answers or clear grading rubrics. You should be able to run it against any model in an afternoon. The best ones come from production traffic. (If you're reading this, you already know to scrub the personal data out first.) When a new model ships, your evals tell you within hours whether it's better at your work. Public benchmarks tell you how it does on someone else's.


Cost per task. Token price is the sticker on the shelf. The number that matters is what you pay to finish a task correctly. That includes input and output tokens, reasoning tokens, retries, tool calls, and the human time spent cleaning up bad answers. Reasoning models make this math harder, since a model can think through thousands of tokens before it writes a word. A cheap model that needs three attempts can easily cost more than an expensive one that gets it right the first time. The steward tracks cost per task for every feature, every month.

Drift. Model behavior moves even when your code stays put. Aliases like "latest" get pointed at new snapshots. Providers change their serving infrastructure. In September 2025, Anthropic published a postmortem describing three separate infrastructure bugs that had degraded some Claude responses for weeks. I give them real credit for publishing it. It also proves the point. If a lab that careful can ship a regression, your product will eventually inherit one. The steward runs evals on a schedule against the models in production and samples live outputs. They watch the boring signals too: output length, refusal rates, format failures, and tool call errors. They also own the deprecation calendar, because every provider retires old models on its own timeline.

Routing. Once you know what each model is good at and what it costs per task, you can send every request to the right one. Classification and extraction go to a small, inexpensive model. Hard reasoning goes to the frontier. I wrote about the gateway plumbing for this back in June, and the plumbing is the easy part. The routing rules go stale every time a new model ships, so someone has to keep them current.

How this goes wrong

I see two failures.

The first is the silent regression. Nothing throws an error. The JSON still parses and the latency looks fine. The answers just get slightly worse... or even worse, they stay the same but your competition is using better models. A summarization feature starts dropping the one clause that mattered. A classifier slips a few points on the categories your customers care about most. In a compliance workflow, that kind of slippage can run for a month before a customer finds the missed clause. By then you have a trust problem on top of a quality problem.

The second is overpaying for the flagship. The default move is to wire every feature to the best model available and never look at it again. It feels safe. A year later most of that traffic is extraction and formatting that a far cheaper model handles just as well. Data has shown that, roughly 70 percent of requests never need a frontier model. Nobody has the numbers because running the numbers isn't anyone's job.

In both cases, nobody owns the models, so nobody notices.

Where it reports

This is the hard part, because model decisions are financial, product, and engineering decisions all at once. Every choice about which model runs a feature changes what it costs, how well it works for customers, and what engineering has to maintain.

Some companies will put this role under FinOps, the cloud cost discipline that usually lives on the CFO's team. I understand the instinct. The invoice is the most visible symptom, and FinOps teams are already good at tagging spend and finding waste.

I think that's the wrong home. A team measured on spend will reliably fix the second failure mode above and quietly create the first. Cheaper models look great on a cost dashboard right up until a missed clause turns into a customer escalation. Cost per task only means something when the same person owns the definition of a task done correctly.

My inclination is a direct report to the CTO, with a standing seat in product reviews and a close working relationship with finance. Finance should see the numbers every month and help set the budget. The call on which model runs which feature belongs to someone who answers for how that feature behaves. The exact reporting line will depend on org size, and I expect it to move around for a few years as companies figure this out.

Two clocks

The software development lifecycle has split into two tracks, and they run at very different speeds.

Code maintenance has slowed down. The platform is stable, the frameworks are settled, and agents absorb a lot of the routine upkeep. The quarterly review and the annual upgrade plan still fit that work well.

Model maintenance has sped up to something close to a weekly cadence. New releases, new prices, new failure modes, and new chances to spend less for better results. A quarterly review on this track means every decision you make is already out of date.

Most engineering orgs are still staffed for the first scenario.

Google ran into a version of this in 2003. Ben Treynor Sloss was asked to run a production team, and he built what became site reliability engineering. It took years for the rest of the industry to copy the idea and put the title on org charts. I expect model stewardship to follow the same path, faster.

If you lead an engineering team, you don't have to wait for the title to catch on. Pick the person now. Give them an eval suite, a cost dashboard, and the authority to change which model a feature uses. The models underneath your product are going to keep getting better, and someone on your team should be watching when they do.