Ask a Founder what they are hiring and you usually get the title first. “We need a machine learning engineer.” Sometimes two.
Ask what that person will actually do in their first six months and the answer splits into four different jobs.
One trains models. One makes an existing model work against a product problem. One keeps the training and serving infrastructure standing. One builds the product around a model somebody else trained.
Same title on the job board. Almost no overlap in the hiring bar.
The Title Stopped Being Descriptive Around 2023
For most of the last decade, Machine Learning Engineer meant roughly one thing: someone who could take a modelling problem, assemble a dataset, train something, and get it into production. Narrow enough that the title carried information.
Foundation models broke that. When a capable general model is a call away, the scarce work moves. Fewer companies train from scratch. Far more companies need people who can evaluate, adapt, serve and productise models they did not build.
The title did not move with the work. So it now covers four jobs at once, and a brief written around the title inherits all four.
Four Profiles Behind One Title
Treat these as distinct searches, because that is what they are.
- Research and research engineering. Advances model capability itself: architectures, training regimes, pre-training and post-training, evaluation methodology. Publication record and lab pedigree matter here because the work is genuinely closer to science than to software. This is the scarcest of the four by a wide margin, and the smallest share of open roles.
- Applied ML engineering. Takes capability that already exists and makes it work on your problem: fine-tuning, retrieval, prompting and orchestration, and above all evaluation. The defining skill is building an eval harness that tells you whether a change actually helped. Most AI product companies think they are hiring research and need this.
- ML platform and infrastructure engineering. Training and serving infrastructure, GPU scheduling and utilisation, data pipelines, inference latency and cost. Closer to distributed systems than to modelling. At infrastructure companies this is often the first ML hire, and the one that unblocks everyone else.
- ML-informed product engineering. Builds the product surface around model behaviour: the interface, the fallbacks, the handling of a system that is wrong some percentage of the time. Strong product engineers who understand model failure modes. Usually the highest-volume genuine need, and almost never the title on the brief.
A candidate can be exceptional at one of these and unhireable for another. That is not a gap in their profile. It is four professions sharing a name.
What Getting It Wrong Actually Costs
The failure is rarely a bad hire. It is a search that never converges.
A company writes a research-flavoured brief because research sounds like the ambitious version of the role. The interview loop follows the brief, so it tests derivations and training internals. Strong applied engineers fail a loop that has nothing to do with the job, and the few researchers who pass turn down an offer to build retrieval pipelines, because that is not the career they are in.
Three months later the conclusion is that the talent market is impossible.
The market was not the problem. The company ran a search for a profile it did not need, using a loop that screened out the profile it did.
What to Test For, by Profile
The interview loop should look different for each of the four. If your loop is identical regardless of which profile you are hiring, it is calibrated for none of them.
- Research. Depth on a problem they own. Ask them to walk through something that did not work and what it changed in their thinking. Reasoning quality under pushback beats breadth of coverage.
- Applied. Give them a real, ambiguous product problem and ask how they would know whether a change improved it. Listen for evaluation design before model choice. Anyone who reaches for a bigger model before an eval set is answering a different question.
- Platform and infrastructure. Systems design under cost and latency constraints. Ask about a time they cut inference cost or training time and what the tradeoff was. Vagueness about numbers here is diagnostic.
- Product engineering. Ask how they would design a feature that is confidently wrong two times in ten. The good answers are about surfacing uncertainty and building recoverable paths, not about accuracy.
Why the Market Feels Tighter Than It Is
Genuine scarcity is concentrated in one profile. Pre-training researchers with real lab experience are few and are being competed for by organisations with unmatched budgets. If that is your need, expect a hard search and price it honestly.
The other three are far more available than the discourse suggests. Applied engineers, ML infrastructure engineers and product engineers who understand model behaviour exist in reasonable numbers, and a good share of them are currently doing a version of the job under a different title.
So most perceived scarcity is self-inflicted. It comes from writing a research brief for an applied job, then searching a population of a few thousand people when the right population was a hundred times larger.
Compensation Follows the Profile, Not the Title
Benchmarking by title produces an average across four jobs, which describes none of them. Research commands a premium that applied and product engineering do not, and platform engineering prices closer to senior infrastructure work than to modelling.
Pay the profile you are hiring. Benchmarking to the highest of the four inflates the offer and attracts the wrong candidates. Benchmarking to the lowest loses the search quietly, at the offer stage, months in.
How to Write the Brief
Before opening the search, answer four questions honestly.
- What does this person ship in the first ninety days? If the answer involves a model that does not exist yet, you are hiring research. If it involves a feature, you are not.
- Who decides whether the model is good enough? If nobody currently owns evaluation, that is the job, and it is applied.
- What is standing between the model and production today? If the honest answer is serving cost, latency or data plumbing, hire platform first.
- Would the strongest candidate for this role want it in two years? Profile and ambition have to match, or you win the search and lose the person.
Most briefs get sharper under those four questions, and some resolve into a different role entirely.
The Broader Principle
Job titles are inherited. Job definitions have to be written.
The companies that hire machine learning talent well are not the ones with the biggest budgets or the best logos. They are the ones who worked out which of the four jobs they were actually filling before anyone saw a CV.
It is not a shortage of machine learning engineers. It is a shortage of briefs precise enough to find them.
If you are scoping an ML hire and want a second view on which profile the role really is, I would be happy to talk it through. You can see how we approach this on our page on hiring machine learning engineers for AI startups, and more on sequencing in Sequencing GTM Hires at Seed and Series A.
Vector is a specialist recruiting agency helping VC-backed AI and infrastructure startups build their GTM, product, and engineering teams.