Skip to content
Getting Digital

MLOps

Also: ML operations, machine learning operations, LLMOps

MLOps is the engineering practice of keeping a trained model useful in production: recording what produced it, serving it, watching for the day its answers stop matching the world, and retraining or reverting before anyone downstream is harmed by the gap.

Assessment. Getting a model into production stopped being the hard problem years ago; realising it has gone wrong never was easy and still is not. A model whose answers have drifted out of date still replies promptly and still passes every health check you own, so spend the opening week of any budget on the signal that tells you the inputs have moved, and the week after on being able to restore the previous version without an incident review.

The part after the model works

  • MLOps is everything that happens once the model is good enough. Training a workable machine-learning model is a demonstration; running that model for years is a discipline. It takes the DevOps habits, automation, source control, monitoring, an unglamorous release process, and stretches them over a system whose behaviour came from a dataset as much as from code, which means the dataset now needs the same custody the code has always had.
  • Reproducibility is a question you get asked afterwards. Sooner or later somebody wants to know why a particular decision came out as it did, and answering requires knowing which weights were live, which code served them and which rows trained them. Binding those three together at build time is nearly free. Reconstructing them once a complaint has landed is expensive, slow and frequently impossible.
  • Serving is the easy half, and saying so out loud helps. A model behind an HTTP endpoint or inside a scheduled batch job is ordinary engineering with ordinary tools, and most teams pour their first month into it because it offers a satisfying finish line. Budget accordingly. Projects rarely come apart here, and the month spent polishing it is a month not spent on the parts that do.
  • Watch what arrives, not only whether the service answers. An uptime dashboard cannot tell you the population changed. Compare the shape of today's inputs against the data the model learned from, feature by feature, and raise an alarm when they separate. That catches the upstream edit which turns a column into noise, long before anybody gets round to reading outcome figures.
  • Ground truth turns up late, or never. Whether last month's churn prediction was right becomes clear next month; whether a rejected applicant would have repaid may never become clear at all. Where the truth lags, input monitoring is the only instrument you have in the interval, and where it never comes, a sampled human review deserves the seriousness of an audit rather than a spare Friday.
  • Retraining belongs on a trigger, not in a calendar. A scheduled refresh renews a model that was fine and misses the week that mattered. Tie it to the monitoring signal, keep the outgoing model loaded and make reverting a routine act rather than a crisis. The same applies to systems built on large language models, often billed as LLMOps, where the versioned artefacts are prompts and evaluation sets, and the upstream change is a supplier improving a model underneath you without asking.

Write the alert before you ship

Before a model serves its first real request, write down the precise signal that would reveal it had stopped being right, the threshold that fires it and the person it reaches. Should it turn out that a customer would tell you first, or that a quarterly report would, what you are running is a prototype with excellent uptime. Most of those signals get built by data engineering, because each one is a query over a pipeline somebody has to keep honest.

In practice

A marketing team scores inbound leads using a model built from past enquiries, and among its strongest features is the company-size band a visitor picks on the contact form. The website gets redesigned, the form is rebuilt, and that fixed set of bands becomes a free-text box. Nothing breaks. The endpoint stays up, response times are unchanged, a score comes back for every lead exactly as before, and the field the model leaned on now arrives as prose it cannot read, so every lead is scored as though the information were missing. Sales works a list no better than alphabetical for an entire quarter, and the conversation that finally exposes it opens in a pipeline review rather than on a pager.

  • A daily check on how often each feature arrives empty or unparseable, measured against the training baseline. This one fires on the afternoon of the deploy.
  • A week-on-week comparison of the score distribution. Losing a strong feature flattens the spread of predictions well before any outcome data reveals the damage.
  • A line in the form's release checklist noting that the field feeds a model. Cheapest control of the three, and the least likely to exist anywhere.

Often confused with

Data Engineering
Data engineering owns the pipelines that manufacture the inputs. MLOps owns the model consuming them, and it fails most often because of a change on the other side of that line.
Machine Learning
The modelling work stops once something performs well enough in a notebook. This is the discipline that assumes that day has already passed and asks what the next three years look like.

Key takeaways

  • →Deployment is solved engineering. Detection is not, and the gap between those two is where the discipline earns its name.
  • →Version the data, the code and the weights as one object, because the question of what produced a given answer always arrives later and always arrives urgently.
  • →A model can be perfectly healthy and thoroughly wrong at the same time, which is a failure mode ordinary operations tooling was never built to see.

Related concepts

  • Broader topicMachine Learning

    MLOps is the operational sub-discipline of machine learning.

  • Every production model sits downstream of data pipelines; the two disciplines share most of their plumbing.

Courses that teach this

Where this concept sits in the field

FAQ

Is this simply DevOps for models?
DevOps looks after systems whose behaviour is fully determined by code somebody wrote and somebody reviewed. Here the behaviour was inferred from data, which adds artefacts to version, datasets, weights and feature definitions, and introduces one failure mode ordinary operations never had to handle: a service that is entirely healthy and silently incorrect.
Does a single model need a platform?
No, you need the habits. Training data under version control, a pipeline that runs end to end without anyone's laptop, a check on incoming features, and a documented route back to the previous model. Platforms start paying for themselves when several models and several teams begin competing for the same pipelines.
What is drift?
Two different problems wearing one name. Either the world moved so the inputs no longer resemble the training data, or the inputs look familiar while the relationship between them and the outcome has changed underneath. The first is detectable without waiting for outcomes. The second usually is not, and the second is the expensive one.
Who should own it?
In a small team, whoever trained the model, which is a decent reason to keep the model simple. In a larger one it belongs with the engineers who own the serving path, with the data scientist on the alert rota, because the person able to interpret a drift signal is rarely the person carrying the pager at midnight.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 26 September 2026 · Getting Digital