The part after the model works
- MLOps is everything that happens once the model is good enough. Training a workable machine-learning model is a demonstration; running that model for years is a discipline. It takes the DevOps habits, automation, source control, monitoring, an unglamorous release process, and stretches them over a system whose behaviour came from a dataset as much as from code, which means the dataset now needs the same custody the code has always had.
- Reproducibility is a question you get asked afterwards. Sooner or later somebody wants to know why a particular decision came out as it did, and answering requires knowing which weights were live, which code served them and which rows trained them. Binding those three together at build time is nearly free. Reconstructing them once a complaint has landed is expensive, slow and frequently impossible.
- Serving is the easy half, and saying so out loud helps. A model behind an HTTP endpoint or inside a scheduled batch job is ordinary engineering with ordinary tools, and most teams pour their first month into it because it offers a satisfying finish line. Budget accordingly. Projects rarely come apart here, and the month spent polishing it is a month not spent on the parts that do.
- Watch what arrives, not only whether the service answers. An uptime dashboard cannot tell you the population changed. Compare the shape of today's inputs against the data the model learned from, feature by feature, and raise an alarm when they separate. That catches the upstream edit which turns a column into noise, long before anybody gets round to reading outcome figures.
- Ground truth turns up late, or never. Whether last month's churn prediction was right becomes clear next month; whether a rejected applicant would have repaid may never become clear at all. Where the truth lags, input monitoring is the only instrument you have in the interval, and where it never comes, a sampled human review deserves the seriousness of an audit rather than a spare Friday.
- Retraining belongs on a trigger, not in a calendar. A scheduled refresh renews a model that was fine and misses the week that mattered. Tie it to the monitoring signal, keep the outgoing model loaded and make reverting a routine act rather than a crisis. The same applies to systems built on large language models, often billed as LLMOps, where the versioned artefacts are prompts and evaluation sets, and the upstream change is a supplier improving a model underneath you without asking.
Write the alert before you ship
Before a model serves its first real request, write down the precise signal that would reveal it had stopped being right, the threshold that fires it and the person it reaches. Should it turn out that a customer would tell you first, or that a quarterly report would, what you are running is a prototype with excellent uptime. Most of those signals get built by data engineering, because each one is a query over a pipeline somebody has to keep honest.
