#mlops#machine learning#ai#devops#model monitoring

MLOps in Practice: Reproducible Pipelines, Safe Releases, and Drift Monitoring That Works

A practical MLOps guide: versioning, training-serving skew, release patterns, drift metrics with a worked PSI example, LLM evaluation, and 2026 rules.

📅 January 9, 2026✏️ Updated: September 27, 2026⏱ 9 min read✍ Web3 Listicle Editorial Team

An engineering team reviewing an automated machine learning pipeline and model monitoring metrics.

Training a model that scores well on a test set is the easy part. Keeping it accurate in production, where data arrives late or malformed, customers change behavior, and the code that computes features drifts away from the code used in training, is where most of the work sits. Google engineers made the point in their 2015 paper Hidden Technical Debt in Machine Learning Systems: the model code is a small part of a production ML system, surrounded by much larger amounts of data collection, feature extraction, serving, and monitoring infrastructure.

MLOps is the discipline of building that surrounding system. This guide covers what to version, how to avoid training-serving skew, a sensible order for automation, release patterns, drift monitoring with a worked example, what changes for LLM applications, and the rules that apply in 2026.

A failure to learn from

Zillow shut down its home-buying business in November 2021. In its third-quarter shareholder letter, the company said it had been unable to forecast home prices accurately, with errors in both directions larger than it had modeled as possible. It recorded a $304 million inventory write-down after buying homes above its current estimates of their resale value, and it announced cuts of about 25% of its workforce. The stock fell 25% the next day.

The letter also noted that conversion rates rose above what Zillow had seen before: more sellers accepted its offers, which in hindsight signaled that it was paying too much. The lessons for MLOps are practical. Monitor business outcomes as well as model metrics, treat a sudden improvement in a downstream metric as something to investigate, and cap the exposure a model can create before its predictions can be checked against reality. The letter also cited renovation and resale capacity constraints, so the failure was not only a modeling problem.

What to version, and how to keep it reproducible

A diagram of data flowing from ingestion through feature pipelines to training and a model registry.

To reproduce a model, or explain one to an auditor, you need to know exactly what produced it:

  • Code for data preparation, feature computation, training, and serving, in version control.
  • The training data, as a snapshot or a query against immutable, dated data, with its schema.
  • Configuration and hyperparameters, the random seeds, and the library and container versions.
  • The trained model artifact, its evaluation results on fixed test sets, and who approved it.

A model registry ties these together and records which version is in production. MLflow is a widely used open-source option; its 3.0 release in June 2025 added tracking for generative AI applications. Cloud platforms offer managed registries. The ecosystem is consolidating: CoreWeave completed its acquisition of Weights & Biases in May 2025, and Databricks acquired the assets of Tecton, a feature store vendor, later that year. Favor tools that export in open formats so a vendor change does not strand your history.

Training-serving skew

The most common silent failure in production ML is a model receiving features in production that differ from the ones it was trained on. Typical causes:

  • Two code paths: features computed in a batch SQL job for training and reimplemented in a service for real-time predictions.
  • Leakage: training data that includes information not available at prediction time, such as a field updated after the event being predicted.
  • Point-in-time errors: joining the latest value of a slowly changing attribute instead of the value as of the prediction date.
  • Defaults and missing values handled differently in training and serving.

Google's Rules of Machine Learning recommend logging the features used at serving time and training on those logs, and reusing the same code in both pipelines. A feature store does this at scale by computing each feature once, storing its history for point-in-time-correct training sets, and serving the same values online; Feast is the main open-source option. A feature store is worth its operating cost once several models share features or need real-time values. For a single batch model, a well-tested shared library and logged features usually suffice.

Notebooks are fine for exploration, but code that runs in production should be packaged as tested modules, with unit tests on feature transformations and a check that the served model reproduces its offline predictions on a sample of logged requests.

Automate in a sensible order

Google Cloud's MLOps maturity levels describe a progression from manual training and deployment (level 0), to an automated training pipeline (level 1), to automated building, testing, and deployment of the pipeline itself (level 2). Most teams should move through them in that order and stop where the benefits stop justifying the work:

  1. Data validation first. Check schemas, value ranges, null rates, and row counts on every batch before it reaches training or scoring. Catching bad input data here is cheaper than diagnosing a degraded model later.
  2. A reproducible training pipeline that runs from versioned code and data to a registered model with evaluation results.
  3. Automated gates before promotion: the new model must beat the current one on fixed test sets and on important segments, not only on average, and pass fairness or policy checks where they apply.
  4. Automated deployment with rollback.
  5. Scheduled or triggered retraining, once monitoring shows how quickly performance decays. Retraining on a fixed schedule without monitoring can train a model on corrupted data.

Google's ML Test Score paper offers a checklist of tests for data, models, infrastructure, and monitoring that teams can score themselves against.

Releasing new models

  • Shadow deployment runs the candidate on live traffic without acting on its outputs. It catches skew, latency, and errors before any customer is affected, and it is especially useful when outcomes arrive slowly.
  • A canary release sends a small share of traffic to the candidate and widens it while guardrail metrics (errors, latency, prediction distribution, business outcomes) hold.
  • An A/B test measures the business effect with a proper experiment design and enough traffic to detect the difference that matters.
  • Rollback should take minutes and be tested. Keep the previous model version and its feature pipeline deployable until the new one has run through at least one full business cycle.

Monitoring drift: a worked example

A monitoring dashboard showing model latency, error rates, and drift indicators.

Monitor three layers: the data going in, the predictions coming out, and the outcomes when they arrive. Outcome labels often lag. Fraud chargebacks and loan defaults can take weeks to months, so input and prediction drift serve as early warnings.

The population stability index is a simple, widely used measure of input drift. An illustration: a feature is split into five bins that each held 20% of the training data. Last week's production data shows 10%, 15%, 20%, 25%, and 30%.

Bin Training share Production share (Production - training) x ln(production / training)
1 20% 10% 0.0693
2 20% 15% 0.0144
3 20% 20% 0.0000
4 20% 25% 0.0112
5 20% 30% 0.0405
PSI 0.135

A common rule of thumb from credit scoring reads under 0.1 as stable, 0.1 to 0.25 as a moderate shift to investigate, and above 0.25 as a major shift. At 0.135, this feature warrants a look: is the change real (a new customer segment, a marketing campaign), or a data problem (a unit change, a broken join)? If it is real, check whether the model performs worse on the growing segment.

Two cautions about drift tests:

  • Statistical significance is not business relevance. With a million observations in each sample, a Kolmogorov-Smirnov test flags a shift of one-hundredth of a standard deviation, because the critical value falls to about 0.002. Alert on effect sizes, such as PSI or changes in segment performance, rather than p-values.
  • Drift alerts need an owner and a playbook: who investigates, what evidence triggers retraining or rollback, and how quickly.

When labels arrive, track performance by segment and over time against the metrics that matter to the business, such as losses avoided or approval rates, alongside accuracy measures. Our guides to AI fraud detection and AI financial forecasting cover domain-specific metrics.

What changes for LLM applications

Applications built on large language models need the same discipline with different artifacts:

  • Version prompts, retrieval configurations, model identifiers, and tool definitions as code, since any of them can change behavior.
  • Keep evaluation sets of real questions with expected answers or grading rubrics, and run them on every change and whenever the provider updates a model.
  • Trace each request through retrieval, model calls, and tool calls, with latency and cost per request.
  • Monitor for failure modes specific to LLMs: answers not supported by the retrieved sources, prompt injection through retrieved content, and leaks of sensitive data. Our guide to generative AI data governance covers the data side.

Governance and the 2026 rules

  • U.S. banks apply the Federal Reserve's SR 11-7 model risk guidance, which expects independent validation, documentation, and ongoing monitoring of models, including machine learning models.
  • The EU AI Act requires high-risk AI systems to log events automatically over their lifetime and requires providers to run post-market monitoring. After the Digital Omnibus amendments, obligations for most high-risk uses listed in Annex III apply from December 2, 2027.
  • The NIST AI Risk Management Framework is voluntary in the U.S. but widely used to organize controls.

Good MLOps practice produces most of what these frameworks ask for: lineage, evaluation records, monitoring logs, and change approvals. Our guide to AI governance frameworks covers the policy side, and cloud cost governance covers controlling training and inference spend.

A starting checklist

  1. Put training code, feature code, and serving code in version control, and snapshot training data with its schema.
  2. Register every production model with its data version, evaluation results, and approver.
  3. Log the features used at serving time and compare them with training features on a sample.
  4. Validate every incoming batch before training or scoring.
  5. Release new models in shadow first, then as a canary, with a tested rollback.
  6. Monitor inputs, predictions, and outcomes by segment, with owners and playbooks for alerts.
  7. Retrain on evidence from monitoring, not only on a calendar.

This guide is for informational purposes only. Tools, regulations, and deadlines change; regulatory dates are as of September 2026. Organizations subject to model risk or AI regulations should confirm their obligations with qualified legal and compliance advisers.

Frequently Asked Questions

MLOps is the set of engineering practices for getting machine learning models into production and keeping them working there: versioning code, data, and models together, automating training and deployment, releasing new models safely, monitoring inputs and outcomes, and retraining or rolling back when performance slips. It borrows from DevOps but adds the problem that a model's behavior depends on data that keeps changing.
It is a difference between how a model performs in training and how it performs in production, usually because features are computed differently in the two places, for example by separate code paths or from data that was not available at prediction time. Google's Rules of Machine Learning recommend logging the features actually used at serving time and training on those logs, and reusing the same code for training and serving.
Data drift is a change in the distribution of a model's inputs, such as a new customer mix. Concept drift is a change in the relationship between inputs and the outcome, such as fraudsters changing tactics so the same signals mean something different. Data drift can be measured immediately from inputs; concept drift usually shows up only when outcome labels arrive, which can take weeks or months.
The population stability index (PSI) compares how a variable is distributed across bins in production versus a reference period. It sums, over the bins, the difference in shares multiplied by the log of their ratio. A common rule of thumb from credit scoring treats values under 0.1 as stable, 0.1 to 0.25 as a moderate shift worth investigating, and above 0.25 as a major shift. In our example, a feature whose production mix moved from 20% in each of five bins to 10%, 15%, 20%, 25%, and 30% has a PSI of about 0.135.
Shadow deployment runs the new model on live traffic without acting on its predictions, so its outputs can be compared with the current model's. A canary release sends a small share of traffic to the new model and expands it if metrics hold. An A/B test splits users to measure the business effect. All three need a fast, tested way to roll back to the previous model version.
Banks in the U.S. follow the Federal Reserve's SR 11-7 model risk guidance, which expects validation and ongoing monitoring. The EU AI Act requires high-risk AI systems to keep automatic logs and providers to run post-market monitoring; after the Digital Omnibus amendments, the obligations for most high-risk uses apply from December 2, 2027. The NIST AI Risk Management Framework is a voluntary U.S. reference that many organizations map their controls to.

Share this article