MLOps in Practice: Reproducible Pipelines, Safe Releases, and Drift Monitoring That Works
A practical MLOps guide: versioning, training-serving skew, release patterns, drift metrics with a worked PSI example, LLM evaluation, and 2026 rules.

Training a model that scores well on a test set is the easy part. Keeping it accurate in production, where data arrives late or malformed, customers change behavior, and the code that computes features drifts away from the code used in training, is where most of the work sits. Google engineers made the point in their 2015 paper Hidden Technical Debt in Machine Learning Systems: the model code is a small part of a production ML system, surrounded by much larger amounts of data collection, feature extraction, serving, and monitoring infrastructure.
MLOps is the discipline of building that surrounding system. This guide covers what to version, how to avoid training-serving skew, a sensible order for automation, release patterns, drift monitoring with a worked example, what changes for LLM applications, and the rules that apply in 2026.
A failure to learn from
Zillow shut down its home-buying business in November 2021. In its third-quarter shareholder letter, the company said it had been unable to forecast home prices accurately, with errors in both directions larger than it had modeled as possible. It recorded a $304 million inventory write-down after buying homes above its current estimates of their resale value, and it announced cuts of about 25% of its workforce. The stock fell 25% the next day.
The letter also noted that conversion rates rose above what Zillow had seen before: more sellers accepted its offers, which in hindsight signaled that it was paying too much. The lessons for MLOps are practical. Monitor business outcomes as well as model metrics, treat a sudden improvement in a downstream metric as something to investigate, and cap the exposure a model can create before its predictions can be checked against reality. The letter also cited renovation and resale capacity constraints, so the failure was not only a modeling problem.
What to version, and how to keep it reproducible

To reproduce a model, or explain one to an auditor, you need to know exactly what produced it:
- Code for data preparation, feature computation, training, and serving, in version control.
- The training data, as a snapshot or a query against immutable, dated data, with its schema.
- Configuration and hyperparameters, the random seeds, and the library and container versions.
- The trained model artifact, its evaluation results on fixed test sets, and who approved it.
A model registry ties these together and records which version is in production. MLflow is a widely used open-source option; its 3.0 release in June 2025 added tracking for generative AI applications. Cloud platforms offer managed registries. The ecosystem is consolidating: CoreWeave completed its acquisition of Weights & Biases in May 2025, and Databricks acquired the assets of Tecton, a feature store vendor, later that year. Favor tools that export in open formats so a vendor change does not strand your history.
Training-serving skew
The most common silent failure in production ML is a model receiving features in production that differ from the ones it was trained on. Typical causes:
- Two code paths: features computed in a batch SQL job for training and reimplemented in a service for real-time predictions.
- Leakage: training data that includes information not available at prediction time, such as a field updated after the event being predicted.
- Point-in-time errors: joining the latest value of a slowly changing attribute instead of the value as of the prediction date.
- Defaults and missing values handled differently in training and serving.
Google's Rules of Machine Learning recommend logging the features used at serving time and training on those logs, and reusing the same code in both pipelines. A feature store does this at scale by computing each feature once, storing its history for point-in-time-correct training sets, and serving the same values online; Feast is the main open-source option. A feature store is worth its operating cost once several models share features or need real-time values. For a single batch model, a well-tested shared library and logged features usually suffice.
Notebooks are fine for exploration, but code that runs in production should be packaged as tested modules, with unit tests on feature transformations and a check that the served model reproduces its offline predictions on a sample of logged requests.
Automate in a sensible order
Google Cloud's MLOps maturity levels describe a progression from manual training and deployment (level 0), to an automated training pipeline (level 1), to automated building, testing, and deployment of the pipeline itself (level 2). Most teams should move through them in that order and stop where the benefits stop justifying the work:
- Data validation first. Check schemas, value ranges, null rates, and row counts on every batch before it reaches training or scoring. Catching bad input data here is cheaper than diagnosing a degraded model later.
- A reproducible training pipeline that runs from versioned code and data to a registered model with evaluation results.
- Automated gates before promotion: the new model must beat the current one on fixed test sets and on important segments, not only on average, and pass fairness or policy checks where they apply.
- Automated deployment with rollback.
- Scheduled or triggered retraining, once monitoring shows how quickly performance decays. Retraining on a fixed schedule without monitoring can train a model on corrupted data.
Google's ML Test Score paper offers a checklist of tests for data, models, infrastructure, and monitoring that teams can score themselves against.
Releasing new models
- Shadow deployment runs the candidate on live traffic without acting on its outputs. It catches skew, latency, and errors before any customer is affected, and it is especially useful when outcomes arrive slowly.
- A canary release sends a small share of traffic to the candidate and widens it while guardrail metrics (errors, latency, prediction distribution, business outcomes) hold.
- An A/B test measures the business effect with a proper experiment design and enough traffic to detect the difference that matters.
- Rollback should take minutes and be tested. Keep the previous model version and its feature pipeline deployable until the new one has run through at least one full business cycle.
Monitoring drift: a worked example

Monitor three layers: the data going in, the predictions coming out, and the outcomes when they arrive. Outcome labels often lag. Fraud chargebacks and loan defaults can take weeks to months, so input and prediction drift serve as early warnings.
The population stability index is a simple, widely used measure of input drift. An illustration: a feature is split into five bins that each held 20% of the training data. Last week's production data shows 10%, 15%, 20%, 25%, and 30%.
| Bin | Training share | Production share | (Production - training) x ln(production / training) |
|---|---|---|---|
| 1 | 20% | 10% | 0.0693 |
| 2 | 20% | 15% | 0.0144 |
| 3 | 20% | 20% | 0.0000 |
| 4 | 20% | 25% | 0.0112 |
| 5 | 20% | 30% | 0.0405 |
| PSI | 0.135 |
A common rule of thumb from credit scoring reads under 0.1 as stable, 0.1 to 0.25 as a moderate shift to investigate, and above 0.25 as a major shift. At 0.135, this feature warrants a look: is the change real (a new customer segment, a marketing campaign), or a data problem (a unit change, a broken join)? If it is real, check whether the model performs worse on the growing segment.
Two cautions about drift tests:
- Statistical significance is not business relevance. With a million observations in each sample, a Kolmogorov-Smirnov test flags a shift of one-hundredth of a standard deviation, because the critical value falls to about 0.002. Alert on effect sizes, such as PSI or changes in segment performance, rather than p-values.
- Drift alerts need an owner and a playbook: who investigates, what evidence triggers retraining or rollback, and how quickly.
When labels arrive, track performance by segment and over time against the metrics that matter to the business, such as losses avoided or approval rates, alongside accuracy measures. Our guides to AI fraud detection and AI financial forecasting cover domain-specific metrics.
What changes for LLM applications
Applications built on large language models need the same discipline with different artifacts:
- Version prompts, retrieval configurations, model identifiers, and tool definitions as code, since any of them can change behavior.
- Keep evaluation sets of real questions with expected answers or grading rubrics, and run them on every change and whenever the provider updates a model.
- Trace each request through retrieval, model calls, and tool calls, with latency and cost per request.
- Monitor for failure modes specific to LLMs: answers not supported by the retrieved sources, prompt injection through retrieved content, and leaks of sensitive data. Our guide to generative AI data governance covers the data side.
Governance and the 2026 rules
- U.S. banks apply the Federal Reserve's SR 11-7 model risk guidance, which expects independent validation, documentation, and ongoing monitoring of models, including machine learning models.
- The EU AI Act requires high-risk AI systems to log events automatically over their lifetime and requires providers to run post-market monitoring. After the Digital Omnibus amendments, obligations for most high-risk uses listed in Annex III apply from December 2, 2027.
- The NIST AI Risk Management Framework is voluntary in the U.S. but widely used to organize controls.
Good MLOps practice produces most of what these frameworks ask for: lineage, evaluation records, monitoring logs, and change approvals. Our guide to AI governance frameworks covers the policy side, and cloud cost governance covers controlling training and inference spend.
A starting checklist
- Put training code, feature code, and serving code in version control, and snapshot training data with its schema.
- Register every production model with its data version, evaluation results, and approver.
- Log the features used at serving time and compare them with training features on a sample.
- Validate every incoming batch before training or scoring.
- Release new models in shadow first, then as a canary, with a tested rollback.
- Monitor inputs, predictions, and outcomes by segment, with owners and playbooks for alerts.
- Retrain on evidence from monitoring, not only on a calendar.
This guide is for informational purposes only. Tools, regulations, and deadlines change; regulatory dates are as of September 2026. Organizations subject to model risk or AI regulations should confirm their obligations with qualified legal and compliance advisers.



