MLOps Pipeline Design: From Experiment to Production Deployment
The 87% of Models That Never Ship
Gartner's frequently cited statistic that 87% of ML projects never make it to production reflects a real infrastructure failure, not a modeling failure. The models work. They achieve strong offline metrics on held-out evaluation sets. But the bridge from a Jupyter notebook to a production endpoint — reproducible training, consistent feature computation, version-controlled deployment, automated monitoring — doesn't exist in most organizations.
MLOps is the engineering discipline that builds that bridge. It's not about making ML experiments faster (that's ML engineering). It's about making the path from a successful experiment to a reliable production system repeatable, automated, and reversible. The core insight is that deploying a model is deploying a data pipeline, not deploying software. Models degrade when data changes, not when code changes.
After building MLOps platforms at three different organizations — from a 10-person startup to a 500-person ML team — I've converged on a reference architecture that balances infrastructure investment against practical value. Each component earns its complexity by preventing a specific class of production failure.
Feature Stores: Solving Training-Serving Skew
Training-serving skew is the silent killer of ML systems. A model trained on features computed one way in a batch pipeline will silently produce wrong predictions when those features are computed slightly differently in a real-time serving path. The discrepancy might be as subtle as a timestamp timezone mismatch, a different null-handling convention, or a windowed aggregation computed over a slightly different time range.
A feature store solves this by providing a single definition for each feature that executes identically in both batch (training) and online (serving) contexts. The definition specifies the transformation logic, the data source, and the materialization schedule. For batch features, the store precomputes and stores historical values; for real-time features, it computes values on demand from streaming data.
from feast import Entity, FeatureView, Field
from feast.types import Float32, Int64
user = Entity(name="user_id", join_keys=["user_id"])
user_features = FeatureView(
name="user_activity_features",
entities=[user],
schema=[
Field(name="purchase_count_7d", dtype=Int64),
Field(name="avg_session_duration_30d", dtype=Float32),
Field(name="last_active_hours_ago", dtype=Float32),
],
source=bigquery_source,
online=True,
ttl=timedelta(hours=1),
)
Point-in-time correctness is the feature store's most critical capability. When constructing training data, the store retrieves the feature values as they existed at the time of each training example — not the current values. Without this, training data leaks future information into the model, inflating offline metrics and producing models that perform worse in production than expected. The feature store connects to broader data quality management practices that catch these temporal leakage issues.
Experiment Tracking: Making ML Reproducible
An ML experiment is defined by its code, data, configuration, environment, and random seeds. Change any of these and you get a different result. Experiment tracking systems record all five dimensions for every training run, making it possible to reproduce any result and compare experiments systematically.
The tracking system should capture metrics at multiple granularities: per-step training loss for debugging, per-epoch validation metrics for hyperparameter selection, and final evaluation metrics for model comparison. It should also log artifacts — the trained model weights, the data version used, the configuration file, and any preprocessing artifacts — versioned and linked to the experiment run.
Tagging conventions matter more than the tool choice. We use a three-level hierarchy: project (the business problem), experiment group (the modeling approach), and run (the specific configuration). Runs within an experiment group share the same architecture and differ only in hyperparameters. Experiment groups within a project may use entirely different approaches (a gradient boosted tree vs. a neural network for the same task). This structure is maintained through the experiment tracking platform.
Model Registry: The Version Control for Models
The model registry sits between experiment tracking and deployment. Its job is to manage model lifecycle stages — from experimental candidate to staging to production — with clear promotion criteria and rollback capabilities. Every model in the registry has a version number, a link to the experiment run that produced it, the evaluation metrics, and a set of metadata tags describing its intended use.
Promotion from staging to production requires passing validation gates. At minimum, these gates check that the model meets accuracy thresholds on the latest evaluation set, that inference latency stays within the serving budget, that the model size doesn't exceed memory limits, and that performance is stable across critical data slices (age groups, geographic regions, device types). The validation is automated — human approval is the backstop, not the primary gate.
import mlflow
client = mlflow.tracking.MlflowClient()
# Register model from experiment run
model_uri = f"runs:/{run_id}/model"
mv = client.create_model_version(
name="fraud_detector",
source=model_uri,
run_id=run_id,
tags={"data_version": "v2.4", "architecture": "xgboost"}
)
# Promote after validation
client.transition_model_version_stage(
name="fraud_detector",
version=mv.version,
stage="Production",
archive_existing_versions=True
)
Model lineage — tracking which data, features, and code produced each model — is the registry's most underappreciated capability. When a production incident occurs, lineage lets you trace from the failing prediction back to the training data that caused it. This traceability is increasingly required by AI governance frameworks and audit trail requirements.
CI/CD for ML: Beyond Software Testing
Traditional CI/CD tests code correctness. ML CI/CD must also test model quality, data quality, and behavioral correctness. The testing pyramid for ML has four layers:
- Unit tests: Verify individual transformation functions, feature engineering logic, and preprocessing steps. These run in seconds and catch coding errors.
- Integration tests: Run the full pipeline end-to-end on a small data sample. These verify that components connect correctly and produce outputs in the expected format. Typical runtime: 2-10 minutes.
- Model quality tests: Train the model on a subset of data and verify that evaluation metrics meet minimum thresholds. These catch data issues, feature bugs, and configuration errors that unit tests miss. Runtime: 15-60 minutes.
- Behavioral tests: Check specific input-output pairs that encode domain knowledge. For a sentiment classifier, these might verify that "I love this product" is positive and "terrible experience, never again" is negative. They catch subtle regressions that aggregate metrics miss.
The challenge is speed. Model quality tests that train a full model take hours. We address this with a tiered approach: fast tests run on every commit (unit + integration, under 10 minutes), quality tests run on PR merge (model training on a data sample, under 1 hour), and full evaluation runs nightly on the complete dataset. The CI/CD pipeline design must accommodate these different time horizons.
Serving Infrastructure: Deployment Patterns
Model serving is not web application serving. The computational profile is fundamentally different: ML inference involves large matrix operations on GPU or specialized hardware, with deterministic compute per request and no I/O-bound waiting. This means the optimization strategies differ from traditional web services.
We deploy models using three patterns depending on latency requirements and traffic volume. For real-time inference under 100ms, the model runs as a gRPC service behind a load balancer with request batching to improve GPU utilization. For near-real-time (100ms-10s), the model processes requests from a queue with horizontal scaling based on queue depth. For batch inference, an orchestrated pipeline processes the full dataset on a scheduled cadence and writes results to a serving table.
Deployment strategy matters as much as serving infrastructure. Canary deployments route 5% of traffic to the new model version while the remaining 95% continues to the incumbent. If the canary's online metrics (accuracy, latency, error rate) stay within bounds for 24 hours, traffic gradually shifts to 100%. Shadow deployments run the new model on all traffic in parallel without serving its predictions, collecting comparison metrics without any user impact. These patterns mirror standard canary deployment approaches adapted for ML-specific concerns.
Monitoring: Detecting Degradation Before Users Do
Models degrade silently. Unlike software bugs that produce errors, model degradation manifests as subtly wrong predictions that no monitoring system catches unless you're specifically looking for them. Three types of monitoring are essential.
Data drift monitoring compares the distribution of incoming features against the training distribution using statistical tests (Kolmogorov-Smirnov for continuous features, chi-squared for categorical features). When drift exceeds thresholds, it triggers an alert and potentially an automated retraining pipeline. The key subtlety is distinguishing benign drift (seasonal patterns, gradual trends) from harmful drift (data pipeline bugs, schema changes).
Model performance monitoring tracks prediction quality using delayed ground truth labels. For a fraud detection model, this means comparing predictions against actual fraud labels that arrive 30-90 days later. The monitoring system must handle this label delay by maintaining a prediction log and joining labels as they arrive. Performance is tracked both in aggregate and per data slice.
from evidently import Report
from evidently.metric_preset import DataDriftPreset, TargetDriftPreset
drift_report = Report(metrics=[
DataDriftPreset(stattest="ks", stattest_threshold=0.05),
TargetDriftPreset(),
])
drift_report.run(
reference_data=training_data,
current_data=production_data_last_7d,
)
drift_detected = drift_report.as_dict()["metrics"][0]["result"]["dataset_drift"]
Operational monitoring covers latency, throughput, error rates, GPU utilization, and cost per prediction. This is standard infrastructure monitoring, but with ML-specific thresholds — a 20% increase in p99 latency might indicate that the model received inputs outside its expected range, causing computation to take longer paths. The monitoring infrastructure relates to broader drift detection patterns.
Automated Retraining: Closing the Loop
The final piece of the MLOps pipeline is automated retraining. When monitoring detects data drift or performance degradation, the system automatically triggers a new training run using the latest data, validates the resulting model against quality gates, and deploys it through the standard canary process.
Three retraining triggers complement each other. Scheduled retraining runs daily or weekly regardless of drift signals — it's the baseline that catches gradual distribution shifts. Drift-triggered retraining fires when statistical tests detect significant feature distribution changes, handling sudden data shifts. Performance-triggered retraining activates when online evaluation metrics drop below thresholds, catching degradation regardless of its cause.
The critical safety mechanism is the validation gate between training and deployment. An automatically retrained model that fails quality checks must not deploy — it should alert the team for manual investigation. Without this gate, automated retraining can deploy models trained on corrupted data, creating cascading failures. The validation gates use the same metrics and thresholds as the CI/CD quality tests, ensuring consistency between manual and automated deployments.
Maturity Progression: Where to Start
Most organizations shouldn't build all six components at once. The investment should match the team's maturity and pain points. Start with experiment tracking and a model registry — these provide immediate value by making experiments reproducible and deployments reversible. Add a feature store when training-serving skew becomes a measurable problem. Build monitoring next, because you need to know when models are failing before you can automate the response. Add CI/CD automation when the deployment frequency increases enough that manual validation becomes a bottleneck. Finally, implement automated retraining once all other components are stable.
The temptation is to adopt a platform that includes all components from day one. Resist this. Platform adoption is an organizational change, not a technology change. Teams that adopt too much too fast end up with infrastructure they don't understand, can't debug, and eventually bypass with manual processes — defeating the entire purpose of MLOps. Build incrementally, prove value at each stage, and let the infrastructure grow with the team's capability.