TempatGunting

MLOps Model Monitoring and Drift Detection in Production

A machine learning model that performs well at deployment will degrade over time. This is not speculation — it is a certainty that follows from the nature of statistical models. They learn relationships in training data that reflect a specific moment in time. As the world changes, those relationships weaken. The question is not whether a model will degrade, but how quickly you will detect it and how effectively you will respond.

Model monitoring in production goes beyond tracking prediction latency and error rates. It requires statistical measurement of input data distributions, output distribution stability, and model performance against ground truth when available. The monitoring system must detect three distinct failure modes: data drift (input distribution changes), concept drift (the relationship between inputs and outputs changes), and model degradation (the model's internal representations become less effective even on familiar inputs).

Data Drift Detection

Data drift measures how the distribution of production input features diverges from the training data distribution. When features shift, predictions become unreliable because the model is extrapolating beyond its training regime. Effective drift detection requires choosing the right statistical test, setting meaningful thresholds, and operating at the right granularity.

import numpy as np
from scipy import stats
from evidently.metrics import DataDriftTable
from evidently.report import Report

def compute_drift_metrics(reference_data, production_data, features):
    results = {}
    for feature in features:
        ref = reference_data[feature].dropna()
        prod = production_data[feature].dropna()

        # KS test for continuous features
        ks_stat, ks_pvalue = stats.ks_2samp(ref, prod)

        # Population Stability Index
        psi = calculate_psi(ref, prod, buckets=10)

        # Jensen-Shannon divergence
        js_div = calculate_js_divergence(ref, prod)

        results[feature] = {
            "ks_statistic": ks_stat,
            "ks_pvalue": ks_pvalue,
            "psi": psi,
            "js_divergence": js_div,
            "drifted": psi > 0.2 or ks_pvalue < 0.01
        }
    return results

def calculate_psi(reference, production, buckets=10):
    breakpoints = np.percentile(reference, np.linspace(0, 100, buckets + 1))
    breakpoints[-1] = np.inf
    breakpoints[0] = -np.inf

    ref_counts = np.histogram(reference, breakpoints)[0] / len(reference)
    prod_counts = np.histogram(production, breakpoints)[0] / len(production)

    # Avoid division by zero
    ref_counts = np.clip(ref_counts, 1e-6, None)
    prod_counts = np.clip(prod_counts, 1e-6, None)

    psi = np.sum((prod_counts - ref_counts) * np.log(prod_counts / ref_counts))
    return psi

PSI thresholds have industry-standard interpretations: below 0.1 indicates no significant drift, 0.1-0.2 suggests moderate shift worth monitoring, and above 0.2 signals significant drift requiring investigation. These thresholds originated in credit scoring but generalize well to other domains. Adjust them based on your model's sensitivity — some models tolerate substantial input drift while others degrade rapidly. The techniques used in RAG architecture monitoring share similar principles around detecting when input distributions shift.

Concept Drift Detection

Concept drift is harder to detect because it requires ground truth labels. The input distribution may look identical to training data, but the correct predictions have changed. A fraud detection model trained during normal economic conditions may see the same transaction patterns but with different fraud labels during a recession. The model's accuracy drops not because the inputs changed, but because what those inputs mean has changed.

from nannyml import CBPE, PerformanceEstimator
from nannyml.drift import UnivariateDriftCalculator

# Confidence-Based Performance Estimation (when labels are delayed)
estimator = CBPE(
    y_pred_proba="predicted_probability",
    y_pred="prediction",
    y_true="actual_label",            # used for reference period only
    problem_type="classification_binary",
    metrics=["roc_auc", "f1"],
    chunk_size=5000
)

# Fit on reference period (where labels are available)
estimator.fit(reference_data)

# Estimate performance on production data (no labels needed)
results = estimator.estimate(production_data)

# Alert when estimated performance drops below threshold
for chunk in results.data:
    if chunk["roc_auc"] < 0.75:
        trigger_alert(
            metric="roc_auc",
            estimated_value=chunk["roc_auc"],
            threshold=0.75,
            period=chunk["period"]
        )

NannyML's CBPE method estimates model performance without ground truth by analyzing the relationship between prediction confidence and actual accuracy established during the reference period. When this calibration relationship breaks down, it indicates concept drift. This approach provides earlier detection than waiting for delayed labels, though it should be validated against actual performance once labels arrive.

Production Monitoring Architecture

A production monitoring system needs four components: a data collector that captures model inputs and outputs, a drift calculator that computes statistical metrics on batched data, an alerting system that notifies stakeholders when thresholds are breached, and a dashboard that provides visibility into model health over time.

# Evidently AI monitoring pipeline
from evidently.report import Report
from evidently.metric_preset import (
    DataDriftPreset,
    DataQualityPreset,
    TargetDriftPreset,
)

def generate_monitoring_report(reference_df, production_df):
    report = Report(metrics=[
        DataDriftPreset(
            stattest="ks",
            stattest_threshold=0.05,
            drift_share=0.3     # alert if >30% of features drift
        ),
        DataQualityPreset(),
        TargetDriftPreset(),
    ])
    report.run(
        reference_data=reference_df,
        current_data=production_df,
        column_mapping=column_mapping
    )
    return report

# Schedule hourly monitoring
def hourly_monitoring_job():
    reference = load_reference_data()
    production = load_last_hour_predictions()

    report = generate_monitoring_report(reference, production)
    report_json = report.as_dict()

    # Check for drift
    drift_detected = report_json["metrics"][0]["result"]["dataset_drift"]
    drift_share = report_json["metrics"][0]["result"]["drift_share"]

    if drift_detected:
        send_alert(
            severity="warning",
            message=f"Data drift detected: {drift_share:.1%} of features drifted",
            report_url=save_report(report)
        )

    # Log metrics to monitoring backend
    log_metrics({
        "drift_share": drift_share,
        "data_quality_issues": count_quality_issues(report_json),
        "prediction_volume": len(production),
        "timestamp": datetime.utcnow()
    })

Automated Retraining Triggers

Not every drift signal warrants retraining. Feature drift without performance impact may reflect benign distribution shifts. Automated retraining should trigger only when performance degradation is confirmed or strongly predicted, and the retraining infrastructure should validate the new model against the old one before deployment.

# Retraining decision logic
def should_retrain(monitoring_metrics, history):
    # Hard trigger: confirmed performance drop
    if monitoring_metrics.get("actual_performance"):
        if monitoring_metrics["actual_performance"] < PERFORMANCE_THRESHOLD:
            return True, "confirmed_performance_drop"

    # Soft trigger: estimated performance drop + sustained drift
    if monitoring_metrics.get("estimated_performance"):
        est_perf = monitoring_metrics["estimated_performance"]
        drift_days = count_consecutive_drift_days(history)

        if est_perf < PERFORMANCE_THRESHOLD and drift_days >= 3:
            return True, "estimated_drop_with_sustained_drift"

    # Preventive trigger: severe drift without performance signal
    if monitoring_metrics.get("drift_share", 0) > 0.5:
        drift_days = count_consecutive_drift_days(history)
        if drift_days >= 7:
            return True, "severe_sustained_drift"

    return False, "no_action"

# Retraining pipeline with validation
def retrain_and_validate(trigger_reason):
    # Build training dataset with sliding window
    new_training_data = build_training_set(
        window_days=90,
        include_original_ratio=0.2   # 20% original training data
    )

    # Train candidate model
    candidate = train_model(new_training_data)

    # Validate against current model
    validation = compare_models(
        current_model=load_production_model(),
        candidate_model=candidate,
        test_data=load_holdout_set()
    )

    if validation["candidate_better"]:
        deploy_model(candidate, reason=trigger_reason)
        log_retraining(status="deployed", validation=validation)
    else:
        log_retraining(status="rejected", validation=validation)
        escalate_to_team("Retraining did not improve performance")

The sliding window approach for retraining combines recent data (which reflects current distribution) with a fraction of the original training data (which prevents forgetting important patterns). The 80/20 ratio between new and original data is a starting point — adjust based on how fundamentally the data has shifted. If concept drift is severe, reducing the original data fraction allows the model to fully adapt to the new regime.

Monitoring Tool Comparison

ToolDrift DetectionPerformance EstimationDeploymentBest For
Evidently AIComprehensive stat testsBasicSelf-hosted / cloudOpen-source flexibility
NannyMLUnivariate + multivariateCBPE + DLESelf-hosted / cloudLabel-delay scenarios
ArizeAuto-detectionBuilt-inManaged cloudEnterprise scale
WhylabsStatistical profilingBasicManaged cloudData-centric monitoring
FiddlerDrift + explainabilityBuilt-inManaged cloudRegulated industries

Alert Design

Monitoring alerts should follow the same principles as infrastructure alerting: page on symptoms that affect users, not on every metric movement. A single feature drifting slightly is not actionable. Aggregate drift across feature groups, correlate with performance metrics, and set severity tiers that match your response capabilities.

# Tiered alerting configuration
alert_config = {
    "critical": {
        "conditions": [
            "actual_performance < 0.65",
            "prediction_error_rate > 0.15"
        ],
        "channels": ["pagerduty", "slack-ml-oncall"],
        "response_time": "30 minutes"
    },
    "warning": {
        "conditions": [
            "estimated_performance < 0.75",
            "drift_share > 0.3 for 3 consecutive checks",
            "data_quality_score < 0.9"
        ],
        "channels": ["slack-ml-team"],
        "response_time": "next business day"
    },
    "info": {
        "conditions": [
            "any single feature drift detected",
            "prediction volume deviation > 20%"
        ],
        "channels": ["monitoring-dashboard"],
        "response_time": "weekly review"
    }
}

Over-alerting is worse than under-alerting for model monitoring. Teams that receive too many false positives stop investigating alerts, which means they miss real degradation events. Start with conservative thresholds that produce fewer but higher-precision alerts, then tighten gradually as you build confidence in your monitoring signals. The operational patterns here mirror those in embedding model evaluation, where you need to distinguish signal from noise across many dimensions.

Ground Truth Feedback Loops

The monitoring system is only as good as the feedback loop that connects predictions to outcomes. For some models, ground truth arrives immediately (click prediction → did user click). For others, it arrives days or weeks later (loan default → did the loan default in 12 months). Design the monitoring pipeline to handle delayed labels by backfilling performance metrics when labels arrive and retroactively evaluating whether drift signals correlated with actual degradation.

Track two things: the real-time proxy metrics that give early warning, and the delayed ground truth metrics that confirm or refute those warnings. Over time, calibrate your proxy metrics by measuring how reliably they predict ground truth degradation. A proxy metric that frequently fires without corresponding performance drops is adding noise, not signal. A proxy that consistently misses real degradation events needs different statistical tests or tighter thresholds.

The monitoring infrastructure itself needs monitoring. Track metric computation latency, data pipeline freshness, and alert delivery reliability. A monitoring system that silently fails is worse than no monitoring at all — it provides false confidence that the model is healthy when no one is actually watching.