smaple.tr
MLOps

MLOps Guide: Taking Machine Learning Models to Production [2026]

Mehmet Kurtipek
August 11, 2026
13 min read
MLOps
machine learning
MLflow
Kubeflow
model deployment
CI/CD
ml pipeline

87% of machine learning models built by data science teams never reach production. The models work — they pass cross-validation, they score well on holdout sets, they demonstrate genuine predictive value. The problem is not the modeling. The problem is everything that happens between a notebook experiment and a reliable, monitored, production system. MLOps is the discipline that closes that gap.

This guide covers the full MLOps stack: maturity levels, tooling choices (MLflow, DVC, Kubeflow), CI/CD pipeline design, model serving options, and production monitoring. By the end, you will have a clear map of what MLOps infrastructure looks like at each scale, and which components to build first.

MLOps Guide: Core Components and Tools

MLOps (Machine Learning Operations) applies DevOps principles to the machine learning lifecycle. Where DevOps automates the path from code commit to deployed application, MLOps automates the path from training experiment to deployed model — and then keeps the model healthy once it is live.

The five core components of an MLOps system are:

  • Data versioning — tracking which data was used to train which model
  • Experiment tracking — recording hyperparameters, metrics, and artifacts across training runs
  • Model registry — managing model versions and their promotion stages (staging, production, archived)
  • Automated pipeline — CI/CD for model training, testing, and deployment
  • Production monitoring — detecting data drift, concept drift, and model degradation

Each component can be implemented incrementally. A team does not need all five to start. But teams that skip all five tend to produce models that go stale silently, cannot be reproduced, and require manual intervention to update.

MLOps Maturity Levels

Google Cloud and Gartner both describe MLOps maturity as a staged progression. Four levels capture the range from ad-hoc to fully automated.

Level 0: Manual (Ad-Hoc)

Everything is done by hand. Data preparation happens in notebooks. Training is triggered manually. Deployment is a direct file copy or script run. There is no experiment tracking, no version control for data or models, no monitoring. A model is "in production" when the engineer who built it last remembers how it was deployed.

This describes the majority of ML projects at their start. The risk profile is high: a data scientist leaves the team and nobody can reproduce the model. Input data shifts and nobody knows until users complain.

Level 1: Scripted Pipeline

Training is scripted and reproducible. Data versioning begins (DVC or equivalent). Experiments are tracked (MLflow or equivalent). There is a basic CI/CD pipeline, but it includes manual review and approval gates. Staging and production environments are separated.

This level is achievable by a team of two to three engineers within a month of focused effort. It eliminates the "only one person knows how to deploy this" problem.

Level 2: Automated ML Pipeline

End-to-end training and deployment is automated. Model registry manages version promotion. Automated tests gate deployment: accuracy regression tests, data validation, smoke tests on the serving endpoint. Kubernetes-based model serving is operational. Development and production environments are identical.

At this level, a change to training data or model code triggers an automated pipeline that trains, tests, registers, and deploys — with human approval only at the final production promotion step.

Level 3: Self-Healing Operations

Data drift detection triggers automatic retraining. A/B testing and canary deployment are automated. Performance anomalies are detected and reported automatically. Human intervention is reserved for exceptional cases — policy decisions, budget constraints, or novel failure modes outside the alert thresholds.

Smart Maple's Oplist platform reached Level 2-3 in its ML infrastructure for healthcare scheduling optimization. When patient demand patterns shift in real-time, models move through automated retraining and canary deployment without engineer intervention. The result: model accuracy at 91.2% (up from 73% with manual scheduling), with deployment frequency at 3-5 updates per day instead of once per month.

Tool Comparison: MLflow vs Kubeflow vs Vertex AI vs SageMaker

Choosing your MLOps stack is one of the first decisions that shapes everything else. The table below covers the four most widely adopted platforms:

Feature MLflow Kubeflow Vertex AI SageMaker
Primary focus Experiment tracking, model registry End-to-end pipeline orchestration Google Cloud integration AWS integration
Learning curve Low-Medium High Medium Medium
Data versioning No (use DVC) No Native Native
Model registry Native, strong Limited Native Native
Pipeline orchestration Basic (Projects) Advanced (Kubernetes-native) Native Native
Hosting Self-hosted or MLflow Cloud Requires Kubernetes cluster Google Cloud AWS
Cost model Open source (free) Open source (infrastructure cost) Pay-per-use (~$2K-5K/month) Pay-per-use (~$2K-5K/month)
Enterprise support Databricks Linux Foundation Google AWS

For teams that want cloud-provider independence and are comfortable running Kubernetes, MLflow plus Kubeflow is a proven combination. MLflow handles experiment tracking and the model registry. Kubeflow orchestrates the training and deployment pipeline on your own cluster.

For teams already committed to AWS or Google Cloud, SageMaker and Vertex AI reduce operational overhead by consolidating tooling within the cloud provider's ecosystem.

Data Versioning with DVC

In ML projects, data is as important as code — often more so. If you cannot reproduce the exact dataset that produced a given model, you cannot debug that model, audit its behavior, or safely redeploy it after an incident.

DVC (Data Version Control) solves this by extending Git's version control semantics to large data files and model artifacts. Large files are stored in a remote backend (S3, GCP Cloud Storage, Azure Blob) while the versioning metadata lives in Git. Every Git commit can reference the exact dataset version used for training.

Key DVC capabilities for MLOps pipelines:

  • Pipeline definition — dvc.yaml defines stages (data prep, training, evaluation) and their dependencies
  • Reproducibility — dvc repro replays a pipeline only for stages whose inputs changed
  • Remote storage — large files stay out of Git but remain version-controlled
  • Data diff — compare dataset versions between experiments

When integrated with CI/CD, DVC enables the pipeline to pull the correct dataset version before training, ensuring that any run of the pipeline can be reproduced precisely weeks or months later.

Experiment Tracking with MLflow

During model development, a data scientist runs dozens of experiments: different hyperparameters, different feature engineering choices, different model architectures. Without systematic tracking, the best configuration gets lost. The same experiment gets run twice. The basis for a deployment decision becomes "I think this was the best one."

MLflow Tracking solves this. Every experiment run logs:

  • Parameters — hyperparameters passed to the model
  • Metrics — accuracy, F1, AUC, loss at each epoch
  • Artifacts — the trained model file, plots, confusion matrices

Results are queryable through a central dashboard visible to the entire team. Runs can be compared side-by-side. The artifact for the winning run is ready for promotion to the model registry.

MLflow has three additional components beyond Tracking:

  • Projects — defines how to run training in a reproducible environment (conda, Docker)
  • Models — standardizes model packaging for deployment across serving frameworks
  • Registry — manages model versions and their lifecycle stages

The Registry is what connects experiment tracking to production deployment. A model moves from None (newly registered) to Staging (under test) to Production (live traffic) to Archived (retired). This stage progression gives the deployment workflow a concrete handoff point: the CI/CD pipeline promotes a model to Production only after automated tests pass in Staging.

CI/CD for Machine Learning in Production

Traditional software CI/CD runs tests against code. ML CI/CD has additional steps: validate input data, retrain the model, evaluate model performance, and test the serving endpoint.

A typical ML CI/CD pipeline:

  1. Data pull — DVC fetches the correct dataset version
  2. Data validation — schema checks, distribution checks, feature range checks
  3. Model training — training script runs, MLflow logs results
  4. Performance gate — compare new model accuracy against current production baseline; fail pipeline if regression detected
  5. Build serving artifact — package model for deployment (Docker image, ONNX export, or MLflow model format)
  6. Deploy to staging — deploy to staging endpoint, run integration and load tests
  7. Canary deployment — route 5-10% of production traffic to new model
  8. Monitor canary — watch latency, error rate, prediction distribution for 24-48 hours
  9. Full promotion — if canary metrics pass thresholds, promote to 100%
  10. Rollback on failure — automated rollback to previous model version if metrics degrade

This pipeline structure converts what was a 3-day manual deployment process (training, review, deploy) into a 15-minute automated process with rollback. The key difference from traditional software CI/CD is the performance gate in step 4: you are not just testing that the code runs, you are testing that the resulting model is better than what is currently in production.

Model Serving Options

Once a model passes the CI/CD pipeline, it needs to be served at inference time. The right serving choice depends on your latency requirements, throughput needs, and model type.

Option FastAPI TensorFlow Serving Triton Inference Server
Best for General-purpose, sklearn, XGBoost TensorFlow/Keras deep learning High-performance, multi-framework
Latency 50-200ms 20-100ms 10-50ms
Throughput 100-500 req/s 1000+ req/s 5000+ req/s
Memory Low (~500MB) Medium (~2-4GB) High (~4-8GB)
Scaling Kubernetes pod replicas Model server instances Kubernetes pod replicas
Setup complexity Low Medium High

For most teams starting out with ML in production, FastAPI is the right choice. It requires minimal infrastructure, supports any Python-native model format (pickle, ONNX, SavedModel), and integrates naturally with MLflow's model loading API. A production FastAPI serving endpoint loads the current production model from the MLflow registry at startup, exposes a /predict endpoint, and includes a /health endpoint for the Kubernetes liveness probe.

TensorFlow Serving and Triton make sense once throughput requirements exceed what a single FastAPI instance can handle, or when serving latency is measured in single-digit milliseconds.

Production Monitoring: Detecting Drift Before It Becomes a Failure

Deploying a model is not the end of the MLOps story. In production, three failure modes accumulate silently over time:

Data drift — the distribution of incoming features shifts away from the training distribution. A demand forecasting model trained on 2024 purchasing patterns faces drift after a major market event. Predictions remain structurally valid but numerically miscalibrated.

Concept drift — the relationship between features and the target variable changes. A fraud detection model trained before a new fraud campaign launches faces concept drift: the old patterns that predicted fraud are no longer predictive.

Model degradation — a combination of drift effects that causes measurable accuracy decline. Usually detected after the fact, by comparing model predictions against ground truth labels collected in production.

An effective monitoring setup catches each failure mode:

  • Feature distribution monitoring — compare incoming feature distributions against training distributions using statistical tests (KS test, PSI)
  • Prediction distribution monitoring — track the distribution of model outputs; a shift often precedes measurable accuracy degradation
  • Ground truth comparison — when labels arrive (delayed), compare against model predictions to compute actual production accuracy
  • Latency and throughput monitoring — infrastructure-level signals that catch serving problems before users are affected

Prometheus and Grafana are the standard stack for ML monitoring dashboards. Evidently AI and Whylogs are purpose-built libraries for data drift detection. Arize, Fiddler, and Aporia provide commercial platforms that combine drift detection, explainability, and alerting.

MLOps Infrastructure Costs

Building MLOps infrastructure requires real investment. Understanding cost ranges helps teams budget appropriately and avoid over-engineering for their scale.

Small scale (1-3 models, <10M records/month): Kubernetes cluster (3-node), MLflow server (single instance), S3 storage (50GB), PostgreSQL (RDS t3.small): approximately $370/month. Handles teams of two to three data scientists running weekly retraining cycles.

Medium scale (5-10 models, 100M records/month): Larger Kubernetes cluster (6-node), HA MLflow, expanded storage, monitoring suite: approximately $1,500/month. Handles teams of five to eight engineers with daily retraining.

Enterprise scale (20+ models, 1B+ records/month): Multi-region Kubernetes, dedicated GPU nodes for training, enterprise monitoring, SLA support: approximately $10,000/month and up.

The cost-benefit calculation for MLOps investment typically compares infrastructure spend against the cost of the problems it prevents: a single production model failure that causes degraded user experience for three days often costs more in engineering time (diagnosis, rollback, re-investigation) than a month of MLOps infrastructure.

MLOps Best Practices

Ten practices that experienced teams consistently apply:

  1. Version everything — code, data, model artifacts, and configuration all under version control
  2. Automate testing — unit tests for data pipelines, integration tests for model serving, performance regression tests
  3. Monitor continuously — data drift and performance degradation rarely announce themselves; you need instrumentation to catch them
  4. Separate roles — data engineering, ML engineering, and ML operations are distinct disciplines with distinct tooling
  5. Write model cards — document what a model is trained on, what it is designed to do, and what its failure modes are
  6. Use canary deployment — never route 100% of traffic to an unproven model; canary gives you rollback options
  7. Make rollback fast — the ability to revert to the previous production model in under five minutes is non-negotiable
  8. Review cloud costs — ML infrastructure costs compound; unoptimized training jobs and idle GPU instances accumulate
  9. Enforce access control — model registry access, training data access, and production serving credentials need explicit permission management
  10. Invest in MLOps education — data scientists who understand production concerns produce models that are easier to operationalize

Getting Started: A Practical Sequence

For teams at Level 0 (pure ad-hoc), the practical sequence to reach Level 2 in three months:

Month 1: Add MLflow Tracking to existing training scripts. Every experiment run logs parameters and metrics. Stand up a shared MLflow server (a single EC2 t3.medium or equivalent is sufficient). Stop making deployment decisions based on memory.

Month 2: Introduce DVC for data versioning on your most important dataset. Write a basic CI/CD pipeline that runs training on merge to main, logs results to MLflow, and requires manual approval to promote to production.

Month 3: Add a performance gate to the CI/CD pipeline. Deploy a FastAPI serving endpoint. Add prediction logging and basic distribution monitoring. Define rollback procedure.

By month 3, the team has functional MLOps at Level 1-2. Level 3 features (automated retraining on drift, self-healing operations) come after the foundation is stable.

The 87% failure rate for ML models is not a research curiosity — it reflects what happens when teams treat deployment as an afterthought rather than a first-class engineering concern. MLOps is the answer to that problem. The tooling exists, the patterns are proven, and the investment is justified by the value of models that actually run in production.

Related Articles

August 10, 2026

LLM Fine-Tuning Guide: Custom Model Training with LoRA and QLoRA [2026]

General-purpose LLMs are impressive. They can write code, summarize documents, answer questions, and translate between languages with reasonable accuracy. But "reasonable" is not good enough when your application requires consistent output format, domain-specific terminology, a particular tone, or behavior that the base model was never trained to exhibit. That gap is where fine-tuning matters. Fine-tuning updates a model's weights on your specific data, changing how the model behaves — not

Read More
August 9, 2026

Computer Vision Applications: Object Detection, OCR, and Industrial AI [2026]

Computer vision has moved well past the research phase. The models are trained, the frameworks are mature, the hardware is accessible, and the use cases are generating measurable returns. What was a specialized capability requiring deep expertise in 2018 is now deployable infrastructure — if you know which component to reach for and where the real complexity lives. This guide covers computer vision applications across industrial, medical, logistics, and document processing domains. It expl

Read More
August 8, 2026

Natural Language Processing Guide: Building Production NLP Systems [2026]

Natural language processing has crossed a threshold. For most of the 2010s, production NLP meant rule-based pipelines, hand-crafted feature extractors, and models that were brittle the moment text deviated from the training distribution. Today, transformer-based models — BERT, RoBERTa, and their descendants — handle tasks that once required months of annotation work with a few hours of fine-tuning. This guide covers what NLP actually looks like in production in 2026: the architectures that w

Read More