87% of machine learning models built by data science teams never reach production. The models work — they pass cross-validation, they score well on holdout sets, they demonstrate genuine predictive value. The problem is not the modeling. The problem is everything that happens between a notebook experiment and a reliable, monitored, production system. MLOps is the discipline that closes that gap.
This guide covers the full MLOps stack: maturity levels, tooling choices (MLflow, DVC, Kubeflow), CI/CD pipeline design, model serving options, and production monitoring. By the end, you will have a clear map of what MLOps infrastructure looks like at each scale, and which components to build first.
MLOps Guide: Core Components and Tools
MLOps (Machine Learning Operations) applies DevOps principles to the machine learning lifecycle. Where DevOps automates the path from code commit to deployed application, MLOps automates the path from training experiment to deployed model — and then keeps the model healthy once it is live.
The five core components of an MLOps system are:
- Data versioning — tracking which data was used to train which model
- Experiment tracking — recording hyperparameters, metrics, and artifacts across training runs
- Model registry — managing model versions and their promotion stages (staging, production, archived)
- Automated pipeline — CI/CD for model training, testing, and deployment
- Production monitoring — detecting data drift, concept drift, and model degradation
Each component can be implemented incrementally. A team does not need all five to start. But teams that skip all five tend to produce models that go stale silently, cannot be reproduced, and require manual intervention to update.
MLOps Maturity Levels
Google Cloud and Gartner both describe MLOps maturity as a staged progression. Four levels capture the range from ad-hoc to fully automated.
Level 0: Manual (Ad-Hoc)
Everything is done by hand. Data preparation happens in notebooks. Training is triggered manually. Deployment is a direct file copy or script run. There is no experiment tracking, no version control for data or models, no monitoring. A model is "in production" when the engineer who built it last remembers how it was deployed.
This describes the majority of ML projects at their start. The risk profile is high: a data scientist leaves the team and nobody can reproduce the model. Input data shifts and nobody knows until users complain.
Level 1: Scripted Pipeline
Training is scripted and reproducible. Data versioning begins (DVC or equivalent). Experiments are tracked (MLflow or equivalent). There is a basic CI/CD pipeline, but it includes manual review and approval gates. Staging and production environments are separated.
This level is achievable by a team of two to three engineers within a month of focused effort. It eliminates the "only one person knows how to deploy this" problem.
Level 2: Automated ML Pipeline
End-to-end training and deployment is automated. Model registry manages version promotion. Automated tests gate deployment: accuracy regression tests, data validation, smoke tests on the serving endpoint. Kubernetes-based model serving is operational. Development and production environments are identical.
At this level, a change to training data or model code triggers an automated pipeline that trains, tests, registers, and deploys — with human approval only at the final production promotion step.
Level 3: Self-Healing Operations
Data drift detection triggers automatic retraining. A/B testing and canary deployment are automated. Performance anomalies are detected and reported automatically. Human intervention is reserved for exceptional cases — policy decisions, budget constraints, or novel failure modes outside the alert thresholds.
Smart Maple's Oplist platform reached Level 2-3 in its ML infrastructure for healthcare scheduling optimization. When patient demand patterns shift in real-time, models move through automated retraining and canary deployment without engineer intervention. The result: model accuracy at 91.2% (up from 73% with manual scheduling), with deployment frequency at 3-5 updates per day instead of once per month.
Tool Comparison: MLflow vs Kubeflow vs Vertex AI vs SageMaker
Choosing your MLOps stack is one of the first decisions that shapes everything else. The table below covers the four most widely adopted platforms:
| Feature | MLflow | Kubeflow | Vertex AI | SageMaker |
|---|---|---|---|---|
| Primary focus | Experiment tracking, model registry | End-to-end pipeline orchestration | Google Cloud integration | AWS integration |
| Learning curve | Low-Medium | High | Medium | Medium |
| Data versioning | No (use DVC) | No | Native | Native |
| Model registry | Native, strong | Limited | Native | Native |
| Pipeline orchestration | Basic (Projects) | Advanced (Kubernetes-native) | Native | Native |
| Hosting | Self-hosted or MLflow Cloud | Requires Kubernetes cluster | Google Cloud | AWS |
| Cost model | Open source (free) | Open source (infrastructure cost) | Pay-per-use (~$2K-5K/month) | Pay-per-use (~$2K-5K/month) |
| Enterprise support | Databricks | Linux Foundation | AWS |
For teams that want cloud-provider independence and are comfortable running Kubernetes, MLflow plus Kubeflow is a proven combination. MLflow handles experiment tracking and the model registry. Kubeflow orchestrates the training and deployment pipeline on your own cluster.
For teams already committed to AWS or Google Cloud, SageMaker and Vertex AI reduce operational overhead by consolidating tooling within the cloud provider's ecosystem.
Data Versioning with DVC
In ML projects, data is as important as code — often more so. If you cannot reproduce the exact dataset that produced a given model, you cannot debug that model, audit its behavior, or safely redeploy it after an incident.
DVC (Data Version Control) solves this by extending Git's version control semantics to large data files and model artifacts. Large files are stored in a remote backend (S3, GCP Cloud Storage, Azure Blob) while the versioning metadata lives in Git. Every Git commit can reference the exact dataset version used for training.
Key DVC capabilities for MLOps pipelines:
- Pipeline definition —
dvc.yamldefines stages (data prep, training, evaluation) and their dependencies - Reproducibility —
dvc reproreplays a pipeline only for stages whose inputs changed - Remote storage — large files stay out of Git but remain version-controlled
- Data diff — compare dataset versions between experiments
When integrated with CI/CD, DVC enables the pipeline to pull the correct dataset version before training, ensuring that any run of the pipeline can be reproduced precisely weeks or months later.
Experiment Tracking with MLflow
During model development, a data scientist runs dozens of experiments: different hyperparameters, different feature engineering choices, different model architectures. Without systematic tracking, the best configuration gets lost. The same experiment gets run twice. The basis for a deployment decision becomes "I think this was the best one."
MLflow Tracking solves this. Every experiment run logs:
- Parameters — hyperparameters passed to the model
- Metrics — accuracy, F1, AUC, loss at each epoch
- Artifacts — the trained model file, plots, confusion matrices
Results are queryable through a central dashboard visible to the entire team. Runs can be compared side-by-side. The artifact for the winning run is ready for promotion to the model registry.
MLflow has three additional components beyond Tracking:
- Projects — defines how to run training in a reproducible environment (conda, Docker)
- Models — standardizes model packaging for deployment across serving frameworks
- Registry — manages model versions and their lifecycle stages
The Registry is what connects experiment tracking to production deployment. A model moves from None (newly registered) to Staging (under test) to Production (live traffic) to Archived (retired). This stage progression gives the deployment workflow a concrete handoff point: the CI/CD pipeline promotes a model to Production only after automated tests pass in Staging.
CI/CD for Machine Learning in Production
Traditional software CI/CD runs tests against code. ML CI/CD has additional steps: validate input data, retrain the model, evaluate model performance, and test the serving endpoint.
A typical ML CI/CD pipeline:
- Data pull — DVC fetches the correct dataset version
- Data validation — schema checks, distribution checks, feature range checks
- Model training — training script runs, MLflow logs results
- Performance gate — compare new model accuracy against current production baseline; fail pipeline if regression detected
- Build serving artifact — package model for deployment (Docker image, ONNX export, or MLflow model format)
- Deploy to staging — deploy to staging endpoint, run integration and load tests
- Canary deployment — route 5-10% of production traffic to new model
- Monitor canary — watch latency, error rate, prediction distribution for 24-48 hours
- Full promotion — if canary metrics pass thresholds, promote to 100%
- Rollback on failure — automated rollback to previous model version if metrics degrade
This pipeline structure converts what was a 3-day manual deployment process (training, review, deploy) into a 15-minute automated process with rollback. The key difference from traditional software CI/CD is the performance gate in step 4: you are not just testing that the code runs, you are testing that the resulting model is better than what is currently in production.
Model Serving Options
Once a model passes the CI/CD pipeline, it needs to be served at inference time. The right serving choice depends on your latency requirements, throughput needs, and model type.
| Option | FastAPI | TensorFlow Serving | Triton Inference Server |
|---|---|---|---|
| Best for | General-purpose, sklearn, XGBoost | TensorFlow/Keras deep learning | High-performance, multi-framework |
| Latency | 50-200ms | 20-100ms | 10-50ms |
| Throughput | 100-500 req/s | 1000+ req/s | 5000+ req/s |
| Memory | Low (~500MB) | Medium (~2-4GB) | High (~4-8GB) |
| Scaling | Kubernetes pod replicas | Model server instances | Kubernetes pod replicas |
| Setup complexity | Low | Medium | High |
For most teams starting out with ML in production, FastAPI is the right choice. It requires minimal infrastructure, supports any Python-native model format (pickle, ONNX, SavedModel), and integrates naturally with MLflow's model loading API. A production FastAPI serving endpoint loads the current production model from the MLflow registry at startup, exposes a /predict endpoint, and includes a /health endpoint for the Kubernetes liveness probe.
TensorFlow Serving and Triton make sense once throughput requirements exceed what a single FastAPI instance can handle, or when serving latency is measured in single-digit milliseconds.
Production Monitoring: Detecting Drift Before It Becomes a Failure
Deploying a model is not the end of the MLOps story. In production, three failure modes accumulate silently over time:
Data drift — the distribution of incoming features shifts away from the training distribution. A demand forecasting model trained on 2024 purchasing patterns faces drift after a major market event. Predictions remain structurally valid but numerically miscalibrated.
Concept drift — the relationship between features and the target variable changes. A fraud detection model trained before a new fraud campaign launches faces concept drift: the old patterns that predicted fraud are no longer predictive.
Model degradation — a combination of drift effects that causes measurable accuracy decline. Usually detected after the fact, by comparing model predictions against ground truth labels collected in production.
An effective monitoring setup catches each failure mode:
- Feature distribution monitoring — compare incoming feature distributions against training distributions using statistical tests (KS test, PSI)
- Prediction distribution monitoring — track the distribution of model outputs; a shift often precedes measurable accuracy degradation
- Ground truth comparison — when labels arrive (delayed), compare against model predictions to compute actual production accuracy
- Latency and throughput monitoring — infrastructure-level signals that catch serving problems before users are affected
Prometheus and Grafana are the standard stack for ML monitoring dashboards. Evidently AI and Whylogs are purpose-built libraries for data drift detection. Arize, Fiddler, and Aporia provide commercial platforms that combine drift detection, explainability, and alerting.
MLOps Infrastructure Costs
Building MLOps infrastructure requires real investment. Understanding cost ranges helps teams budget appropriately and avoid over-engineering for their scale.
Small scale (1-3 models, <10M records/month): Kubernetes cluster (3-node), MLflow server (single instance), S3 storage (50GB), PostgreSQL (RDS t3.small): approximately $370/month. Handles teams of two to three data scientists running weekly retraining cycles.
Medium scale (5-10 models, 100M records/month): Larger Kubernetes cluster (6-node), HA MLflow, expanded storage, monitoring suite: approximately $1,500/month. Handles teams of five to eight engineers with daily retraining.
Enterprise scale (20+ models, 1B+ records/month): Multi-region Kubernetes, dedicated GPU nodes for training, enterprise monitoring, SLA support: approximately $10,000/month and up.
The cost-benefit calculation for MLOps investment typically compares infrastructure spend against the cost of the problems it prevents: a single production model failure that causes degraded user experience for three days often costs more in engineering time (diagnosis, rollback, re-investigation) than a month of MLOps infrastructure.
MLOps Best Practices
Ten practices that experienced teams consistently apply:
- Version everything — code, data, model artifacts, and configuration all under version control
- Automate testing — unit tests for data pipelines, integration tests for model serving, performance regression tests
- Monitor continuously — data drift and performance degradation rarely announce themselves; you need instrumentation to catch them
- Separate roles — data engineering, ML engineering, and ML operations are distinct disciplines with distinct tooling
- Write model cards — document what a model is trained on, what it is designed to do, and what its failure modes are
- Use canary deployment — never route 100% of traffic to an unproven model; canary gives you rollback options
- Make rollback fast — the ability to revert to the previous production model in under five minutes is non-negotiable
- Review cloud costs — ML infrastructure costs compound; unoptimized training jobs and idle GPU instances accumulate
- Enforce access control — model registry access, training data access, and production serving credentials need explicit permission management
- Invest in MLOps education — data scientists who understand production concerns produce models that are easier to operationalize
Getting Started: A Practical Sequence
For teams at Level 0 (pure ad-hoc), the practical sequence to reach Level 2 in three months:
Month 1: Add MLflow Tracking to existing training scripts. Every experiment run logs parameters and metrics. Stand up a shared MLflow server (a single EC2 t3.medium or equivalent is sufficient). Stop making deployment decisions based on memory.
Month 2: Introduce DVC for data versioning on your most important dataset. Write a basic CI/CD pipeline that runs training on merge to main, logs results to MLflow, and requires manual approval to promote to production.
Month 3: Add a performance gate to the CI/CD pipeline. Deploy a FastAPI serving endpoint. Add prediction logging and basic distribution monitoring. Define rollback procedure.
By month 3, the team has functional MLOps at Level 1-2. Level 3 features (automated retraining on drift, self-healing operations) come after the foundation is stable.
The 87% failure rate for ML models is not a research curiosity — it reflects what happens when teams treat deployment as an afterthought rather than a first-class engineering concern. MLOps is the answer to that problem. The tooling exists, the patterns are proven, and the investment is justified by the value of models that actually run in production.
Related Articles
LLM Fine-Tuning Guide: Custom Model Training with LoRA and QLoRA [2026]
General-purpose LLMs are impressive. They can write code, summarize documents, answer questions, and translate between languages with reasonable accuracy. But "reasonable" is not good enough when your application requires consistent output format, domain-specific terminology, a particular tone, or behavior that the base model was never trained to exhibit. That gap is where fine-tuning matters. Fine-tuning updates a model's weights on your specific data, changing how the model behaves — not
Read MoreComputer Vision Applications: Object Detection, OCR, and Industrial AI [2026]
Computer vision has moved well past the research phase. The models are trained, the frameworks are mature, the hardware is accessible, and the use cases are generating measurable returns. What was a specialized capability requiring deep expertise in 2018 is now deployable infrastructure — if you know which component to reach for and where the real complexity lives. This guide covers computer vision applications across industrial, medical, logistics, and document processing domains. It expl
Read MoreNatural Language Processing Guide: Building Production NLP Systems [2026]
Natural language processing has crossed a threshold. For most of the 2010s, production NLP meant rule-based pipelines, hand-crafted feature extractors, and models that were brittle the moment text deviated from the training distribution. Today, transformer-based models — BERT, RoBERTa, and their descendants — handle tasks that once required months of annotation work with a few hours of fine-tuning. This guide covers what NLP actually looks like in production in 2026: the architectures that w
Read More
