87% of machine learning models never reach production. Among the ones that do reach production, a significant share degrade within 90 days to accuracy levels that make them operationally useless. The most common cause in both cases is not model architecture — it is data infrastructure. The model is only as good as the data pipeline behind it.
Data engineering MLOps integration is the practice of building data infrastructure specifically designed to support ML system requirements: reproducible training datasets, feature consistency between training and inference, automated drift detection, and data versioning that lets teams trace exactly which data produced which model. This guide covers the data engineering side of MLOps — not ML operations governance or AI workflow management, but the specific data pipeline patterns that make the difference between a model that trains in a notebook and a system that delivers reliable predictions in production for months or years.
Data Engineering MLOps Integration: Why Data Is the Bottleneck
An ML model depends on three things: algorithm architecture, hyperparameters, and data. Data engineers control exactly one of these. But the failure patterns in production ML systems are predominantly data failures:
Training-serving skew — the feature values used to train the model are computed differently than the feature values used at inference time. The model learned from one distribution; it predicts on another. Result: predictions that are technically valid but systematically wrong.
Data drift — the statistical distribution of input data changes after the model is trained. Customer behavior patterns shift, seasonal effects compound, upstream schema changes alter field values. The model's learned patterns become stale. Without detection, the degradation is silent.
Temporal leakage — training data includes future information that would not be available at prediction time. A fraud detection model trained on features computed from the future ("this customer will make 3 more transactions today") appears to perform well in offline evaluation but fails in production.
Dataset irreproducibility — a model is retrained, performance degrades, and the team cannot determine whether the regression is due to model changes or data changes because training dataset composition is not versioned.
These are data engineering problems. Solving them requires purpose-built data infrastructure for ML — not just a data warehouse and a dbt pipeline.
Feature Store: The Central Data Engineering MLOps Component
A feature store is the architectural component that resolves the most persistent data engineering challenge in ML: ensuring that the features used to train a model are identical to the features computed at inference time.
What a Feature Store Solves
Without a feature store, features are typically computed twice: once in a batch training pipeline (using historical data in the warehouse) and once in the inference service (using real-time operational data). Even when the computation logic is identical on paper, subtle differences emerge — timezone handling, null treatment, aggregation window boundaries. These differences become training-serving skew.
A feature store provides one computation definition that serves both purposes. Features are defined once (in Python or YAML), computed by a centralized pipeline, and stored in two backends: an offline store (historical features for training) and an online store (current features for low-latency inference).
Feature Store Architecture
Offline store — a historical record of every feature value for every entity across time. Implemented on Snowflake tables, S3 Parquet files, or BigQuery tables. Used to construct point-in-time correct training datasets: "give me the feature values for customer_id=123 at 2025-01-15T14:30:00" — the values that would have been available at that exact moment, not today's values.
Point-in-time correctness is critical for avoiding temporal leakage. The offline store's time travel capability returns historical feature snapshots, not current values.
Online store — a low-latency key-value store (Redis, DynamoDB, or Cassandra) containing the most recent feature values for each entity. Inference requests query the online store with microsecond latency — the 10–100ms budget for a fraud detection prediction cannot accommodate a Snowflake SQL query.
Feature pipeline — the computation layer that reads from source systems, calculates feature values, and writes to both stores. Runs on a schedule (batch features) or processes a streaming data source (real-time features).
Feature Store Tools
Feast (open source, sponsored by Tecton/Google) — the most widely deployed open-source feature store. Supports multiple offline backends (BigQuery, Snowflake, Parquet), multiple online backends (Redis, DynamoDB, SQLite), and a Python SDK for feature definition and retrieval. Free. Requires operational investment to deploy and maintain.
Tecton — the fully managed, enterprise version of the Feast pattern. Higher reliability, managed infrastructure, better observability. Cost: $3,000–10,000/month depending on feature volume. Appropriate for organizations where feature infrastructure downtime is a P1 incident.
Vertex AI Feature Store (Google Cloud) — integrated with BigQuery and Vertex AI pipelines. Serverless deployment; pricing based on storage and serving requests. Natural choice for teams already in Google Cloud with BigQuery-based data warehouse.
AWS SageMaker Feature Store — native AWS integration with S3, DynamoDB, and SageMaker training infrastructure. The path-of-least-resistance choice for AWS-native organizations.
Selection guidance: start with Feast if cost is a primary constraint and the team has engineering capacity to operate it. Evaluate Tecton or cloud-native options when feature store downtime has measurable business impact.
ML Training Data Pipeline Design
Training data pipelines for ML have requirements that differ from standard analytical pipelines. The key differences:
Point-in-time correctness — historical training data must reflect what was known at each point in time, not the current state of the world. Customer churn models trained on "current" customer attributes instead of "attributes at the time of cancellation" produce models that cannot generalize.
Train-test temporal splits — training and test data must be split by time, not randomly. Random splits leak future patterns into training data. Temporal splits ensure the model is evaluated on data it could not have seen during training.
Reproducible dataset composition — the exact set of records in each training run must be versioned so it can be reproduced. If model performance changes, the team needs to determine whether data changed, code changed, or both.
Feature freshness requirements — some features update daily; others update every 15 minutes. The training pipeline must handle features with different update cadences without inadvertently using stale values for fast-updating features.
In production ML systems at Smart Maple — including a scheduling optimization system that processes multi-dimensional demand patterns for healthcare facilities — we found that the transition from notebook experiments to production models consistently revealed training data pipeline requirements that were invisible in offline evaluation. Temporal leakage and training-serving skew were the two most common causes of models that passed offline metrics but degraded within 60 days of deployment.
Data Versioning with DVC
Data Version Control (DVC) brings Git-like versioning discipline to ML datasets and model artifacts. Datasets stored in S3, GCS, or Azure Blob are tracked by DVC using content-addressed hashing — a DVC pointer file in Git references the exact dataset snapshot by content hash, not by path.
What DVC enables:
- Reproduce any historical training run by checking out the corresponding Git commit (which includes the DVC pointer to the exact dataset version used)
- Track dataset lineage: which feature extraction scripts produced which dataset, which transformations were applied, and what the resulting model performance was
- Branch and merge dataset versions like code: create a dataset variation for an experiment without overwriting the production training set
- Cache expensive computations: DVC pipelines store intermediate outputs so reruns only recompute changed stages
DVC pipeline integration with ML workflows:
A DVC pipeline defines stages: data extraction → feature engineering → training → evaluation. Each stage has defined inputs, outputs, and a command. DVC tracks checksums of all inputs and outputs; unchanged stages are skipped on rerun. The entire pipeline, from raw data to trained model, is reproducible from a single dvc repro command given a specific Git commit.
MLflow for experiment tracking — while DVC tracks data and pipeline reproducibility, MLflow tracks model training experiments: hyperparameters tried, metrics achieved, model artifacts produced. Together, DVC + MLflow provide full lineage from raw data through feature computation to trained model artifact.
Data Drift Detection
Data drift is the statistical divergence between the data distribution a model was trained on and the data distribution it encounters in production. All production ML models experience drift; the question is how fast and how severe.
Types of drift relevant to data engineering:
Covariate shift — the distribution of input feature values changes. The transaction_hour feature, which ranged from 9 AM–5 PM during training, now includes significant evening and weekend transactions due to changed user behavior. The model's feature importance weights were calibrated on the old distribution.
Schema drift — upstream data source changes field names, data types, or allowed values. The payment_method field adds a new value ("BNPL") that was absent during training. The model's encoding for this field produces an out-of-distribution input.
Concept drift — the relationship between features and the target changes. The patterns that predicted high-value customers in 2023 no longer predict high-value customers in 2026 because market conditions changed. This is the hardest drift type to detect because the feature distributions may look healthy while the predictive relationship degrades.
Detection tooling:
- Evidently AI — open-source Python library that generates drift detection reports comparing reference distributions to current distributions across all features. Integrates with Airflow as a pipeline step.
- WhyLabs — managed monitoring platform that logs feature statistics from inference requests and alerts on statistical divergence from baseline.
- Arize AI — ML observability platform with drift detection, prediction monitoring, and explainability analysis.
Data engineering responsibility for drift: The data engineering team owns the infrastructure that collects feature statistics during training (baseline) and during inference (current). Even if a data science team operates the drift detection tooling, the data engineering pipeline must be instrumented to emit the right metrics.
ML Pipeline Checklist for Data Engineers
Before a model transitions from development to production, the data engineering team should verify:
Feature consistency
- Offline and online store compute features with identical logic
- Point-in-time correct training dataset created (no temporal leakage)
- Feature definitions are versioned and documented in the feature registry
Data versioning
- Training dataset version is recorded in model metadata
- DVC or equivalent tool tracks dataset-to-model lineage
- Retraining pipeline can reproduce any historical training run from Git commit
Drift infrastructure
- Feature distribution baseline captured from training data
- Inference pipeline logs feature statistics for drift comparison
- Drift alert thresholds defined and monitoring enabled
Pipeline reliability
- Retraining pipeline runs on schedule, not ad-hoc
- Failed retraining runs alert data engineering team within 15 minutes
- Rollback procedure defined (which model version to reactivate if new model degrades)
Data Engineering vs. ML Engineering: Boundary and Collaboration
The data engineering MLOps integration challenge is partly organizational: who builds what, and who is responsible when the model degrades?
Data engineering owns:
- Feature pipeline implementation and reliability
- Feature store infrastructure and SLA
- Training dataset construction and versioning
- Data drift detection instrumentation
- Feature freshness monitoring
ML engineering owns:
- Model architecture and training implementation
- Hyperparameter optimization
- Offline model evaluation
- Inference service deployment
- Model performance monitoring
Shared ownership:
- Feature definition (data engineering implements; ML engineering specifies)
- Drift response (ML engineering diagnoses; data engineering fixes the pipeline)
- Retraining trigger criteria (joint decision on when degradation justifies retraining cost)
The collaboration model that works: ML engineers file feature requests as specifications ("I need 30-day rolling purchase count per customer, updated daily"); data engineers implement and maintain the feature pipeline and store the result in the feature registry; ML engineers retrieve features from the registry without touching the pipeline. The interface between the two roles is the feature registry.
Conclusion
Data engineering MLOps integration is the unglamorous foundation that makes ML systems reliable in production. Feature stores prevent training-serving skew. Data versioning makes training runs reproducible. Drift detection surfaces model degradation before users notice. The toolchain — Feast, DVC, MLflow, Evidently — is mature and well-documented. Teams working with large language models should also consider LLM fine-tuning to adapt foundation models on domain-specific datasets — a complementary process that benefits from the same data versioning and pipeline discipline described here.
The organizational challenge is as important as the technical one: data engineering teams must understand what ML systems need from data infrastructure, and ML teams must communicate those needs as clear specifications. When the data-ML interface is well-designed, production ML systems are stable, debuggable, and improvable. When it is not, the 87% failure-to-production statistic is the predictable outcome.
Related Articles
MLOps Guide: Taking Machine Learning Models to Production [2026]
87% of machine learning models built by data science teams never reach production. The models work — they pass cross-validation, they score well on holdout sets, they demonstrate genuine predictive value. The problem is not the modeling. The problem is everything that happens between a notebook experiment and a reliable, monitored, production system. MLOps is the discipline that closes that gap. This guide covers the full MLOps stack: maturity levels, tooling choices (MLflow, DVC, Kubeflow
Read MoreLLM Fine-Tuning Guide: Custom Model Training with LoRA and QLoRA [2026]
General-purpose LLMs are impressive. They can write code, summarize documents, answer questions, and translate between languages with reasonable accuracy. But "reasonable" is not good enough when your application requires consistent output format, domain-specific terminology, a particular tone, or behavior that the base model was never trained to exhibit. That gap is where fine-tuning matters. Fine-tuning updates a model's weights on your specific data, changing how the model behaves — not
Read MoreComputer Vision Applications: Object Detection, OCR, and Industrial AI [2026]
Computer vision has moved well past the research phase. The models are trained, the frameworks are mature, the hardware is accessible, and the use cases are generating measurable returns. What was a specialized capability requiring deep expertise in 2018 is now deployable infrastructure — if you know which component to reach for and where the real complexity lives. This guide covers computer vision applications across industrial, medical, logistics, and document processing domains. It expl
Read More
