MLOps: The Bridge Between Research and Production

Essay
MLOps
MLOps is the operational discipline that decides whether an organization’s machine learning ever leaves the notebook and delivers value.
Published

January 15, 2026

The 85% Problem

Most machine learning projects fail. The algorithms usually work. What fails is the organization’s ability to run them in production. One 2024 industry analysis found that 85% of ML initiatives never deliver business value in production, often because of silent failures, data drift, and broken pipelines that nobody catches in time (Galileo). That is a hard number for a technology so many companies are betting on.

The pattern is familiar to anyone who has worked in data science. A team spends months developing a model. The metrics look great and the stakeholders are excited. Six months later everyone has quietly moved on, and the model is still sitting in a Jupyter notebook.

The missing piece is usually infrastructure. The discipline that supplies it, MLOps, now separates organizations that get value from ML from those that don’t.

What Makes ML Different

You could call MLOps DevOps for machine learning, but that undersells the problem. Traditional software is, in a sense, honest. You write code, test it, deploy it, and it does what you told it to do. Machine learning systems are shaped by data in ways that make them unpredictable, even to the people who built them.

The same code can produce very different outputs as data distributions shift. A model that worked yesterday can fail tomorrow with no change to the codebase. Performance degrades as the real world drifts away from the world in the training data. Random seeds, hardware, and dependency versions make reproducibility a constant struggle. In regulated industries you also need explainability and audit trails that traditional software never demanded.

The best known paper on this subject, “Hidden Technical Debt in Machine Learning Systems,” showed that the ML code is only a small fraction of a production system. The rest is data collection, feature extraction, configuration, serving infrastructure, and monitoring. The model is the easy part (Sculley et al., 2015).

Uber ran into this with its arrival time predictions. Traffic behaves differently from city to city and changes over the year, so a model that works in one place and season will drift in another. Uber’s answer was automated pipelines that monitor accuracy and trigger retraining without human intervention. Making that work took feature stores, model registries, and testing infrastructure on the scale of any serious software platform (Uber Michelangelo, DeepETA).

Why It Matters Now

For years, machine learning lived in the back office. It powered internal analytics and batch jobs that ran overnight. That era is over. Chatbots now sit in everyday products. Voice assistants respond in real time (OpenAI Realtime API). Recommendation systems shape what people watch and buy. Search engines have started generating answers instead of pointing to them.

Moving machine learning from analytics into the product changes the requirements. A notebook model can take hours to run. A consumer product has to respond in milliseconds, and it has to keep adapting without going dark.

Netflix processes billions of interactions while retraining the models behind its recommendations. That means coordinating hundreds of models across many services, a level of orchestration that would have seemed absurd a decade ago (Netflix Maestro, Netflix foundation model).

Large language models make all of this more urgent. They power chatbots, coding assistants, and creative tools. They also bring failure modes we are only beginning to understand, such as hallucinations, prompt injection, and answers that sound confident and are wrong. LLMOps, as practitioners now call it, is less about training from scratch and more about prompting, fine tuning, and deploying responsibly.

Regulators have noticed. The EU AI Act brings governance and audit requirements for high risk AI systems. Meanwhile the silent failures behind that 85% figure are not silent in consumer products. Users notice, and revenue suffers.

What You Actually Need

A startup with one model has different needs than a bank running hundreds of risk models. Still, certain components show up everywhere.

Version control has to extend beyond code. Models, data, configurations, and experiments all need versioning, because the relationships between them determine how a model behaves. Tools like DVC, MLflow, and Weights & Biases track the lineage from raw data to deployment. You need that lineage when someone asks, months later, why a model made a particular decision.

Reproducible environments matter more than people expect. A model can produce different outputs after a framework upgrade, even with identical weights. Containers and pinned dependencies are the only reliable defense.

Testing has to happen at several levels because ML systems fail in ways traditional software doesn’t. You need unit tests for data processing, integration tests for pipelines, validation tests for data quality and schema drift, performance tests against accuracy thresholds, and infrastructure tests for deployment configuration.

Monitoring covers three areas. System monitoring tracks CPU, memory, and latency. Data monitoring watches feature distributions, missing values, and schema changes. Model monitoring tracks prediction distributions, accuracy over time, and drift. Tools like Evidently AI and WhyLabs help, though custom dashboards on Prometheus and Grafana remain common.

Airbnb’s pricing work shows why individual predictions are not enough. A price suggestion can look reasonable on its own and still be wrong in aggregate. Judging those models meant watching aggregate patterns and business metrics as well as individual outputs (Airbnb pricing).

A model registry holds metadata, performance metrics, and approval workflows in one place, so you can roll back when something goes wrong. Feature stores stop teams from reimplementing the same feature logic across projects and introducing inconsistencies that are maddening to debug. Uber’s Palette serves features for both training and real time inference, so what the model learned matches what it sees in production (Uber Palette).

The Hard Part

The hardest part of MLOps is organizational.

Data scientists, ML engineers, DevOps engineers, and product managers have to work together in ways that traditional structures discourage. Silos that took years to build have to come down. The culture has to move away from “move fast and break things” toward something closer to moving deliberately and watching everything.

MLOps needs a blend of data science, software engineering, and infrastructure skills that few individuals have. Some organizations respond by creating dedicated ML platform teams. Platform engineers own the tooling and operational complexity, and data scientists get to focus on modeling.

The deepest challenge may be defining success. Uptime and latency don’t capture what matters for ML systems. A recommendation engine can raise click through rates while trapping users in filter bubbles. A fraud detection system can score high on accuracy while flagging so many legitimate transactions that customers leave. These tradeoffs need judgment that spans technology and the business. Organizations that can’t make those calls well build systems that optimize for the wrong things.

Getting Started

If you are starting from scratch, don’t try to build everything at once. Version your models and data first. Even basic versioning pays off and lays the foundation for everything else. Add simple monitoring next, covering prediction distributions, latency, and error rates. Then pick one manual process, such as deployment, retraining, or testing, and automate it. Learn from that before expanding.

The organizations that succeed start with their biggest pain points rather than an idealized architecture. MLOps is iterative. You build what you need, find new problems, and build again.

Takeaways

MLOps is no longer optional. As machine learning moves from research curiosity to production infrastructure, organizations that invest in operations will pull ahead. Sound operations are also becoming a prerequisite for responsible deployment.

The models we build now influence loan approvals, medical diagnoses, hiring decisions, and resource allocation at a scale that would have been unimaginable a decade ago. Deploying them without proper monitoring and governance is reckless.

Better models matter. What matters more is whether they keep working, openly and fairly, once they leave the notebook. MLOps is the bridge that gets them there.

References

  • Sculley, D., et al. (2015). “Hidden Technical Debt in Machine Learning Systems.” Advances in Neural Information Processing Systems, 28. Paper

  • Netflix Engineering. (2021). “RecSysOps: Best Practices for Operating a Large-Scale Recommender System.” Netflix Tech Blog. Article

  • Netflix Engineering. (2023). “Foundation Model for Personalized Recommendation.” Netflix Tech Blog. Article

  • Netflix Engineering. (2020). “Maestro: Netflix’s Workflow Orchestrator.” Netflix Tech Blog. Article

  • Uber Engineering. (2017). “Meet Michelangelo: Uber’s Machine Learning Platform.” Uber Engineering Blog. Article

  • Uber Engineering. (2022). “DeepETA: How Uber Predicts Arrival Times Using Deep Learning.” Uber Blog. Article

  • Uber Engineering. (2021). “Palette Meta Store Journey.” Uber Blog. Article

  • Airbnb Engineering. (2018). “Customized Regression Model for Airbnb Dynamic Pricing.” KDD 2018. Paper

  • Zalando Research. (2017). “Practical Lessons from Developing a Large-Scale Recommender System at Zalando.” RecSys 2017. Paper

  • Stitch Fix Engineering. (2023). “Accelerating AI: Implementing Multi-GPU Distributed Training for Personalized Recommendations.” Multithreaded Blog. Article

  • McKinsey Digital. (2024). “MLOps so AI can scale.” McKinsey & Company. Article

  • Google Cloud. (2024). “How L’Oréal’s tech accelerator built its end-to-end MLOps platform.” Google Cloud Blog. Article

  • Galileo AI. (2024). “A Guide to Reliable ML Observability to Prevent 85% of Production Failures.” Galileo AI Blog. Article

  • OpenAI. (2025). “Introducing gpt-realtime and Realtime API updates for production voice agents.” OpenAI Blog. Article

Tools and Documentation

Also covered
Machine Learning EngineeringData ScienceModel MonitoringFeature Stores