Model Deployment
Deploying a model means making it available to make predictions on new data in a production environment.
Definition
A deployed model takes features as input and returns a prediction — via REST API, batch prediction job, or embedded in an application.
Key requirements: low latency, high availability, monitoring for data drift and model staleness.
Intuition
A model that can't be deployed is just an experiment. Deployment exposes the gap between notebooks and production systems.
Data drift (input distribution changes) and concept drift (relationship between and changes) degrade models over time.
Worked example
A fraud model called via API: receive transaction features, run model, return fraud score in <100ms at 99.9% uptime.
Batch scoring: nightly job reads today's transactions, outputs risk scores for review the next morning.
The math
Model serialization: pickle (Python), ONNX (cross-platform), or framework-specific formats (TensorFlow SavedModel, PyTorch jit).
A/B testing a new model: route 10% of traffic to the new model and compare metrics before full rollout.
In practice
Monitor prediction distributions — a sudden shift in average fraud score may indicate a new fraud pattern.
Shadow mode: new model runs alongside old one without affecting responses, building confidence before switching.
Go deeper
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.