🚢 Model Deployment: From Development to Production
📚 What You'll Learn
By the end of this lesson, you will be able to:
- Compare deployment patterns — REST APIs, batch, and streaming serving
- Wrap a trained model in a prediction API using Flask and FastAPI
- Containerize a model service with Docker for reproducible deployment
- Choose serving strategies such as blue-green and canary releases
- Load-test an endpoint and reason about latency and throughput
- Apply a production deployment checklist covering validation, versioning, and rollback
⏱️ Estimated Time: 45–60 minutes
🎯 Project: Take a saved model, expose it through a FastAPI endpoint, containerize it with Docker, and load-test the running service to measure its latency.
Introduction
Model deployment is the process of making trained machine learning models available for use in production environments. This involves creating APIs, containerization, scaling strategies, and monitoring systems. This lesson covers different deployment patterns including REST APIs, batch processing, streaming, and edge deployment, along with practical implementations using Flask, FastAPI, Docker, and cloud platforms.
Deployment Patterns Overview
Deployment Architecture
Setting Up Deployment Environment
import os
import json
import pickle
import joblib
import numpy as np
import pandas as pd
from datetime import datetime
import logging
# Setup logging
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)
# Create deployment directories
deployment_dirs = ['models', 'configs', 'logs', 'tests']
for dir_name in deployment_dirs:
os.makedirs(dir_name, exist_ok=True)
print("Deployment Environment Setup Complete")
print("=" * 50)
REST API Deployment with Flask
# flask_app.py
from flask import Flask, request, jsonify
import joblib
import numpy as np
import pandas as pd
import traceback
import logging
from datetime import datetime
# Initialize Flask app
app = Flask(__name__)
# Configure logging
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)
# Global model variable
MODEL = None
MODEL_METADATA = {}
class ModelService:
"""Model service for handling predictions"""
def __init__(self, model_path):
self.model_path = model_path
self.model = None
self.metadata = {}
self.load_model()
def load_model(self):
"""Load model and metadata"""
try:
self.model = joblib.load(self.model_path)
self.metadata = {
'model_path': self.model_path,
'loaded_at': datetime.now().isoformat(),
'model_type': type(self.model).__name__
}
logger.info(f"Model loaded successfully from {self.model_path}")
except Exception as e:
logger.error(f"Error loading model: {str(e)}")
raise
def predict(self, data):
"""Make predictions"""
try:
# Convert to DataFrame if needed
if isinstance(data, dict):
df = pd.DataFrame([data])
elif isinstance(data, list):
df = pd.DataFrame(data)
else:
df = data
# Make prediction
prediction = self.model.predict(df)
# Get prediction probability if available
proba = None
if hasattr(self.model, 'predict_proba'):
proba = self.model.predict_proba(df).tolist()
return {
'prediction': prediction.tolist(),
'probability': proba
}
except Exception as e:
logger.error(f"Prediction error: {str(e)}")
raise
def get_model_info(self):
"""Get model information"""
return self.metadata
# Initialize model service (example path)
# In production, load from environment variable or config
# model_service = ModelService('models/model.pkl')
# Flask Routes
@app.route('/health', methods=['GET'])
def health_check():
"""Health check endpoint"""
return jsonify({
'status': 'healthy',
'timestamp': datetime.now().isoformat()
})
@app.route('/predict', methods=['POST'])
def predict():
"""Prediction endpoint"""
try:
# Get data from request
data = request.json
# Log request
logger.info(f"Prediction request received: {len(data)} samples")
# Make prediction
# result = model_service.predict(data)
# For demo, return mock prediction
result = {
'prediction': [1],
'probability': [[0.3, 0.7]],
'model_version': '1.0.0',
'timestamp': datetime.now().isoformat()
}
return jsonify(result)
except Exception as e:
error_msg = f"Prediction error: {str(e)}\n{traceback.format_exc()}"
logger.error(error_msg)
return jsonify({'error': str(e)}), 500
@app.route('/model/info', methods=['GET'])
def model_info():
"""Get model information"""
try:
# info = model_service.get_model_info()
info = {
'model_type': 'RandomForestClassifier',
'version': '1.0.0',
'features': ['feature_1', 'feature_2', 'feature_3'],
'last_updated': datetime.now().isoformat()
}
return jsonify(info)
except Exception as e:
return jsonify({'error': str(e)}), 500
# Example Flask app code
print("\n" + "=" * 50)
print("FLASK API EXAMPLE")
print("=" * 50)
print("""
To run the Flask app:
1. Save the code to 'app.py'
2. Install Flask: pip install flask
3. Run: python app.py
4. Test: curl -X POST http://localhost:5000/predict -H "Content-Type: application/json" -d '{"feature1": 1.0}'
""")
FastAPI Deployment (Modern Alternative)
# fastapi_app.py
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
from typing import List, Optional
import joblib
import numpy as np
import pandas as pd
from datetime import datetime
import logging
# Create FastAPI app
app = FastAPI(
title="ML Model API",
description="Production ML model serving with FastAPI",
version="1.0.0"
)
# Configure logging
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)
# Pydantic models for request/response validation
class PredictionRequest(BaseModel):
"""Request model for predictions"""
features: List[float]
feature_names: Optional[List[str]] = None
class Config:
schema_extra = {
"example": {
"features": [5.1, 3.5, 1.4, 0.2],
"feature_names": ["sepal_length", "sepal_width", "petal_length", "petal_width"]
}
}
class PredictionResponse(BaseModel):
"""Response model for predictions"""
prediction: List[int]
probability: Optional[List[List[float]]]
model_version: str
timestamp: str
class ModelInfo(BaseModel):
"""Model information response"""
model_type: str
version: str
features: List[str]
performance_metrics: dict
last_updated: str
# API Endpoints
@app.get("/")
async def root():
"""Root endpoint"""
return {
"message": "ML Model API",
"version": "1.0.0",
"endpoints": ["/docs", "/health", "/predict", "/model/info"]
}
@app.get("/health")
async def health_check():
"""Health check endpoint"""
return {
"status": "healthy",
"timestamp": datetime.now().isoformat()
}
@app.post("/predict", response_model=PredictionResponse)
async def predict(request: PredictionRequest):
"""
Make predictions with the model
"""
try:
# Prepare data
features = np.array(request.features).reshape(1, -1)
# Make prediction (mock for demo)
prediction = [1]
probability = [[0.3, 0.7]]
return PredictionResponse(
prediction=prediction,
probability=probability,
model_version="1.0.0",
timestamp=datetime.now().isoformat()
)
except Exception as e:
logger.error(f"Prediction error: {str(e)}")
raise HTTPException(status_code=500, detail=str(e))
@app.get("/model/info", response_model=ModelInfo)
async def get_model_info():
"""
Get model information and metadata
"""
return ModelInfo(
model_type="RandomForestClassifier",
version="1.0.0",
features=["feature_1", "feature_2", "feature_3", "feature_4"],
performance_metrics={
"accuracy": 0.95,
"precision": 0.94,
"recall": 0.96,
"f1_score": 0.95
},
last_updated=datetime.now().isoformat()
)
@app.post("/batch_predict")
async def batch_predict(requests: List[PredictionRequest]):
"""
Batch prediction endpoint
"""
results = []
for req in requests:
result = await predict(req)
results.append(result)
return results
print("\n" + "=" * 50)
print("FASTAPI DEPLOYMENT EXAMPLE")
print("=" * 50)
print("""
FastAPI advantages:
1. Automatic API documentation (Swagger UI)
2. Type validation with Pydantic
3. Async support for better performance
4. Modern Python features
To run:
1. Install: pip install fastapi uvicorn
2. Run: uvicorn app:app --reload
3. View docs: http://localhost:8000/docs
""")
Docker Containerization
# Dockerfile
FROM python:3.9-slim
# Set working directory
WORKDIR /app
# Copy requirements first for better caching
COPY requirements.txt .
# Install dependencies
RUN pip install --no-cache-dir -r requirements.txt
# Copy application code
COPY app.py .
COPY models/ ./models/
COPY configs/ ./configs/
# Create non-root user for security
RUN useradd -m -u 1000 mluser && chown -R mluser:mluser /app
USER mluser
# Expose port
EXPOSE 8000
# Health check
HEALTHCHECK --interval=30s --timeout=3s --start-period=5s --retries=3 \
CMD python -c "import requests; requests.get('http://localhost:8000/health')" || exit 1
# Run application
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]
# docker-compose.yml
version: '3.8'
services:
model-api:
build: .
container_name: ml-model-api
ports:
- "8000:8000"
environment:
- MODEL_PATH=/app/models/model.pkl
- LOG_LEVEL=INFO
- MAX_BATCH_SIZE=100
volumes:
- ./models:/app/models:ro
- ./logs:/app/logs
restart: unless-stopped
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 40s
nginx:
image: nginx:alpine
container_name: ml-nginx
ports:
- "80:80"
volumes:
- ./nginx.conf:/etc/nginx/nginx.conf:ro
depends_on:
- model-api
restart: unless-stopped
prometheus:
image: prom/prometheus
container_name: ml-prometheus
ports:
- "9090:9090"
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
restart: unless-stopped
print("\n" + "=" * 50)
print("DOCKER DEPLOYMENT")
print("=" * 50)
print("""
Docker commands:
1. Build image: docker build -t ml-model:latest .
2. Run container: docker run -p 8000:8000 ml-model:latest
3. Docker Compose: docker-compose up -d
4. View logs: docker logs ml-model-api
5. Stop: docker-compose down
""")
Model Serving Patterns
class ModelServer:
"""
Advanced model serving with caching and monitoring
"""
def __init__(self, model_path, cache_size=1000):
self.model = joblib.load(model_path)
self.cache = {}
self.cache_size = cache_size
self.metrics = {
'total_predictions': 0,
'cache_hits': 0,
'cache_misses': 0,
'avg_latency': 0,
'errors': 0
}
def _get_cache_key(self, features):
"""Generate cache key from features"""
return hash(tuple(features.flatten()))
def predict_with_cache(self, features):
"""Predict with caching"""
import time
start_time = time.time()
cache_key = self._get_cache_key(features)
# Check cache
if cache_key in self.cache:
self.metrics['cache_hits'] += 1
logger.info("Cache hit")
return self.cache[cache_key]
# Make prediction
try:
prediction = self.model.predict(features)
# Update cache (FIFO if full)
if len(self.cache) >= self.cache_size:
# Remove oldest entry
oldest = next(iter(self.cache))
del self.cache[oldest]
self.cache[cache_key] = prediction
self.metrics['cache_misses'] += 1
# Update metrics
latency = time.time() - start_time
self.metrics['total_predictions'] += 1
self.metrics['avg_latency'] = (
(self.metrics['avg_latency'] * (self.metrics['total_predictions'] - 1) + latency)
/ self.metrics['total_predictions']
)
return prediction
except Exception as e:
self.metrics['errors'] += 1
logger.error(f"Prediction error: {str(e)}")
raise
def get_metrics(self):
"""Get server metrics"""
cache_hit_rate = (
self.metrics['cache_hits'] /
max(self.metrics['cache_hits'] + self.metrics['cache_misses'], 1)
)
return {
**self.metrics,
'cache_hit_rate': cache_hit_rate,
'cache_size': len(self.cache)
}
def clear_cache(self):
"""Clear prediction cache"""
self.cache.clear()
logger.info("Cache cleared")
# Batch processing
class BatchProcessor:
"""
Batch processing for offline scoring
"""
def __init__(self, model_path, batch_size=1000):
self.model = joblib.load(model_path)
self.batch_size = batch_size
def process_file(self, input_file, output_file):
"""Process predictions from file"""
# Read data
data = pd.read_csv(input_file)
# Process in batches
predictions = []
for i in range(0, len(data), self.batch_size):
batch = data.iloc[i:i + self.batch_size]
batch_pred = self.model.predict(batch)
predictions.extend(batch_pred)
# Save results
data['prediction'] = predictions
data.to_csv(output_file, index=False)
logger.info(f"Processed {len(data)} records")
return len(data)
def process_stream(self, data_generator):
"""Process streaming data"""
for batch in data_generator:
predictions = self.model.predict(batch)
yield predictions
# Demonstrate serving patterns
print("\n" + "=" * 50)
print("MODEL SERVING PATTERNS")
print("=" * 50)
# Create sample model (mock)
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import make_classification
X, y = make_classification(n_samples=100, n_features=4, random_state=42)
model = RandomForestClassifier(n_estimators=10, random_state=42)
model.fit(X, y)
# Save model
joblib.dump(model, 'models/demo_model.pkl')
# Test server with caching
server = ModelServer('models/demo_model.pkl', cache_size=10)
# Make predictions
test_features = X[:5]
for i, features in enumerate(test_features):
pred = server.predict_with_cache(features.reshape(1, -1))
print(f"Prediction {i+1}: {pred[0]}")
# Make same predictions again (should hit cache)
print("\nRepeating predictions (should hit cache):")
for i, features in enumerate(test_features[:3]):
pred = server.predict_with_cache(features.reshape(1, -1))
print(f"Cached prediction {i+1}: {pred[0]}")
# Show metrics
print("\nServer Metrics:")
metrics = server.get_metrics()
for key, value in metrics.items():
print(f" {key}: {value}")
Load Testing and Performance
# load_test.py
import asyncio
import aiohttp
import time
import statistics
from typing import List
class LoadTester:
"""
Load testing for deployed models
"""
def __init__(self, base_url="http://localhost:8000"):
self.base_url = base_url
self.results = []
async def make_request(self, session, data):
"""Make single async request"""
start_time = time.time()
try:
async with session.post(
f"{self.base_url}/predict",
json=data
) as response:
result = await response.json()
latency = time.time() - start_time
return {
'status': response.status,
'latency': latency,
'success': response.status == 200
}
except Exception as e:
return {
'status': 0,
'latency': time.time() - start_time,
'success': False,
'error': str(e)
}
async def run_load_test(self, num_requests=100, concurrent=10):
"""Run load test with concurrent requests"""
print(f"\nRunning load test: {num_requests} requests, {concurrent} concurrent")
# Prepare test data
test_data = {'features': [1.0, 2.0, 3.0, 4.0]}
# Create semaphore for concurrency control
semaphore = asyncio.Semaphore(concurrent)
async def bounded_request(session):
async with semaphore:
return await self.make_request(session, test_data)
# Run requests
async with aiohttp.ClientSession() as session:
start_time = time.time()
tasks = [bounded_request(session) for _ in range(num_requests)]
self.results = await asyncio.gather(*tasks)
total_time = time.time() - start_time
# Calculate statistics
successful = [r for r in self.results if r['success']]
latencies = [r['latency'] for r in successful]
stats = {
'total_requests': num_requests,
'successful_requests': len(successful),
'failed_requests': num_requests - len(successful),
'success_rate': len(successful) / num_requests * 100,
'total_time': total_time,
'requests_per_second': num_requests / total_time,
'avg_latency': statistics.mean(latencies) if latencies else 0,
'min_latency': min(latencies) if latencies else 0,
'max_latency': max(latencies) if latencies else 0,
'p50_latency': statistics.median(latencies) if latencies else 0,
'p95_latency': statistics.quantiles(latencies, n=20)[18] if len(latencies) > 20 else 0,
'p99_latency': statistics.quantiles(latencies, n=100)[98] if len(latencies) > 100 else 0
}
return stats
def print_results(self, stats):
"""Print load test results"""
print("\n" + "=" * 50)
print("LOAD TEST RESULTS")
print("=" * 50)
print(f"Total Requests: {stats['total_requests']}")
print(f"Successful: {stats['successful_requests']} ({stats['success_rate']:.1f}%)")
print(f"Failed: {stats['failed_requests']}")
print(f"Total Time: {stats['total_time']:.2f}s")
print(f"Requests/Second: {stats['requests_per_second']:.1f}")
print("\nLatency Statistics (seconds):")
print(f" Average: {stats['avg_latency']:.3f}")
print(f" Min: {stats['min_latency']:.3f}")
print(f" Max: {stats['max_latency']:.3f}")
print(f" P50: {stats['p50_latency']:.3f}")
print(f" P95: {stats['p95_latency']:.3f}")
print(f" P99: {stats['p99_latency']:.3f}")
# Example load test
print("\n" + "=" * 50)
print("LOAD TESTING EXAMPLE")
print("=" * 50)
print("""
To run load test:
1. Start your model API server
2. Run the load test:
import asyncio
tester = LoadTester()
stats = asyncio.run(tester.run_load_test(num_requests=1000, concurrent=50))
tester.print_results(stats)
""")
Production Deployment Checklist
✅ Deployment Checklist
- Model: Versioned, tested, documented
- API: RESTful endpoints, input validation, error handling
- Container: Docker image, health checks, resource limits
- Security: Authentication, HTTPS, input sanitization
- Monitoring: Metrics, logging, alerting
- Scaling: Load balancer, auto-scaling, caching
- Testing: Unit tests, integration tests, load tests
- Documentation: API docs, deployment guide, runbook
- Rollback: Version control, quick rollback strategy
- Compliance: Data privacy, audit logs, regulations
Best Practices
🎯 Deployment Best Practices
- Blue-Green Deployment: Zero-downtime deployments
- Canary Releases: Gradual rollout to subset of users
- Feature Flags: Enable/disable features without deployment
- Circuit Breakers: Prevent cascade failures
- Rate Limiting: Protect against abuse
- Request Validation: Validate all inputs
- Async Processing: Use queues for heavy computations
- Model Warm-up: Pre-load models before serving
- Graceful Shutdown: Handle in-flight requests
- Observability: Structured logging, distributed tracing
Practice Exercises
Exercise 1: Build Production API
Create a production-ready model API:
- Implement FastAPI with proper validation
- Add authentication and rate limiting
- Implement caching layer
- Add comprehensive logging
- Create Docker container
Exercise 2: Deploy to Cloud
Deploy model to cloud platform:
- Choose cloud provider (AWS/GCP/Azure)
- Set up container registry
- Configure auto-scaling
- Implement monitoring dashboard
- Set up CI/CD pipeline
Summary
✅ You've Learned
- Different deployment patterns and when to use them
- Building REST APIs with Flask and FastAPI
- Docker containerization for ML models
- Model serving with caching and optimization
- Load testing and performance monitoring
- Production deployment best practices
- Security and scaling considerations
📓 Learning Journal
Keep a learning journal — digital or physical. After this lesson, take a few minutes to write down:
- Key concepts you learned
- Techniques that clicked for you
- Questions or confusion points to revisit
- Ideas you want to try
- Your progress and feelings about learning this
✍️ This lesson's prompt: A model in a notebook helps no one until it's deployed. Which part of taking a model to production — the API, the container, the release strategy — felt most unfamiliar, and how would you build confidence in it before real users depend on it?
📝 Lesson Summary
🎓 Key Takeaways
- Deployment turns a saved model into a service others can call — most commonly a REST API for real-time predictions, or batch/streaming jobs.
- Flask is simple and familiar; FastAPI adds async performance, type validation, and automatic docs.
- Docker packages the model, code, and dependencies together so it runs identically everywhere.
- Safe rollouts (blue-green, canary), load testing, and a deployment checklist protect users from bad releases.
🎉 What You've Accomplished
You can now take a trained model out of the notebook and stand it up as a containerized, testable prediction service — and choose a rollout strategy that lets you ship changes without risking downtime.
❓ Common Questions at This Stage
Flask or FastAPI for a model API?
FastAPI is the modern default: async support, request validation via type hints, and auto-generated docs. Flask is fine for quick prototypes or when you already have a Flask stack.
Why bother with Docker?
It eliminates "works on my machine" problems by bundling the exact Python version, libraries, and model artifact. The same image runs on your laptop, CI, and the cloud unchanged.
What's the difference between blue-green and canary deployment?
Blue-green keeps two full environments and switches all traffic at once (easy instant rollback). Canary sends a small slice of traffic to the new version first, expanding only if metrics stay healthy.
🔭 Looking Ahead
Shipping the model is only the beginning — next comes watching it in production, catching drift and degradation before they hurt users.
✅ Before the Next Lesson
- Serve a saved model behind a FastAPI endpoint and call it with a sample request.
- Write a Dockerfile for the service, build the image, and run the container locally.
- Write your Learning Journal entry for this lesson
🌟 Encouragement for the Journey
Crossing the gap from "it works in training" to "it serves real requests" is where many models stall — and you just built that bridge. This is the skill that makes your ML work actually count. Well done!