🔍 Anomaly Detection: Finding the Unusual
📚 What You'll Learn
By the end of this lesson, you will be able to:
- Distinguish point, contextual, and collective anomalies and give real-world examples of each
- Apply statistical methods (z-score, IQR, and distribution-based tests) to flag outliers
- Use machine-learning detectors such as Isolation Forest, One-Class SVM, and Local Outlier Factor
- Build an end-to-end anomaly detection pipeline and evaluate it on imbalanced data
- Detect anomalies in time-series data and interpret the results for a stakeholder
⏱️ Estimated Time: 45–60 minutes
🎯 Project: Build an anomaly detection pipeline that flags fraudulent or unusual records and compare statistical vs machine-learning detectors.
Introduction
Anomaly detection is the identification of rare items, events, or observations that deviate significantly from the majority of the data. These anomalies, also called outliers, can indicate critical incidents such as fraud, system failures, or defects. This lesson covers various anomaly detection techniques from statistical methods to machine learning approaches, with practical applications in fraud detection, network security, and quality control.
Anomaly Detection Overview
Types of Anomalies
Deviates from Normal] C --> C1[Normal in Different Context
Abnormal in Current] D --> D1[Group of Instances
Collectively Abnormal] B1 --> E[Example: Fraud Transaction] C1 --> F[Example: Hot Winter Day] D1 --> G[Example: Cyber Attack Pattern] style B fill:#ffebee style C fill:#fff3e0 style D fill:#e8f5e9
Setting Up the Environment
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report, confusion_matrix
import warnings
warnings.filterwarnings('ignore')
# Set style for better visualizations
plt.style.use('seaborn-v0_8-darkgrid')
sns.set_palette("husl")
# Set random seed for reproducibility
np.random.seed(42)
print("Anomaly Detection Methods Comparison")
print("=" * 50)
Statistical Methods
class StatisticalAnomalyDetector:
"""
Statistical methods for anomaly detection
"""
def __init__(self):
self.methods = {}
def z_score_detection(self, data, threshold=3):
"""
Detect anomalies using Z-score method
"""
from scipy import stats
# Calculate Z-scores
z_scores = np.abs(stats.zscore(data))
# Identify anomalies
anomalies = z_scores > threshold
return anomalies, z_scores
def iqr_detection(self, data, multiplier=1.5):
"""
Detect anomalies using Interquartile Range (IQR)
"""
Q1 = np.percentile(data, 25)
Q3 = np.percentile(data, 75)
IQR = Q3 - Q1
# Define bounds
lower_bound = Q1 - multiplier * IQR
upper_bound = Q3 + multiplier * IQR
# Identify anomalies
anomalies = (data < lower_bound) | (data > upper_bound)
return anomalies, (lower_bound, upper_bound)
def mahalanobis_distance(self, data):
"""
Calculate Mahalanobis distance for multivariate data
"""
from scipy.spatial.distance import mahalanobis
# Calculate covariance matrix and mean
cov_matrix = np.cov(data.T)
inv_cov_matrix = np.linalg.inv(cov_matrix)
mean = np.mean(data, axis=0)
# Calculate Mahalanobis distance for each point
distances = []
for point in data:
dist = mahalanobis(point, mean, inv_cov_matrix)
distances.append(dist)
return np.array(distances)
def visualize_methods(self, data):
"""
Visualize different statistical detection methods
"""
fig, axes = plt.subplots(2, 2, figsize=(12, 10))
# Generate sample data with anomalies
np.random.seed(42)
normal_data = np.random.normal(0, 1, 1000)
anomalies_data = np.random.uniform(-6, 6, 50)
all_data = np.concatenate([normal_data, anomalies_data])
np.random.shuffle(all_data)
# Z-score method
z_anomalies, z_scores = self.z_score_detection(all_data)
axes[0, 0].scatter(range(len(all_data)), all_data,
c=['red' if a else 'blue' for a in z_anomalies],
alpha=0.5, s=10)
axes[0, 0].axhline(y=3, color='r', linestyle='--', label='Z-score threshold')
axes[0, 0].axhline(y=-3, color='r', linestyle='--')
axes[0, 0].set_title('Z-Score Method')
axes[0, 0].set_xlabel('Index')
axes[0, 0].set_ylabel('Value')
axes[0, 0].legend()
# IQR method
iqr_anomalies, (lower, upper) = self.iqr_detection(all_data)
axes[0, 1].scatter(range(len(all_data)), all_data,
c=['red' if a else 'blue' for a in iqr_anomalies],
alpha=0.5, s=10)
axes[0, 1].axhline(y=upper, color='r', linestyle='--', label='Upper bound')
axes[0, 1].axhline(y=lower, color='r', linestyle='--', label='Lower bound')
axes[0, 1].set_title('IQR Method')
axes[0, 1].set_xlabel('Index')
axes[0, 1].set_ylabel('Value')
axes[0, 1].legend()
# Distribution plot
axes[1, 0].hist(all_data, bins=50, edgecolor='black', alpha=0.7)
axes[1, 0].set_title('Data Distribution')
axes[1, 0].set_xlabel('Value')
axes[1, 0].set_ylabel('Frequency')
# Box plot
axes[1, 1].boxplot(all_data)
axes[1, 1].set_title('Box Plot')
axes[1, 1].set_ylabel('Value')
plt.suptitle('Statistical Anomaly Detection Methods', fontsize=14)
plt.tight_layout()
plt.show()
# Print statistics
print("Statistical Summary:")
print(f"Total points: {len(all_data)}")
print(f"Z-score anomalies: {np.sum(z_anomalies)} ({np.sum(z_anomalies)/len(all_data)*100:.1f}%)")
print(f"IQR anomalies: {np.sum(iqr_anomalies)} ({np.sum(iqr_anomalies)/len(all_data)*100:.1f}%)")
# Demonstrate statistical methods
stat_detector = StatisticalAnomalyDetector()
print("\nVisualizing Statistical Methods:")
stat_detector.visualize_methods(None)
Machine Learning Approaches
from sklearn.ensemble import IsolationForest
from sklearn.svm import OneClassSVM
from sklearn.neighbors import LocalOutlierFactor
from sklearn.covariance import EllipticEnvelope
class MLAnomalyDetector:
"""
Machine Learning methods for anomaly detection
"""
def __init__(self):
self.models = {}
def generate_sample_data(self, n_samples=1000, contamination=0.1):
"""
Generate sample data with anomalies
"""
from sklearn.datasets import make_blobs
# Generate normal data (clusters)
n_inliers = int(n_samples * (1 - contamination))
n_outliers = n_samples - n_inliers
# Normal data
X_normal, _ = make_blobs(n_samples=n_inliers,
centers=2,
n_features=2,
cluster_std=0.5,
random_state=42)
# Anomalies (uniform random)
X_outliers = np.random.uniform(low=-6, high=6,
size=(n_outliers, 2))
# Combine data
X = np.vstack([X_normal, X_outliers])
y = np.hstack([np.ones(n_inliers), -np.ones(n_outliers)])
# Shuffle
shuffle_idx = np.random.permutation(n_samples)
X = X[shuffle_idx]
y = y[shuffle_idx]
return X, y
def train_models(self, X, contamination=0.1):
"""
Train different anomaly detection models
"""
models = {
'Isolation Forest': IsolationForest(
contamination=contamination,
random_state=42
),
'One-Class SVM': OneClassSVM(
nu=contamination,
kernel='rbf',
gamma='auto'
),
'Local Outlier Factor': LocalOutlierFactor(
contamination=contamination,
novelty=True
),
'Robust Covariance': EllipticEnvelope(
contamination=contamination,
random_state=42
)
}
# Train models
for name, model in models.items():
model.fit(X)
self.models[name] = model
return models
def compare_models(self, X, y_true):
"""
Compare performance of different models
"""
results = {}
fig, axes = plt.subplots(2, 3, figsize=(15, 10))
axes = axes.ravel()
# Plot original data
axes[0].scatter(X[:, 0], X[:, 1], c=y_true, cmap='coolwarm',
edgecolor='black', linewidth=0.5, s=30)
axes[0].set_title('True Labels')
axes[0].set_xlabel('Feature 1')
axes[0].set_ylabel('Feature 2')
# Evaluate each model
for idx, (name, model) in enumerate(self.models.items(), 1):
# Predict
y_pred = model.predict(X)
# Calculate metrics
accuracy = np.mean(y_pred == y_true)
tp = np.sum((y_pred == -1) & (y_true == -1))
fp = np.sum((y_pred == -1) & (y_true == 1))
fn = np.sum((y_pred == 1) & (y_true == -1))
if (tp + fp) > 0:
precision = tp / (tp + fp)
else:
precision = 0
if (tp + fn) > 0:
recall = tp / (tp + fn)
else:
recall = 0
results[name] = {
'accuracy': accuracy,
'precision': precision,
'recall': recall
}
# Plot predictions
axes[idx].scatter(X[:, 0], X[:, 1], c=y_pred, cmap='coolwarm',
edgecolor='black', linewidth=0.5, s=30)
axes[idx].set_title(f'{name}\nAcc: {accuracy:.2f}, Prec: {precision:.2f}')
axes[idx].set_xlabel('Feature 1')
axes[idx].set_ylabel('Feature 2')
# Remove empty subplot
if len(self.models) < 5:
fig.delaxes(axes[-1])
plt.suptitle('Anomaly Detection Models Comparison', fontsize=14)
plt.tight_layout()
plt.show()
return results
def plot_decision_boundaries(self, X, model_name='Isolation Forest'):
"""
Plot decision boundaries for a specific model
"""
model = self.models[model_name]
# Create mesh
h = 0.02
x_min, x_max = X[:, 0].min() - 1, X[:, 0].max() + 1
y_min, y_max = X[:, 1].min() - 1, X[:, 1].max() + 1
xx, yy = np.meshgrid(np.arange(x_min, x_max, h),
np.arange(y_min, y_max, h))
# Predict on mesh
Z = model.predict(np.c_[xx.ravel(), yy.ravel()])
Z = Z.reshape(xx.shape)
# Plot
plt.figure(figsize=(10, 8))
plt.contourf(xx, yy, Z, cmap=plt.cm.coolwarm, alpha=0.3)
# Plot data points
y_pred = model.predict(X)
plt.scatter(X[:, 0], X[:, 1], c=y_pred, cmap='coolwarm',
edgecolor='black', linewidth=1, s=50)
plt.xlabel('Feature 1')
plt.ylabel('Feature 2')
plt.title(f'{model_name} Decision Boundary')
plt.colorbar(label='Prediction (-1: Anomaly, 1: Normal)')
plt.grid(True, alpha=0.3)
plt.show()
# Demonstrate ML methods
ml_detector = MLAnomalyDetector()
# Generate data
print("\nGenerating sample data...")
X, y_true = ml_detector.generate_sample_data(n_samples=500, contamination=0.1)
# Train models
print("Training anomaly detection models...")
ml_detector.train_models(X, contamination=0.1)
# Compare models
print("\nComparing model performance:")
results = ml_detector.compare_models(X, y_true)
# Display results
print("\n" + "="*50)
print("Model Performance Summary:")
print("-"*50)
for name, metrics in results.items():
print(f"\n{name}:")
for metric, value in metrics.items():
print(f" {metric}: {value:.3f}")
# Show decision boundary for best model
print("\nPlotting decision boundary for Isolation Forest:")
ml_detector.plot_decision_boundaries(X, 'Isolation Forest')
Anomaly Detection Pipeline
Real-World Application: Credit Card Fraud Detection
class FraudDetectionSystem:
"""
Credit card fraud detection system using anomaly detection
"""
def __init__(self):
self.model = None
self.scaler = StandardScaler()
self.threshold = None
def prepare_features(self, transactions_df):
"""
Prepare transaction features for fraud detection
"""
# Feature engineering
features = pd.DataFrame()
# Transaction amount statistics
features['amount'] = transactions_df['amount']
features['amount_log'] = np.log1p(transactions_df['amount'])
# Time-based features (if timestamp available)
if 'timestamp' in transactions_df.columns:
features['hour'] = pd.to_datetime(transactions_df['timestamp']).dt.hour
features['day_of_week'] = pd.to_datetime(transactions_df['timestamp']).dt.dayofweek
features['is_weekend'] = features['day_of_week'].isin([5, 6]).astype(int)
# Merchant category features
if 'merchant_category' in transactions_df.columns:
features['is_online'] = (transactions_df['merchant_category'] == 'online').astype(int)
features['is_travel'] = (transactions_df['merchant_category'] == 'travel').astype(int)
# Location features
if 'location_risk' in transactions_df.columns:
features['location_risk'] = transactions_df['location_risk']
return features
def train(self, X_train, contamination=0.01):
"""
Train fraud detection model
"""
# Scale features
X_scaled = self.scaler.fit_transform(X_train)
# Train Isolation Forest (good for fraud detection)
self.model = IsolationForest(
contamination=contamination,
max_samples='auto',
random_state=42,
n_estimators=100
)
self.model.fit(X_scaled)
# Calculate anomaly scores
scores = self.model.score_samples(X_scaled)
self.threshold = np.percentile(scores, contamination * 100)
print(f"Model trained with contamination rate: {contamination}")
print(f"Anomaly threshold: {self.threshold:.3f}")
def predict(self, X):
"""
Predict if transactions are fraudulent
"""
X_scaled = self.scaler.transform(X)
# Get predictions (-1 for anomaly/fraud, 1 for normal)
predictions = self.model.predict(X_scaled)
# Get anomaly scores
scores = self.model.score_samples(X_scaled)
# Convert to fraud probability (0-1 scale)
fraud_probability = 1 - (scores - scores.min()) / (scores.max() - scores.min())
return predictions, fraud_probability
def evaluate(self, X_test, y_true):
"""
Evaluate model performance
"""
predictions, fraud_probs = self.predict(X_test)
# Convert predictions to binary (1 for fraud, 0 for normal)
y_pred = (predictions == -1).astype(int)
y_true_binary = (y_true == -1).astype(int)
# Calculate metrics
from sklearn.metrics import precision_recall_curve, roc_curve, auc
# Confusion matrix
cm = confusion_matrix(y_true_binary, y_pred)
# ROC curve
fpr, tpr, _ = roc_curve(y_true_binary, fraud_probs)
roc_auc = auc(fpr, tpr)
# Precision-Recall curve
precision, recall, _ = precision_recall_curve(y_true_binary, fraud_probs)
# Visualization
fig, axes = plt.subplots(1, 3, figsize=(15, 5))
# Confusion Matrix
sns.heatmap(cm, annot=True, fmt='d', cmap='Blues', ax=axes[0])
axes[0].set_title('Confusion Matrix')
axes[0].set_xlabel('Predicted')
axes[0].set_ylabel('Actual')
# ROC Curve
axes[1].plot(fpr, tpr, color='blue', lw=2,
label=f'ROC curve (AUC = {roc_auc:.2f})')
axes[1].plot([0, 1], [0, 1], color='red', lw=2, linestyle='--')
axes[1].set_xlabel('False Positive Rate')
axes[1].set_ylabel('True Positive Rate')
axes[1].set_title('ROC Curve')
axes[1].legend(loc="lower right")
axes[1].grid(True, alpha=0.3)
# Precision-Recall Curve
axes[2].plot(recall, precision, color='green', lw=2)
axes[2].set_xlabel('Recall')
axes[2].set_ylabel('Precision')
axes[2].set_title('Precision-Recall Curve')
axes[2].grid(True, alpha=0.3)
plt.suptitle('Fraud Detection Model Evaluation', fontsize=14)
plt.tight_layout()
plt.show()
# Print metrics
print("\nModel Performance:")
print(f"ROC AUC: {roc_auc:.3f}")
print(f"True Positives (Frauds caught): {cm[1, 1]}")
print(f"False Positives (False alarms): {cm[0, 1]}")
print(f"True Negatives (Normal correctly identified): {cm[0, 0]}")
print(f"False Negatives (Frauds missed): {cm[1, 0]}")
# Demonstrate fraud detection
print("\n" + "="*50)
print("CREDIT CARD FRAUD DETECTION EXAMPLE")
print("="*50)
# Generate synthetic transaction data
np.random.seed(42)
n_transactions = 10000
fraud_rate = 0.02
# Normal transactions
normal_transactions = pd.DataFrame({
'amount': np.random.lognormal(3, 1.5, int(n_transactions * (1 - fraud_rate))),
'hour': np.random.normal(14, 6, int(n_transactions * (1 - fraud_rate))).clip(0, 23).astype(int),
'location_risk': np.random.beta(2, 5, int(n_transactions * (1 - fraud_rate)))
})
# Fraudulent transactions (different patterns)
fraud_transactions = pd.DataFrame({
'amount': np.random.lognormal(5, 2, int(n_transactions * fraud_rate)), # Higher amounts
'hour': np.random.choice([3, 4, 5, 22, 23], int(n_transactions * fraud_rate)), # Odd hours
'location_risk': np.random.beta(5, 2, int(n_transactions * fraud_rate)) # Higher risk locations
})
# Combine and label
normal_transactions['is_fraud'] = 0
fraud_transactions['is_fraud'] = 1
all_transactions = pd.concat([normal_transactions, fraud_transactions]).reset_index(drop=True)
# Shuffle
all_transactions = all_transactions.sample(frac=1).reset_index(drop=True)
print(f"Total transactions: {len(all_transactions)}")
print(f"Fraudulent transactions: {all_transactions['is_fraud'].sum()} ({all_transactions['is_fraud'].mean()*100:.1f}%)")
# Prepare features
fraud_detector = FraudDetectionSystem()
X = all_transactions[['amount', 'hour', 'location_risk']]
y = all_transactions['is_fraud'].map({0: 1, 1: -1}) # Convert to anomaly detection format
# Split data
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
# Train model
print("\nTraining fraud detection model...")
fraud_detector.train(X_train, contamination=fraud_rate)
# Evaluate
print("\nEvaluating model performance...")
fraud_detector.evaluate(X_test, y_test)
Time Series Anomaly Detection
on Windows] D --> D2[LSTM Autoencoder] D --> D3[Prophet] C1 --> E[Detect Anomalies] D1 --> E E --> F[Point Anomalies] E --> G[Subsequence Anomalies] style A fill:#e3f2fd style F fill:#ffcdd2 style G fill:#ffcdd2
Best Practices
🎯 Anomaly Detection Guidelines
- Understand Normal: Define what constitutes normal behavior clearly
- Feature Engineering: Create relevant features that capture anomalous patterns
- Contamination Rate: Set based on domain knowledge or business requirements
- Multiple Methods: Combine different techniques for robust detection
- Handle Imbalance: Anomalies are rare - use appropriate evaluation metrics
- Continuous Learning: Update models as new patterns emerge
- False Positive Trade-off: Balance between catching anomalies and false alarms
- Interpretability: Understand why something is flagged as anomalous
Practice Exercises
Exercise 1: Network Intrusion Detection
Build an anomaly detection system for network security:
- Process network traffic features
- Implement real-time detection
- Compare different algorithms
- Minimize false positives
Exercise 2: Manufacturing Quality Control
Detect defective products in manufacturing:
- Use sensor data from production line
- Implement multivariate anomaly detection
- Set up alerting system
- Track anomaly trends over time
Summary
✅ You've Learned
- Types of anomalies: point, contextual, and collective
- Statistical methods: Z-score, IQR, Mahalanobis distance
- Machine learning approaches: Isolation Forest, One-Class SVM, LOF
- How to build a fraud detection system
- Evaluation metrics for imbalanced anomaly detection
- Best practices for production deployment
📓 Learning Journal
Keep a learning journal — digital or physical. After this lesson, take a few minutes to write down:
- Key concepts you learned
- Techniques that clicked for you
- Questions or confusion points to revisit
- Ideas you want to try
- Your progress and feelings about learning this
✍️ This lesson's prompt: Anomaly detection is largely about defining "normal." Pick a system you know well (a website, a machine, your own routine) — what would "normal" look like, and what kinds of anomalies would matter most to catch?
📝 Lesson Summary
🎓 Key Takeaways
- Anomalies come in three flavors — point, contextual, and collective — and the type shapes the method you choose.
- Statistical methods (z-score, IQR) are fast and interpretable; ML methods (Isolation Forest, One-Class SVM, LOF) capture more complex structure.
- Anomalies are rare, so accuracy is misleading — evaluate with precision, recall, and the precision–recall curve.
- Anomaly detection is often unsupervised or semi-supervised because labeled anomalies are scarce.
🎉 What You've Accomplished
You can now build an anomaly detection pipeline, choose between statistical and machine-learning detectors, and evaluate them honestly on imbalanced data — skills that transfer directly to fraud, intrusion, and quality-control problems.
❓ Common Questions at This Stage
Should I use a supervised classifier or an anomaly detector?
If you have plenty of labeled examples of the anomaly, a supervised classifier often wins. When anomalies are rare, unlabeled, or constantly evolving, dedicated anomaly detectors that model "normal" are the better fit.
How do I set the contamination or threshold?
Start from a domain estimate of how often anomalies occur, then adjust the threshold using a precision–recall trade-off that matches the cost of false alarms versus missed anomalies in your setting.
Why does Isolation Forest work so well?
It isolates points by random splits; anomalies, being few and different, get isolated in fewer splits than normal points. That makes it fast, scalable, and effective in higher dimensions.
🔭 Looking Ahead
The imbalanced-data mindset you practiced here — rare positives, careful metrics, tuned thresholds — carries straight into applied case studies such as fraud and churn, where finding the rare, costly cases is the whole point.
✅ Before the Next Lesson
- Run Isolation Forest and Local Outlier Factor on the same dataset and compare which points each flags.
- Plot a precision–recall curve for your detector and pick a threshold for a chosen recall target.
- Write your Learning Journal entry for this lesson
🌟 Encouragement for the Journey
Finding the rare, important signal in a sea of normal data is one of the most valuable things ML can do. You now have a toolkit for exactly that — keep hunting for the meaningful outliers.