Skip to main content

🔍 Anomaly Detection: Finding the Unusual

📚 What You'll Learn

By the end of this lesson, you will be able to:

  • Distinguish point, contextual, and collective anomalies and give real-world examples of each
  • Apply statistical methods (z-score, IQR, and distribution-based tests) to flag outliers
  • Use machine-learning detectors such as Isolation Forest, One-Class SVM, and Local Outlier Factor
  • Build an end-to-end anomaly detection pipeline and evaluate it on imbalanced data
  • Detect anomalies in time-series data and interpret the results for a stakeholder

⏱️ Estimated Time: 45–60 minutes

🎯 Project: Build an anomaly detection pipeline that flags fraudulent or unusual records and compare statistical vs machine-learning detectors.

Introduction

Anomaly detection is the identification of rare items, events, or observations that deviate significantly from the majority of the data. These anomalies, also called outliers, can indicate critical incidents such as fraud, system failures, or defects. This lesson covers various anomaly detection techniques from statistical methods to machine learning approaches, with practical applications in fraud detection, network security, and quality control.

Anomaly Detection Overview

flowchart TD A[Data Stream] --> B{Anomaly Detection} B --> C[Statistical Methods] C --> C1[Z-Score] C --> C2[IQR Method] C --> C3[Grubbs Test] B --> D[Machine Learning] D --> D1[Isolation Forest] D --> D2[One-Class SVM] D --> D3[Local Outlier Factor] D --> D4[Autoencoders] B --> E[Clustering Based] E --> E1[DBSCAN] E --> E2[K-Means Distance] C1 --> F[Anomalies Detected] D1 --> F E1 --> F F --> G[Alert/Action] style A fill:#e3f2fd style F fill:#ffccbc style G fill:#c8e6c9

Types of Anomalies

graph LR A[Anomaly Types] --> B[Point Anomaly] A --> C[Contextual Anomaly] A --> D[Collective Anomaly] B --> B1[Single Instance
Deviates from Normal] C --> C1[Normal in Different Context
Abnormal in Current] D --> D1[Group of Instances
Collectively Abnormal] B1 --> E[Example: Fraud Transaction] C1 --> F[Example: Hot Winter Day] D1 --> G[Example: Cyber Attack Pattern] style B fill:#ffebee style C fill:#fff3e0 style D fill:#e8f5e9

Setting Up the Environment

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report, confusion_matrix
import warnings
warnings.filterwarnings('ignore')

# Set style for better visualizations
plt.style.use('seaborn-v0_8-darkgrid')
sns.set_palette("husl")

# Set random seed for reproducibility
np.random.seed(42)

print("Anomaly Detection Methods Comparison")
print("=" * 50)

Statistical Methods

class StatisticalAnomalyDetector:
    """
    Statistical methods for anomaly detection
    """
    
    def __init__(self):
        self.methods = {}
        
    def z_score_detection(self, data, threshold=3):
        """
        Detect anomalies using Z-score method
        """
        from scipy import stats
        
        # Calculate Z-scores
        z_scores = np.abs(stats.zscore(data))
        
        # Identify anomalies
        anomalies = z_scores > threshold
        
        return anomalies, z_scores
    
    def iqr_detection(self, data, multiplier=1.5):
        """
        Detect anomalies using Interquartile Range (IQR)
        """
        Q1 = np.percentile(data, 25)
        Q3 = np.percentile(data, 75)
        IQR = Q3 - Q1
        
        # Define bounds
        lower_bound = Q1 - multiplier * IQR
        upper_bound = Q3 + multiplier * IQR
        
        # Identify anomalies
        anomalies = (data < lower_bound) | (data > upper_bound)
        
        return anomalies, (lower_bound, upper_bound)
    
    def mahalanobis_distance(self, data):
        """
        Calculate Mahalanobis distance for multivariate data
        """
        from scipy.spatial.distance import mahalanobis
        
        # Calculate covariance matrix and mean
        cov_matrix = np.cov(data.T)
        inv_cov_matrix = np.linalg.inv(cov_matrix)
        mean = np.mean(data, axis=0)
        
        # Calculate Mahalanobis distance for each point
        distances = []
        for point in data:
            dist = mahalanobis(point, mean, inv_cov_matrix)
            distances.append(dist)
        
        return np.array(distances)
    
    def visualize_methods(self, data):
        """
        Visualize different statistical detection methods
        """
        fig, axes = plt.subplots(2, 2, figsize=(12, 10))
        
        # Generate sample data with anomalies
        np.random.seed(42)
        normal_data = np.random.normal(0, 1, 1000)
        anomalies_data = np.random.uniform(-6, 6, 50)
        all_data = np.concatenate([normal_data, anomalies_data])
        np.random.shuffle(all_data)
        
        # Z-score method
        z_anomalies, z_scores = self.z_score_detection(all_data)
        axes[0, 0].scatter(range(len(all_data)), all_data, 
                          c=['red' if a else 'blue' for a in z_anomalies],
                          alpha=0.5, s=10)
        axes[0, 0].axhline(y=3, color='r', linestyle='--', label='Z-score threshold')
        axes[0, 0].axhline(y=-3, color='r', linestyle='--')
        axes[0, 0].set_title('Z-Score Method')
        axes[0, 0].set_xlabel('Index')
        axes[0, 0].set_ylabel('Value')
        axes[0, 0].legend()
        
        # IQR method
        iqr_anomalies, (lower, upper) = self.iqr_detection(all_data)
        axes[0, 1].scatter(range(len(all_data)), all_data,
                          c=['red' if a else 'blue' for a in iqr_anomalies],
                          alpha=0.5, s=10)
        axes[0, 1].axhline(y=upper, color='r', linestyle='--', label='Upper bound')
        axes[0, 1].axhline(y=lower, color='r', linestyle='--', label='Lower bound')
        axes[0, 1].set_title('IQR Method')
        axes[0, 1].set_xlabel('Index')
        axes[0, 1].set_ylabel('Value')
        axes[0, 1].legend()
        
        # Distribution plot
        axes[1, 0].hist(all_data, bins=50, edgecolor='black', alpha=0.7)
        axes[1, 0].set_title('Data Distribution')
        axes[1, 0].set_xlabel('Value')
        axes[1, 0].set_ylabel('Frequency')
        
        # Box plot
        axes[1, 1].boxplot(all_data)
        axes[1, 1].set_title('Box Plot')
        axes[1, 1].set_ylabel('Value')
        
        plt.suptitle('Statistical Anomaly Detection Methods', fontsize=14)
        plt.tight_layout()
        plt.show()
        
        # Print statistics
        print("Statistical Summary:")
        print(f"Total points: {len(all_data)}")
        print(f"Z-score anomalies: {np.sum(z_anomalies)} ({np.sum(z_anomalies)/len(all_data)*100:.1f}%)")
        print(f"IQR anomalies: {np.sum(iqr_anomalies)} ({np.sum(iqr_anomalies)/len(all_data)*100:.1f}%)")

# Demonstrate statistical methods
stat_detector = StatisticalAnomalyDetector()
print("\nVisualizing Statistical Methods:")
stat_detector.visualize_methods(None)

Machine Learning Approaches

from sklearn.ensemble import IsolationForest
from sklearn.svm import OneClassSVM
from sklearn.neighbors import LocalOutlierFactor
from sklearn.covariance import EllipticEnvelope

class MLAnomalyDetector:
    """
    Machine Learning methods for anomaly detection
    """
    
    def __init__(self):
        self.models = {}
        
    def generate_sample_data(self, n_samples=1000, contamination=0.1):
        """
        Generate sample data with anomalies
        """
        from sklearn.datasets import make_blobs
        
        # Generate normal data (clusters)
        n_inliers = int(n_samples * (1 - contamination))
        n_outliers = n_samples - n_inliers
        
        # Normal data
        X_normal, _ = make_blobs(n_samples=n_inliers, 
                                 centers=2, 
                                 n_features=2,
                                 cluster_std=0.5,
                                 random_state=42)
        
        # Anomalies (uniform random)
        X_outliers = np.random.uniform(low=-6, high=6, 
                                       size=(n_outliers, 2))
        
        # Combine data
        X = np.vstack([X_normal, X_outliers])
        y = np.hstack([np.ones(n_inliers), -np.ones(n_outliers)])
        
        # Shuffle
        shuffle_idx = np.random.permutation(n_samples)
        X = X[shuffle_idx]
        y = y[shuffle_idx]
        
        return X, y
    
    def train_models(self, X, contamination=0.1):
        """
        Train different anomaly detection models
        """
        models = {
            'Isolation Forest': IsolationForest(
                contamination=contamination,
                random_state=42
            ),
            'One-Class SVM': OneClassSVM(
                nu=contamination,
                kernel='rbf',
                gamma='auto'
            ),
            'Local Outlier Factor': LocalOutlierFactor(
                contamination=contamination,
                novelty=True
            ),
            'Robust Covariance': EllipticEnvelope(
                contamination=contamination,
                random_state=42
            )
        }
        
        # Train models
        for name, model in models.items():
            model.fit(X)
            self.models[name] = model
            
        return models
    
    def compare_models(self, X, y_true):
        """
        Compare performance of different models
        """
        results = {}
        
        fig, axes = plt.subplots(2, 3, figsize=(15, 10))
        axes = axes.ravel()
        
        # Plot original data
        axes[0].scatter(X[:, 0], X[:, 1], c=y_true, cmap='coolwarm', 
                       edgecolor='black', linewidth=0.5, s=30)
        axes[0].set_title('True Labels')
        axes[0].set_xlabel('Feature 1')
        axes[0].set_ylabel('Feature 2')
        
        # Evaluate each model
        for idx, (name, model) in enumerate(self.models.items(), 1):
            # Predict
            y_pred = model.predict(X)
            
            # Calculate metrics
            accuracy = np.mean(y_pred == y_true)
            tp = np.sum((y_pred == -1) & (y_true == -1))
            fp = np.sum((y_pred == -1) & (y_true == 1))
            fn = np.sum((y_pred == 1) & (y_true == -1))
            
            if (tp + fp) > 0:
                precision = tp / (tp + fp)
            else:
                precision = 0
                
            if (tp + fn) > 0:
                recall = tp / (tp + fn)
            else:
                recall = 0
            
            results[name] = {
                'accuracy': accuracy,
                'precision': precision,
                'recall': recall
            }
            
            # Plot predictions
            axes[idx].scatter(X[:, 0], X[:, 1], c=y_pred, cmap='coolwarm',
                            edgecolor='black', linewidth=0.5, s=30)
            axes[idx].set_title(f'{name}\nAcc: {accuracy:.2f}, Prec: {precision:.2f}')
            axes[idx].set_xlabel('Feature 1')
            axes[idx].set_ylabel('Feature 2')
        
        # Remove empty subplot
        if len(self.models) < 5:
            fig.delaxes(axes[-1])
        
        plt.suptitle('Anomaly Detection Models Comparison', fontsize=14)
        plt.tight_layout()
        plt.show()
        
        return results
    
    def plot_decision_boundaries(self, X, model_name='Isolation Forest'):
        """
        Plot decision boundaries for a specific model
        """
        model = self.models[model_name]
        
        # Create mesh
        h = 0.02
        x_min, x_max = X[:, 0].min() - 1, X[:, 0].max() + 1
        y_min, y_max = X[:, 1].min() - 1, X[:, 1].max() + 1
        xx, yy = np.meshgrid(np.arange(x_min, x_max, h),
                            np.arange(y_min, y_max, h))
        
        # Predict on mesh
        Z = model.predict(np.c_[xx.ravel(), yy.ravel()])
        Z = Z.reshape(xx.shape)
        
        # Plot
        plt.figure(figsize=(10, 8))
        plt.contourf(xx, yy, Z, cmap=plt.cm.coolwarm, alpha=0.3)
        
        # Plot data points
        y_pred = model.predict(X)
        plt.scatter(X[:, 0], X[:, 1], c=y_pred, cmap='coolwarm',
                   edgecolor='black', linewidth=1, s=50)
        
        plt.xlabel('Feature 1')
        plt.ylabel('Feature 2')
        plt.title(f'{model_name} Decision Boundary')
        plt.colorbar(label='Prediction (-1: Anomaly, 1: Normal)')
        plt.grid(True, alpha=0.3)
        plt.show()

# Demonstrate ML methods
ml_detector = MLAnomalyDetector()

# Generate data
print("\nGenerating sample data...")
X, y_true = ml_detector.generate_sample_data(n_samples=500, contamination=0.1)

# Train models
print("Training anomaly detection models...")
ml_detector.train_models(X, contamination=0.1)

# Compare models
print("\nComparing model performance:")
results = ml_detector.compare_models(X, y_true)

# Display results
print("\n" + "="*50)
print("Model Performance Summary:")
print("-"*50)
for name, metrics in results.items():
    print(f"\n{name}:")
    for metric, value in metrics.items():
        print(f"  {metric}: {value:.3f}")

# Show decision boundary for best model
print("\nPlotting decision boundary for Isolation Forest:")
ml_detector.plot_decision_boundaries(X, 'Isolation Forest')

Anomaly Detection Pipeline

flowchart LR A[Raw Data] --> B[Preprocessing] B --> B1[Handle Missing] B --> B2[Scale Features] B --> B3[Feature Engineering] B --> C[Model Selection] C --> D{Data Type} D -->|Low Dimensional| E[Statistical Methods] D -->|High Dimensional| F[ML Methods] D -->|Time Series| G[LSTM/ARIMA] E --> H[Train Model] F --> H G --> H H --> I[Set Threshold] I --> J[Detect Anomalies] J --> K[Validate Results] K --> L{Performance OK?} L -->|No| M[Adjust Parameters] M --> H L -->|Yes| N[Deploy] style A fill:#e3f2fd style N fill:#c8e6c9

Real-World Application: Credit Card Fraud Detection

class FraudDetectionSystem:
    """
    Credit card fraud detection system using anomaly detection
    """
    
    def __init__(self):
        self.model = None
        self.scaler = StandardScaler()
        self.threshold = None
        
    def prepare_features(self, transactions_df):
        """
        Prepare transaction features for fraud detection
        """
        # Feature engineering
        features = pd.DataFrame()
        
        # Transaction amount statistics
        features['amount'] = transactions_df['amount']
        features['amount_log'] = np.log1p(transactions_df['amount'])
        
        # Time-based features (if timestamp available)
        if 'timestamp' in transactions_df.columns:
            features['hour'] = pd.to_datetime(transactions_df['timestamp']).dt.hour
            features['day_of_week'] = pd.to_datetime(transactions_df['timestamp']).dt.dayofweek
            features['is_weekend'] = features['day_of_week'].isin([5, 6]).astype(int)
        
        # Merchant category features
        if 'merchant_category' in transactions_df.columns:
            features['is_online'] = (transactions_df['merchant_category'] == 'online').astype(int)
            features['is_travel'] = (transactions_df['merchant_category'] == 'travel').astype(int)
        
        # Location features
        if 'location_risk' in transactions_df.columns:
            features['location_risk'] = transactions_df['location_risk']
        
        return features
    
    def train(self, X_train, contamination=0.01):
        """
        Train fraud detection model
        """
        # Scale features
        X_scaled = self.scaler.fit_transform(X_train)
        
        # Train Isolation Forest (good for fraud detection)
        self.model = IsolationForest(
            contamination=contamination,
            max_samples='auto',
            random_state=42,
            n_estimators=100
        )
        
        self.model.fit(X_scaled)
        
        # Calculate anomaly scores
        scores = self.model.score_samples(X_scaled)
        self.threshold = np.percentile(scores, contamination * 100)
        
        print(f"Model trained with contamination rate: {contamination}")
        print(f"Anomaly threshold: {self.threshold:.3f}")
        
    def predict(self, X):
        """
        Predict if transactions are fraudulent
        """
        X_scaled = self.scaler.transform(X)
        
        # Get predictions (-1 for anomaly/fraud, 1 for normal)
        predictions = self.model.predict(X_scaled)
        
        # Get anomaly scores
        scores = self.model.score_samples(X_scaled)
        
        # Convert to fraud probability (0-1 scale)
        fraud_probability = 1 - (scores - scores.min()) / (scores.max() - scores.min())
        
        return predictions, fraud_probability
    
    def evaluate(self, X_test, y_true):
        """
        Evaluate model performance
        """
        predictions, fraud_probs = self.predict(X_test)
        
        # Convert predictions to binary (1 for fraud, 0 for normal)
        y_pred = (predictions == -1).astype(int)
        y_true_binary = (y_true == -1).astype(int)
        
        # Calculate metrics
        from sklearn.metrics import precision_recall_curve, roc_curve, auc
        
        # Confusion matrix
        cm = confusion_matrix(y_true_binary, y_pred)
        
        # ROC curve
        fpr, tpr, _ = roc_curve(y_true_binary, fraud_probs)
        roc_auc = auc(fpr, tpr)
        
        # Precision-Recall curve
        precision, recall, _ = precision_recall_curve(y_true_binary, fraud_probs)
        
        # Visualization
        fig, axes = plt.subplots(1, 3, figsize=(15, 5))
        
        # Confusion Matrix
        sns.heatmap(cm, annot=True, fmt='d', cmap='Blues', ax=axes[0])
        axes[0].set_title('Confusion Matrix')
        axes[0].set_xlabel('Predicted')
        axes[0].set_ylabel('Actual')
        
        # ROC Curve
        axes[1].plot(fpr, tpr, color='blue', lw=2, 
                    label=f'ROC curve (AUC = {roc_auc:.2f})')
        axes[1].plot([0, 1], [0, 1], color='red', lw=2, linestyle='--')
        axes[1].set_xlabel('False Positive Rate')
        axes[1].set_ylabel('True Positive Rate')
        axes[1].set_title('ROC Curve')
        axes[1].legend(loc="lower right")
        axes[1].grid(True, alpha=0.3)
        
        # Precision-Recall Curve
        axes[2].plot(recall, precision, color='green', lw=2)
        axes[2].set_xlabel('Recall')
        axes[2].set_ylabel('Precision')
        axes[2].set_title('Precision-Recall Curve')
        axes[2].grid(True, alpha=0.3)
        
        plt.suptitle('Fraud Detection Model Evaluation', fontsize=14)
        plt.tight_layout()
        plt.show()
        
        # Print metrics
        print("\nModel Performance:")
        print(f"ROC AUC: {roc_auc:.3f}")
        print(f"True Positives (Frauds caught): {cm[1, 1]}")
        print(f"False Positives (False alarms): {cm[0, 1]}")
        print(f"True Negatives (Normal correctly identified): {cm[0, 0]}")
        print(f"False Negatives (Frauds missed): {cm[1, 0]}")

# Demonstrate fraud detection
print("\n" + "="*50)
print("CREDIT CARD FRAUD DETECTION EXAMPLE")
print("="*50)

# Generate synthetic transaction data
np.random.seed(42)
n_transactions = 10000
fraud_rate = 0.02

# Normal transactions
normal_transactions = pd.DataFrame({
    'amount': np.random.lognormal(3, 1.5, int(n_transactions * (1 - fraud_rate))),
    'hour': np.random.normal(14, 6, int(n_transactions * (1 - fraud_rate))).clip(0, 23).astype(int),
    'location_risk': np.random.beta(2, 5, int(n_transactions * (1 - fraud_rate)))
})

# Fraudulent transactions (different patterns)
fraud_transactions = pd.DataFrame({
    'amount': np.random.lognormal(5, 2, int(n_transactions * fraud_rate)),  # Higher amounts
    'hour': np.random.choice([3, 4, 5, 22, 23], int(n_transactions * fraud_rate)),  # Odd hours
    'location_risk': np.random.beta(5, 2, int(n_transactions * fraud_rate))  # Higher risk locations
})

# Combine and label
normal_transactions['is_fraud'] = 0
fraud_transactions['is_fraud'] = 1
all_transactions = pd.concat([normal_transactions, fraud_transactions]).reset_index(drop=True)

# Shuffle
all_transactions = all_transactions.sample(frac=1).reset_index(drop=True)

print(f"Total transactions: {len(all_transactions)}")
print(f"Fraudulent transactions: {all_transactions['is_fraud'].sum()} ({all_transactions['is_fraud'].mean()*100:.1f}%)")

# Prepare features
fraud_detector = FraudDetectionSystem()
X = all_transactions[['amount', 'hour', 'location_risk']]
y = all_transactions['is_fraud'].map({0: 1, 1: -1})  # Convert to anomaly detection format

# Split data
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)

# Train model
print("\nTraining fraud detection model...")
fraud_detector.train(X_train, contamination=fraud_rate)

# Evaluate
print("\nEvaluating model performance...")
fraud_detector.evaluate(X_test, y_test)

Time Series Anomaly Detection

graph TD A[Time Series Data] --> B{Method} B --> C[Statistical] C --> C1[Moving Average] C --> C2[ARIMA Residuals] C --> C3[Seasonal Decomposition] B --> D[Machine Learning] D --> D1[Isolation Forest
on Windows] D --> D2[LSTM Autoencoder] D --> D3[Prophet] C1 --> E[Detect Anomalies] D1 --> E E --> F[Point Anomalies] E --> G[Subsequence Anomalies] style A fill:#e3f2fd style F fill:#ffcdd2 style G fill:#ffcdd2

Best Practices

🎯 Anomaly Detection Guidelines

  • Understand Normal: Define what constitutes normal behavior clearly
  • Feature Engineering: Create relevant features that capture anomalous patterns
  • Contamination Rate: Set based on domain knowledge or business requirements
  • Multiple Methods: Combine different techniques for robust detection
  • Handle Imbalance: Anomalies are rare - use appropriate evaluation metrics
  • Continuous Learning: Update models as new patterns emerge
  • False Positive Trade-off: Balance between catching anomalies and false alarms
  • Interpretability: Understand why something is flagged as anomalous

Practice Exercises

Exercise 1: Network Intrusion Detection

Build an anomaly detection system for network security:

  • Process network traffic features
  • Implement real-time detection
  • Compare different algorithms
  • Minimize false positives

Exercise 2: Manufacturing Quality Control

Detect defective products in manufacturing:

  • Use sensor data from production line
  • Implement multivariate anomaly detection
  • Set up alerting system
  • Track anomaly trends over time

Summary

✅ You've Learned

  • Types of anomalies: point, contextual, and collective
  • Statistical methods: Z-score, IQR, Mahalanobis distance
  • Machine learning approaches: Isolation Forest, One-Class SVM, LOF
  • How to build a fraud detection system
  • Evaluation metrics for imbalanced anomaly detection
  • Best practices for production deployment

📓 Learning Journal

Keep a learning journal — digital or physical. After this lesson, take a few minutes to write down:

  • Key concepts you learned
  • Techniques that clicked for you
  • Questions or confusion points to revisit
  • Ideas you want to try
  • Your progress and feelings about learning this

✍️ This lesson's prompt: Anomaly detection is largely about defining "normal." Pick a system you know well (a website, a machine, your own routine) — what would "normal" look like, and what kinds of anomalies would matter most to catch?

📝 Lesson Summary

🎓 Key Takeaways

  • Anomalies come in three flavors — point, contextual, and collective — and the type shapes the method you choose.
  • Statistical methods (z-score, IQR) are fast and interpretable; ML methods (Isolation Forest, One-Class SVM, LOF) capture more complex structure.
  • Anomalies are rare, so accuracy is misleading — evaluate with precision, recall, and the precision–recall curve.
  • Anomaly detection is often unsupervised or semi-supervised because labeled anomalies are scarce.

🎉 What You've Accomplished

You can now build an anomaly detection pipeline, choose between statistical and machine-learning detectors, and evaluate them honestly on imbalanced data — skills that transfer directly to fraud, intrusion, and quality-control problems.

❓ Common Questions at This Stage

Should I use a supervised classifier or an anomaly detector?

If you have plenty of labeled examples of the anomaly, a supervised classifier often wins. When anomalies are rare, unlabeled, or constantly evolving, dedicated anomaly detectors that model "normal" are the better fit.

How do I set the contamination or threshold?

Start from a domain estimate of how often anomalies occur, then adjust the threshold using a precision–recall trade-off that matches the cost of false alarms versus missed anomalies in your setting.

Why does Isolation Forest work so well?

It isolates points by random splits; anomalies, being few and different, get isolated in fewer splits than normal points. That makes it fast, scalable, and effective in higher dimensions.

🔭 Looking Ahead

The imbalanced-data mindset you practiced here — rare positives, careful metrics, tuned thresholds — carries straight into applied case studies such as fraud and churn, where finding the rare, costly cases is the whole point.

✅ Before the Next Lesson

🌟 Encouragement for the Journey

Finding the rare, important signal in a sea of normal data is one of the most valuable things ML can do. You now have a toolkit for exactly that — keep hunting for the meaningful outliers.