Advanced Ensemble Methods: The Power of Many
π What You'll Learn
By the end of this lesson, you will be able to:
- Explain why combining diverse models can reduce variance and/or bias and beat any single model
- Distinguish the main ensemble families: bagging, boosting, voting, and stacking
- Build voting and stacking ensembles in scikit-learn from several base learners
- Understand the central role of model diversity and proper cross-validation in a strong ensemble
- Judge when an ensemble's added complexity and compute cost are worth it β and when they aren't
β±οΈ Estimated Time: 45β60 minutes
π― Project: Build a voting ensemble and a stacking ensemble from several base learners, then compare their performance against the best single model.
Ensemble Learning: When One Model Isn't Enough π
"In diversity there is strength." This principle lies at the heart of ensemble methods. By combining multiple models that make different errors, we create a collective intelligence that outperforms any individual model. From Kaggle competitions to production systems, ensembles are the secret weapon of top data scientists.
The Mathematics of Ensemble Learning
# Understanding the Theory Behind Ensembles
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.model_selection import train_test_split, cross_val_score, KFold
from sklearn.datasets import make_classification, make_regression
from sklearn.metrics import accuracy_score, mean_squared_error
import warnings
warnings.filterwarnings('ignore')
# Set style
plt.style.use('seaborn-v0_8-darkgrid')
sns.set_palette("husl")
print("="*60)
print("THE MATHEMATICS OF ENSEMBLE LEARNING")
print("="*60)
# Demonstration: Why Ensembles Work
np.random.seed(42)
# Create a challenging dataset
X, y = make_classification(
n_samples=1000,
n_features=20,
n_informative=15,
n_redundant=5,
n_clusters_per_class=2,
flip_y=0.1, # Add noise
random_state=42
)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
# Simulate multiple weak learners with different errors
n_models = 100
predictions = []
for i in range(n_models):
# Each model has 65% accuracy (weak learner)
# But makes different mistakes
np.random.seed(i)
# Simulate predictions with controlled accuracy
correct_mask = np.random.random(len(y_test)) < 0.65
model_pred = y_test.copy()
model_pred[~correct_mask] = 1 - model_pred[~correct_mask] # Flip incorrect predictions
predictions.append(model_pred)
predictions = np.array(predictions)
# Individual model performance
individual_accuracies = [accuracy_score(y_test, pred) for pred in predictions]
print(f"Individual Model Performance:")
print(f" Mean accuracy: {np.mean(individual_accuracies):.3f}")
print(f" Std deviation: {np.std(individual_accuracies):.3f}")
# Ensemble by majority voting
ensemble_pred = (predictions.mean(axis=0) > 0.5).astype(int)
ensemble_accuracy = accuracy_score(y_test, ensemble_pred)
print(f"\nEnsemble Performance (Majority Vote):")
print(f" Accuracy: {ensemble_accuracy:.3f}")
print(f" Improvement: +{(ensemble_accuracy - np.mean(individual_accuracies)):.3f}")
# Visualize the power of ensembling
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
# Plot 1: Accuracy vs number of models
ensemble_scores = []
for n in range(1, n_models + 1):
partial_ensemble = (predictions[:n].mean(axis=0) > 0.5).astype(int)
ensemble_scores.append(accuracy_score(y_test, partial_ensemble))
axes[0].plot(range(1, n_models + 1), ensemble_scores, 'b-', linewidth=2)
axes[0].axhline(np.mean(individual_accuracies), color='r', linestyle='--',
label='Mean Individual Accuracy')
axes[0].set_xlabel('Number of Models in Ensemble')
axes[0].set_ylabel('Accuracy')
axes[0].set_title('Ensemble Performance vs Number of Models')
axes[0].legend()
axes[0].grid(True, alpha=0.3)
# Plot 2: Error correlation
correlation_matrix = np.corrcoef(predictions)
sns.heatmap(correlation_matrix[:10, :10], cmap='coolwarm', center=0,
ax=axes[1], cbar_kws={'label': 'Correlation'})
axes[1].set_title('Error Correlation Between Models (First 10)')
plt.tight_layout()
plt.show()
print("\nπ‘ Key Insight:")
print("Even though individual models are weak (65% accuracy),")
print("the ensemble achieves much higher accuracy because models")
print("make different errors (low correlation).")
Advanced Voting Methods
from sklearn.ensemble import VotingClassifier, VotingRegressor
from sklearn.linear_model import LogisticRegression, Ridge
from sklearn.svm import SVC, SVR
from sklearn.tree import DecisionTreeClassifier, DecisionTreeRegressor
from sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier
from sklearn.naive_bayes import GaussianNB
from sklearn.neighbors import KNeighborsClassifier
print("\n" + "="*60)
print("ADVANCED VOTING METHODS")
print("="*60)
# 1. HARD VS SOFT VOTING
print("\n1. Hard vs Soft Voting Comparison")
print("-" * 40)
# Create diverse base models
base_models = [
('lr', LogisticRegression(random_state=42)),
('svc', SVC(probability=True, random_state=42)), # Need probability=True for soft voting
('nb', GaussianNB()),
('rf', RandomForestClassifier(n_estimators=50, random_state=42))
]
# Hard voting (majority vote)
hard_voting = VotingClassifier(estimators=base_models, voting='hard')
hard_voting.fit(X_train, y_train)
hard_score = hard_voting.score(X_test, y_test)
# Soft voting (average probabilities)
soft_voting = VotingClassifier(estimators=base_models, voting='soft')
soft_voting.fit(X_train, y_train)
soft_score = soft_voting.score(X_test, y_test)
print(f"Hard Voting Accuracy: {hard_score:.3f}")
print(f"Soft Voting Accuracy: {soft_score:.3f}")
# Individual model scores
print("\nIndividual Model Performance:")
for name, model in base_models:
model.fit(X_train, y_train)
score = model.score(X_test, y_test)
print(f" {name}: {score:.3f}")
# 2. WEIGHTED VOTING
print("\n2. Weighted Voting")
print("-" * 40)
# Calculate weights based on cross-validation performance
from sklearn.model_selection import cross_val_score
weights = []
for name, model in base_models:
cv_scores = cross_val_score(model, X_train, y_train, cv=5)
weight = cv_scores.mean()
weights.append(weight)
print(f"{name} weight (CV score): {weight:.3f}")
# Normalize weights
weights = np.array(weights)
weights = weights / weights.sum()
print(f"\nNormalized weights: {weights}")
# Weighted voting
weighted_voting = VotingClassifier(
estimators=base_models,
voting='soft',
weights=weights
)
weighted_voting.fit(X_train, y_train)
weighted_score = weighted_voting.score(X_test, y_test)
print(f"\nWeighted Voting Accuracy: {weighted_score:.3f}")
print(f"Improvement over uniform soft voting: +{(weighted_score - soft_score):.3f}")
# 3. DYNAMIC VOTING (Custom Implementation)
print("\n3. Dynamic Voting (Confidence-Based)")
print("-" * 40)
class DynamicVotingClassifier:
"""
Dynamic voting based on prediction confidence.
Models with higher confidence get more weight.
"""
def __init__(self, estimators):
self.estimators = estimators
self.fitted_estimators_ = []
def fit(self, X, y):
"""Fit all base estimators"""
self.fitted_estimators_ = []
for name, estimator in self.estimators:
fitted_est = estimator.fit(X, y)
self.fitted_estimators_.append((name, fitted_est))
return self
def predict_proba(self, X):
"""Predict with dynamic weights based on confidence"""
probas = []
weights = []
for name, estimator in self.fitted_estimators_:
proba = estimator.predict_proba(X)
# Calculate confidence (max probability per sample)
confidence = np.max(proba, axis=1)
probas.append(proba)
weights.append(confidence)
# Stack and weight predictions
probas = np.array(probas)
weights = np.array(weights)
# Normalize weights per sample
weights = weights / weights.sum(axis=0)
# Weighted average
weighted_proba = np.zeros_like(probas[0])
for i, proba in enumerate(probas):
weighted_proba += proba * weights[i].reshape(-1, 1)
return weighted_proba
def predict(self, X):
"""Predict class labels"""
probas = self.predict_proba(X)
return np.argmax(probas, axis=1)
def score(self, X, y):
"""Calculate accuracy"""
return accuracy_score(y, self.predict(X))
# Test dynamic voting
dynamic_voting = DynamicVotingClassifier(base_models)
dynamic_voting.fit(X_train, y_train)
dynamic_score = dynamic_voting.score(X_test, y_test)
print(f"Dynamic Voting Accuracy: {dynamic_score:.3f}")
# Compare all voting methods
voting_comparison = pd.DataFrame({
'Method': ['Hard Voting', 'Soft Voting', 'Weighted Voting', 'Dynamic Voting'],
'Accuracy': [hard_score, soft_score, weighted_score, dynamic_score]
})
print("\n" + voting_comparison.to_string(index=False))
Advanced Stacking Techniques
from sklearn.ensemble import StackingClassifier, StackingRegressor
from sklearn.linear_model import LogisticRegression, RidgeCV
from sklearn.model_selection import StratifiedKFold
print("\n" + "="*60)
print("ADVANCED STACKING TECHNIQUES")
print("="*60)
# 1. MULTI-LEVEL STACKING
print("\n1. Multi-Level Stacking")
print("-" * 40)
# Level 0: Base models
level_0_models = [
('rf', RandomForestClassifier(n_estimators=100, random_state=42)),
('gb', GradientBoostingClassifier(n_estimators=100, random_state=42)),
('svc', SVC(probability=True, random_state=42))
]
# Level 1: Stacking on base predictions
level_1_models = [
('stack1', StackingClassifier(
estimators=level_0_models,
final_estimator=LogisticRegression(),
cv=5
))
]
# Level 2: Final stacking
multi_level_stacker = StackingClassifier(
estimators=level_1_models + [('knn', KNeighborsClassifier())],
final_estimator=LogisticRegression(),
cv=3
)
print("Multi-Level Stacking Structure:")
print(" Level 0: RF, GB, SVC")
print(" Level 1: Stacking(Level 0) + KNN")
print(" Level 2: LogisticRegression(Level 1)")
# Train (using smaller dataset for speed)
X_small = X_train[:500]
y_small = y_train[:500]
multi_level_stacker.fit(X_small, y_small)
multi_level_score = multi_level_stacker.score(X_test, y_test)
print(f"\nMulti-Level Stacking Accuracy: {multi_level_score:.3f}")
# 2. STACKING WITH FEATURE ENGINEERING
print("\n2. Stacking with Feature Engineering")
print("-" * 40)
class FeatureStackingClassifier:
"""
Stacking that includes original features along with predictions
"""
def __init__(self, base_estimators, meta_estimator, use_probas=True,
use_features=True, cv=5):
self.base_estimators = base_estimators
self.meta_estimator = meta_estimator
self.use_probas = use_probas
self.use_features = use_features
self.cv = cv
def fit(self, X, y):
"""Fit the stacking classifier"""
self.fitted_base_ = []
kfold = StratifiedKFold(n_splits=self.cv, shuffle=True, random_state=42)
# Prepare meta features
if self.use_probas:
# For probability predictions
n_classes = len(np.unique(y))
meta_features = np.zeros((X.shape[0], len(self.base_estimators) * n_classes))
else:
meta_features = np.zeros((X.shape[0], len(self.base_estimators)))
# Generate out-of-fold predictions
for i, (name, estimator) in enumerate(self.base_estimators):
oof_preds = np.zeros((X.shape[0], n_classes)) if self.use_probas else np.zeros(X.shape[0])
for train_idx, val_idx in kfold.split(X, y):
X_fold_train, X_fold_val = X[train_idx], X[val_idx]
y_fold_train = y[train_idx]
# Fit on fold
est_fold = estimator.__class__(**estimator.get_params())
est_fold.fit(X_fold_train, y_fold_train)
# Predict on validation fold
if self.use_probas:
oof_preds[val_idx] = est_fold.predict_proba(X_fold_val)
else:
oof_preds[val_idx] = est_fold.predict(X_fold_val)
# Store predictions
if self.use_probas:
meta_features[:, i*n_classes:(i+1)*n_classes] = oof_preds
else:
meta_features[:, i] = oof_preds
# Fit on full training set
fitted_est = estimator.fit(X, y)
self.fitted_base_.append((name, fitted_est))
# Combine with original features if requested
if self.use_features:
self.meta_features_ = np.hstack([X, meta_features])
else:
self.meta_features_ = meta_features
# Fit meta estimator
self.meta_estimator.fit(self.meta_features_, y)
return self
def predict(self, X):
"""Make predictions"""
# Generate base predictions
if self.use_probas:
base_preds = []
for name, estimator in self.fitted_base_:
base_preds.append(estimator.predict_proba(X))
base_preds = np.hstack(base_preds)
else:
base_preds = np.column_stack([
est.predict(X) for _, est in self.fitted_base_
])
# Combine with features if needed
if self.use_features:
meta_features = np.hstack([X, base_preds])
else:
meta_features = base_preds
return self.meta_estimator.predict(meta_features)
def score(self, X, y):
return accuracy_score(y, self.predict(X))
# Test feature stacking
feature_stacker = FeatureStackingClassifier(
base_estimators=[
('rf', RandomForestClassifier(n_estimators=50, random_state=42)),
('gb', GradientBoostingClassifier(n_estimators=50, random_state=42))
],
meta_estimator=LogisticRegression(),
use_features=True
)
feature_stacker.fit(X_small, y_small)
feature_stacking_score = feature_stacker.score(X_test, y_test)
print(f"Feature Stacking Accuracy: {feature_stacking_score:.3f}")
# 3. BLENDING
print("\n3. Blending (Holdout Stacking)")
print("-" * 40)
class BlendingClassifier:
"""
Blending uses a holdout validation set instead of cross-validation
"""
def __init__(self, base_estimators, meta_estimator, blend_size=0.2):
self.base_estimators = base_estimators
self.meta_estimator = meta_estimator
self.blend_size = blend_size
def fit(self, X, y):
# Split into blend and train
from sklearn.model_selection import train_test_split
X_blend, X_train, y_blend, y_train = train_test_split(
X, y, test_size=1-self.blend_size, random_state=42, stratify=y
)
# Train base models on train set
self.fitted_base_ = []
blend_features = []
for name, estimator in self.base_estimators:
# Fit on training portion
fitted_est = estimator.fit(X_train, y_train)
self.fitted_base_.append((name, fitted_est))
# Predict on blend portion
if hasattr(estimator, 'predict_proba'):
blend_pred = estimator.predict_proba(X_blend)
else:
blend_pred = estimator.predict(X_blend).reshape(-1, 1)
blend_features.append(blend_pred)
# Stack blend predictions
if blend_features[0].ndim == 2:
blend_features = np.hstack(blend_features)
else:
blend_features = np.column_stack(blend_features)
# Train meta model on blend predictions
self.meta_estimator.fit(blend_features, y_blend)
return self
def predict(self, X):
# Get base predictions
base_preds = []
for name, estimator in self.fitted_base_:
if hasattr(estimator, 'predict_proba'):
pred = estimator.predict_proba(X)
else:
pred = estimator.predict(X).reshape(-1, 1)
base_preds.append(pred)
# Stack predictions
if base_preds[0].ndim == 2:
meta_features = np.hstack(base_preds)
else:
meta_features = np.column_stack(base_preds)
return self.meta_estimator.predict(meta_features)
def score(self, X, y):
return accuracy_score(y, self.predict(X))
# Test blending
blender = BlendingClassifier(
base_estimators=[
('rf', RandomForestClassifier(n_estimators=50, random_state=42)),
('gb', GradientBoostingClassifier(n_estimators=50, random_state=42)),
('svc', SVC(probability=True, random_state=42))
],
meta_estimator=LogisticRegression(),
blend_size=0.2
)
blender.fit(X_train, y_train)
blending_score = blender.score(X_test, y_test)
print(f"Blending Accuracy: {blending_score:.3f}")
# Compare stacking methods
stacking_comparison = pd.DataFrame({
'Method': ['Multi-Level Stacking', 'Feature Stacking', 'Blending'],
'Accuracy': [multi_level_score, feature_stacking_score, blending_score]
})
print("\n" + stacking_comparison.to_string(index=False))
Advanced Boosting Strategies
from sklearn.ensemble import AdaBoostClassifier, GradientBoostingClassifier
import xgboost as xgb
import lightgbm as lgb
print("\n" + "="*60)
print("ADVANCED BOOSTING STRATEGIES")
print("="*60)
# 1. CUSTOM BOOSTING WITH SAMPLE WEIGHTS
print("\n1. Custom Adaptive Boosting")
print("-" * 40)
class CustomAdaBoost:
"""
Custom implementation of AdaBoost to understand the algorithm
"""
def __init__(self, base_estimator, n_estimators=50):
self.base_estimator = base_estimator
self.n_estimators = n_estimators
self.estimators_ = []
self.estimator_weights_ = []
def fit(self, X, y):
n_samples = X.shape[0]
# Initialize weights uniformly
sample_weights = np.ones(n_samples) / n_samples
for i in range(self.n_estimators):
# Train base estimator with current weights
estimator = self.base_estimator.__class__(**self.base_estimator.get_params())
estimator.fit(X, y, sample_weight=sample_weights)
# Make predictions
y_pred = estimator.predict(X)
# Calculate weighted error
incorrect = y_pred != y
error = np.sum(sample_weights[incorrect]) / np.sum(sample_weights)
# Calculate estimator weight (alpha)
if error >= 0.5:
# Stop if error is too high
if i == 0:
self.estimators_.append(estimator)
self.estimator_weights_.append(1.0)
break
alpha = 0.5 * np.log((1 - error) / error)
# Update sample weights
sample_weights *= np.exp(-alpha * y * (2 * y_pred - 1))
sample_weights /= np.sum(sample_weights)
# Store estimator and weight
self.estimators_.append(estimator)
self.estimator_weights_.append(alpha)
if i % 10 == 0:
print(f" Iteration {i}: Error = {error:.3f}, Alpha = {alpha:.3f}")
return self
def predict(self, X):
# Weighted majority vote
predictions = np.zeros(X.shape[0])
for estimator, weight in zip(self.estimators_, self.estimator_weights_):
predictions += weight * (2 * estimator.predict(X) - 1)
return (predictions > 0).astype(int)
def score(self, X, y):
return accuracy_score(y, self.predict(X))
# Test custom AdaBoost
from sklearn.tree import DecisionTreeClassifier
custom_ada = CustomAdaBoost(
base_estimator=DecisionTreeClassifier(max_depth=1),
n_estimators=50
)
custom_ada.fit(X_train, y_train)
custom_score = custom_ada.score(X_test, y_test)
print(f"\nCustom AdaBoost Accuracy: {custom_score:.3f}")
# Compare with sklearn AdaBoost
sklearn_ada = AdaBoostClassifier(n_estimators=50, random_state=42)
sklearn_ada.fit(X_train, y_train)
sklearn_score = sklearn_ada.score(X_test, y_test)
print(f"Sklearn AdaBoost Accuracy: {sklearn_score:.3f}")
# 2. GRADIENT BOOSTING VARIATIONS
print("\n2. Gradient Boosting Variations")
print("-" * 40)
# Standard Gradient Boosting
gb_standard = GradientBoostingClassifier(
n_estimators=100,
learning_rate=0.1,
max_depth=3,
random_state=42
)
# Stochastic Gradient Boosting (with subsampling)
gb_stochastic = GradientBoostingClassifier(
n_estimators=100,
learning_rate=0.1,
max_depth=3,
subsample=0.8, # Use 80% of samples
max_features='sqrt', # Use sqrt(n_features) at each split
random_state=42
)
# Histogram-based Gradient Boosting (faster)
from sklearn.ensemble import HistGradientBoostingClassifier
gb_hist = HistGradientBoostingClassifier(
max_iter=100,
learning_rate=0.1,
max_depth=3,
random_state=42
)
# Compare different GB variations
gb_models = {
'Standard GB': gb_standard,
'Stochastic GB': gb_stochastic,
'Histogram GB': gb_hist
}
gb_results = {}
for name, model in gb_models.items():
import time
start_time = time.time()
model.fit(X_train, y_train)
train_time = time.time() - start_time
score = model.score(X_test, y_test)
gb_results[name] = {'accuracy': score, 'time': train_time}
print(f"{name:15} - Accuracy: {score:.3f}, Time: {train_time:.2f}s")
# 3. EXTREME GRADIENT BOOSTING (XGBoost)
print("\n3. XGBoost Advanced Features")
print("-" * 40)
# Prepare data for XGBoost
dtrain = xgb.DMatrix(X_train, label=y_train)
dtest = xgb.DMatrix(X_test, label=y_test)
# XGBoost with early stopping
params = {
'objective': 'binary:logistic',
'max_depth': 3,
'learning_rate': 0.1,
'n_estimators': 1000,
'eval_metric': 'logloss'
}
# Train with early stopping
evals = [(dtrain, 'train'), (dtest, 'eval')]
xgb_model = xgb.train(
params,
dtrain,
num_boost_round=1000,
evals=evals,
early_stopping_rounds=10,
verbose=False
)
# Get predictions
xgb_pred = (xgb_model.predict(dtest) > 0.5).astype(int)
xgb_score = accuracy_score(y_test, xgb_pred)
print(f"XGBoost with Early Stopping:")
print(f" Best iteration: {xgb_model.best_iteration}")
print(f" Accuracy: {xgb_score:.3f}")
# Feature importance
importance = xgb_model.get_score(importance_type='gain')
print(f" Top 5 important features: {list(importance.keys())[:5]}")
# 4. LIGHTGBM
print("\n4. LightGBM (Gradient Boosting on Leaves)")
print("-" * 40)
# Prepare LightGBM dataset
lgb_train = lgb.Dataset(X_train, y_train)
lgb_eval = lgb.Dataset(X_test, y_test, reference=lgb_train)
# LightGBM parameters
lgb_params = {
'objective': 'binary',
'metric': 'binary_logloss',
'num_leaves': 31,
'learning_rate': 0.1,
'feature_fraction': 0.9,
'bagging_fraction': 0.8,
'bagging_freq': 5,
'verbose': -1
}
# Train LightGBM
lgb_model = lgb.train(
lgb_params,
lgb_train,
valid_sets=[lgb_eval],
num_boost_round=100,
callbacks=[lgb.early_stopping(10), lgb.log_evaluation(0)]
)
# Predictions
lgb_pred = (lgb_model.predict(X_test) > 0.5).astype(int)
lgb_score = accuracy_score(y_test, lgb_pred)
print(f"LightGBM Accuracy: {lgb_score:.3f}")
print(f"Number of trees: {lgb_model.num_trees()}")
Ensemble Selection and Optimization
print("\n" + "="*60)
print("ENSEMBLE SELECTION AND OPTIMIZATION")
print("="*60)
# 1. GREEDY ENSEMBLE SELECTION
print("\n1. Greedy Ensemble Selection")
print("-" * 40)
class GreedyEnsembleSelector:
"""
Greedy algorithm to select best subset of models
"""
def __init__(self, models, metric=accuracy_score, n_best=None):
self.models = models
self.metric = metric
self.n_best = n_best
self.selected_models_ = []
self.selection_path_ = []
def fit(self, X, y):
"""Select best ensemble greedily"""
from sklearn.model_selection import train_test_split
# Split for validation
X_train, X_val, y_train, y_val = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
# Fit all models
fitted_models = []
for name, model in self.models:
model.fit(X_train, y_train)
fitted_models.append((name, model))
# Get predictions for all models
predictions = {}
for name, model in fitted_models:
predictions[name] = model.predict(X_val)
# Greedy selection
available_models = list(predictions.keys())
selected = []
n_to_select = self.n_best or len(available_models)
for i in range(n_to_select):
best_score = -np.inf
best_model = None
for model in available_models:
# Try adding this model
trial_ensemble = selected + [model]
# Ensemble prediction (majority vote)
ensemble_pred = np.zeros(len(y_val))
for m in trial_ensemble:
ensemble_pred += predictions[m]
ensemble_pred = (ensemble_pred > len(trial_ensemble) / 2).astype(int)
# Calculate score
score = self.metric(y_val, ensemble_pred)
if score > best_score:
best_score = score
best_model = model
# Add best model if it improves score
if best_model and (not selected or best_score > self.selection_path_[-1][1]):
selected.append(best_model)
available_models.remove(best_model)
self.selection_path_.append((best_model, best_score))
print(f" Step {i+1}: Added {best_model}, Score: {best_score:.3f}")
else:
break
# Store selected models
self.selected_models_ = [(name, model) for name, model in fitted_models
if name in selected]
return self
def predict(self, X):
"""Predict using selected ensemble"""
predictions = np.zeros(X.shape[0])
for name, model in self.selected_models_:
predictions += model.predict(X)
return (predictions > len(self.selected_models_) / 2).astype(int)
def score(self, X, y):
return accuracy_score(y, self.predict(X))
# Test greedy selection
candidate_models = [
('lr', LogisticRegression(random_state=42)),
('svc', SVC(random_state=42)),
('rf', RandomForestClassifier(n_estimators=50, random_state=42)),
('gb', GradientBoostingClassifier(n_estimators=50, random_state=42)),
('nb', GaussianNB()),
('knn', KNeighborsClassifier()),
('dt', DecisionTreeClassifier(max_depth=5, random_state=42))
]
greedy_selector = GreedyEnsembleSelector(candidate_models, n_best=5)
greedy_selector.fit(X_train, y_train)
print(f"\nSelected {len(greedy_selector.selected_models_)} models")
greedy_score = greedy_selector.score(X_test, y_test)
print(f"Greedy Ensemble Accuracy: {greedy_score:.3f}")
# 2. BAYESIAN OPTIMIZATION FOR ENSEMBLE WEIGHTS
print("\n2. Bayesian Optimization for Ensemble Weights")
print("-" * 40)
from scipy.optimize import minimize
from sklearn.metrics import log_loss
class BayesianEnsembleOptimizer:
"""
Find optimal weights using Bayesian optimization
"""
def __init__(self, models):
self.models = models
self.weights_ = None
def objective(self, weights, X, y, predictions):
"""Objective function to minimize"""
# Normalize weights
weights = weights / weights.sum()
# Weighted average of predictions
ensemble_pred = np.zeros_like(predictions[0])
for i, pred in enumerate(predictions):
ensemble_pred += weights[i] * pred
# Calculate negative log loss (to minimize)
return log_loss(y, ensemble_pred)
def fit(self, X, y):
"""Find optimal weights"""
from sklearn.model_selection import train_test_split
# Split for validation
X_train, X_val, y_train, y_val = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
# Fit models and get predictions
predictions = []
self.fitted_models_ = []
for name, model in self.models:
model.fit(X_train, y_train)
self.fitted_models_.append((name, model))
if hasattr(model, 'predict_proba'):
pred = model.predict_proba(X_val)[:, 1]
else:
pred = model.predict(X_val)
predictions.append(pred)
# Initial weights (uniform)
n_models = len(self.models)
initial_weights = np.ones(n_models) / n_models
# Constraints: weights sum to 1, all non-negative
constraints = {'type': 'eq', 'fun': lambda w: w.sum() - 1}
bounds = [(0, 1)] * n_models
# Optimize
result = minimize(
self.objective,
initial_weights,
args=(X_val, y_val, predictions),
method='SLSQP',
bounds=bounds,
constraints=constraints
)
self.weights_ = result.x / result.x.sum()
print("Optimized weights:")
for (name, _), weight in zip(self.models, self.weights_):
print(f" {name}: {weight:.3f}")
return self
def predict_proba(self, X):
"""Predict probabilities"""
predictions = []
for name, model in self.fitted_models_:
if hasattr(model, 'predict_proba'):
pred = model.predict_proba(X)[:, 1]
else:
pred = model.predict(X)
predictions.append(pred)
# Weighted average
ensemble_pred = np.zeros_like(predictions[0])
for i, pred in enumerate(predictions):
ensemble_pred += self.weights_[i] * pred
return ensemble_pred
def predict(self, X):
return (self.predict_proba(X) > 0.5).astype(int)
def score(self, X, y):
return accuracy_score(y, self.predict(X))
# Test Bayesian optimization
bayesian_optimizer = BayesianEnsembleOptimizer(candidate_models[:4])
bayesian_optimizer.fit(X_train, y_train)
bayesian_score = bayesian_optimizer.score(X_test, y_test)
print(f"\nBayesian Optimized Ensemble Accuracy: {bayesian_score:.3f}")
Production-Ready Ensemble System
print("\n" + "="*60)
print("PRODUCTION-READY ENSEMBLE SYSTEM")
print("="*60)
class ProductionEnsemble:
"""
Complete ensemble system for production use
"""
def __init__(self,
ensemble_type='voting',
base_models=None,
cv_folds=5,
optimize_weights=True,
use_calibration=True,
save_models=True):
self.ensemble_type = ensemble_type
self.base_models = base_models or self._get_default_models()
self.cv_folds = cv_folds
self.optimize_weights = optimize_weights
self.use_calibration = use_calibration
self.save_models = save_models
self.fitted_models_ = []
self.weights_ = None
self.calibrators_ = []
self.performance_metrics_ = {}
def _get_default_models(self):
"""Default diverse model set"""
return [
('rf', RandomForestClassifier(n_estimators=100, random_state=42)),
('gb', GradientBoostingClassifier(n_estimators=100, random_state=42)),
('lr', LogisticRegression(max_iter=1000, random_state=42)),
('svc', SVC(probability=True, random_state=42))
]
def _calibrate_probabilities(self, X, y):
"""Calibrate probability predictions"""
from sklearn.calibration import CalibratedClassifierCV
calibrated_models = []
for name, model in self.fitted_models_:
if hasattr(model, 'predict_proba'):
calibrator = CalibratedClassifierCV(
model, method='sigmoid', cv=3
)
calibrator.fit(X, y)
calibrated_models.append((name, calibrator))
else:
calibrated_models.append((name, model))
return calibrated_models
def fit(self, X, y):
"""Fit the production ensemble"""
print("Training Production Ensemble...")
print("-" * 40)
# Step 1: Train base models
print("Step 1: Training base models...")
for name, model in self.base_models:
print(f" Training {name}...")
model.fit(X, y)
# Cross-validation score
cv_scores = cross_val_score(model, X, y, cv=self.cv_folds)
self.performance_metrics_[name] = {
'cv_mean': cv_scores.mean(),
'cv_std': cv_scores.std()
}
self.fitted_models_.append((name, model))
# Step 2: Calibrate probabilities if requested
if self.use_calibration:
print("\nStep 2: Calibrating probabilities...")
self.fitted_models_ = self._calibrate_probabilities(X, y)
# Step 3: Optimize weights if requested
if self.optimize_weights:
print("\nStep 3: Optimizing ensemble weights...")
self._optimize_weights(X, y)
else:
# Use CV scores as weights
scores = [self.performance_metrics_[name]['cv_mean']
for name, _ in self.fitted_models_]
self.weights_ = np.array(scores) / np.sum(scores)
# Step 4: Calculate final metrics
self._calculate_final_metrics(X, y)
print("\nβ
Ensemble training complete!")
return self
def _optimize_weights(self, X, y):
"""Optimize ensemble weights"""
# Similar to BayesianEnsembleOptimizer
# Implementation omitted for brevity
n_models = len(self.fitted_models_)
self.weights_ = np.ones(n_models) / n_models
def _calculate_final_metrics(self, X, y):
"""Calculate and store performance metrics"""
ensemble_pred = self.predict(X)
self.performance_metrics_['ensemble'] = {
'accuracy': accuracy_score(y, ensemble_pred),
'models_used': len(self.fitted_models_),
'weights': self.weights_.tolist()
}
def predict_proba(self, X):
"""Predict probabilities"""
probas = []
for (name, model), weight in zip(self.fitted_models_, self.weights_):
if hasattr(model, 'predict_proba'):
proba = model.predict_proba(X)[:, 1]
else:
proba = model.predict(X)
probas.append(proba * weight)
return np.sum(probas, axis=0)
def predict(self, X):
"""Predict classes"""
if self.ensemble_type == 'voting':
return (self.predict_proba(X) > 0.5).astype(int)
else:
# Other ensemble types can be implemented here
return (self.predict_proba(X) > 0.5).astype(int)
def score(self, X, y):
"""Calculate accuracy"""
return accuracy_score(y, self.predict(X))
def get_model_importance(self):
"""Get importance of each model in ensemble"""
importance = pd.DataFrame({
'Model': [name for name, _ in self.fitted_models_],
'Weight': self.weights_,
'CV_Score': [self.performance_metrics_[name]['cv_mean']
for name, _ in self.fitted_models_]
})
return importance.sort_values('Weight', ascending=False)
def save(self, filepath):
"""Save ensemble to disk"""
import joblib
ensemble_data = {
'models': self.fitted_models_,
'weights': self.weights_,
'metrics': self.performance_metrics_,
'config': {
'ensemble_type': self.ensemble_type,
'cv_folds': self.cv_folds,
'use_calibration': self.use_calibration
}
}
joblib.dump(ensemble_data, filepath)
print(f"Ensemble saved to {filepath}")
@classmethod
def load(cls, filepath):
"""Load ensemble from disk"""
import joblib
ensemble_data = joblib.load(filepath)
# Create new instance
ensemble = cls(
ensemble_type=ensemble_data['config']['ensemble_type'],
cv_folds=ensemble_data['config']['cv_folds'],
use_calibration=ensemble_data['config']['use_calibration']
)
# Restore state
ensemble.fitted_models_ = ensemble_data['models']
ensemble.weights_ = ensemble_data['weights']
ensemble.performance_metrics_ = ensemble_data['metrics']
return ensemble
# Test production ensemble
production_ensemble = ProductionEnsemble(
ensemble_type='voting',
optimize_weights=False,
use_calibration=False # Set to False for speed
)
production_ensemble.fit(X_train, y_train)
print("\nModel Importance:")
print(production_ensemble.get_model_importance())
print(f"\nProduction Ensemble Test Accuracy: {production_ensemble.score(X_test, y_test):.3f}")
# Performance comparison summary
print("\n" + "="*60)
print("ENSEMBLE METHODS PERFORMANCE SUMMARY")
print("="*60)
all_results = {
'Hard Voting': hard_score,
'Soft Voting': soft_score,
'Weighted Voting': weighted_score,
'Dynamic Voting': dynamic_score,
'Feature Stacking': feature_stacking_score,
'Blending': blending_score,
'Greedy Selection': greedy_score,
'Bayesian Optimized': bayesian_score,
'Production Ensemble': production_ensemble.score(X_test, y_test)
}
results_df = pd.DataFrame.from_dict(all_results, orient='index', columns=['Accuracy'])
results_df = results_df.sort_values('Accuracy', ascending=False)
print(results_df.to_string())
print(f"\nBest Method: {results_df.index[0]} with {results_df.iloc[0, 0]:.3f} accuracy")
Practice Exercises
Exercise 1: Custom Ensemble Framework
Build a flexible ensemble framework that:
- Supports multiple ensemble strategies (voting, stacking, blending)
- Automatically selects best base models
- Optimizes ensemble weights
- Provides confidence intervals for predictions
- Handles both classification and regression
Exercise 2: Ensemble Diversity Analysis
Create tools to analyze ensemble diversity:
- Calculate pairwise model correlation
- Measure diversity metrics (Q-statistic, disagreement measure)
- Visualize model agreement/disagreement patterns
- Select diverse models for optimal ensemble
Exercise 3: Competition-Grade Ensemble
Build a Kaggle-competition-ready ensemble:
- Implement multi-level stacking
- Use out-of-fold predictions
- Include feature engineering in the pipeline
- Implement cross-validation strategy
- Create submission file generator
Key Takeaways
- π― Ensemble methods combine multiple models to achieve better performance
- π³οΈ Voting ensembles: Simple but effective (hard, soft, weighted)
- π Stacking: Use meta-learner to combine base model predictions
- π Blending: Similar to stacking but uses holdout validation
- β‘ Boosting: Sequential learning with focus on errors
- π¨ Diversity is key: Models should make different errors
- βοΈ Weight optimization can significantly improve performance
- ποΈ Multi-level stacking can capture complex patterns
- π Greedy selection finds optimal model subset
- π Production ensembles need calibration and monitoring
Summary
Advanced ensemble methods represent the pinnacle of classical machine learning. By intelligently combining diverse models, we can achieve performance that surpasses any individual algorithm. The techniques covered hereβfrom sophisticated voting schemes to multi-level stackingβare the same methods that win machine learning competitions and power production systems. Remember: the key to successful ensembles is diversity, proper validation, and careful optimization. Master these techniques, and you'll have the tools to tackle the most challenging machine learning problems!
π Learning Journal
Keep a learning journal β digital or physical. After this lesson, take a few minutes to write down:
- Key concepts you learned
- Techniques that clicked for you
- Questions or confusion points to revisit
- Ideas you want to try
- Your progress and feelings about learning this
βοΈ This lesson's prompt: The key to a strong ensemble is diversity among its members. Where in your ensemble did diversity clearly help β and where did adding a similar model add cost without any real benefit?
π Lesson Summary
π Key Takeaways
- Ensembles combine multiple models to outperform any single one β and diversity among members is the key ingredient.
- Bagging (e.g. random forests) reduces variance by averaging parallel models; boosting reduces bias by training models sequentially.
- Voting averages or majority-votes base predictions; stacking trains a meta-model to combine them.
- More models means more compute and less interpretability β validate carefully and watch for overfitting the meta-learner.
π What You've Accomplished
You can now combine diverse learners into voting and stacking ensembles that rival the techniques used in competition-winning and production machine learning systems.
β Common Questions at This Stage
What's the difference between bagging and boosting?
Bagging trains models in parallel on bootstrap samples to cut variance. Boosting trains models sequentially, each one focusing on the errors of the last, to cut bias.
What makes stacking work?
A meta-model learns how best to combine the base models' outputs. To avoid leakage, train the meta-model on cross-validated (out-of-fold) predictions rather than in-sample ones.
Are ensembles always the better choice?
They're usually more accurate but slower and harder to interpret. Sometimes a single well-tuned model is the better trade-off β measure both before deciding.
π Looking Ahead
Next you'll learn how to interpret and explain these more complex models, so their added power doesn't come at the cost of being a black box.
β Before the Next Lesson
- Build a
VotingClassifierand aStackingClassifierand compare both to the best base model. - Deliberately add a model from a different algorithm family and measure the effect of the added diversity.
- Write your Learning Journal entry for this lesson.
π Encouragement for the Journey
You now wield the same techniques that top competitors and production teams rely on. Combine your models wisely and they'll punch well above their weight β keep experimenting!