Skip to main content

🎯 Feature Selection: Choosing the Right Features

📚 What You'll Learn

By the end of this lesson, you will be able to:

  • Distinguish feature selection (keeping original features) from feature extraction like PCA, and explain why selection preserves interpretability
  • Apply filter methods such as variance threshold, correlation, the ANOVA F-test, chi-square, and mutual information
  • Use wrapper methods including recursive feature elimination (RFE) and forward/backward selection
  • Use embedded methods such as L1 (Lasso) regularization and tree-based feature importances
  • Choose an appropriate method based on dataset size, model type, and interpretability needs

⏱️ Estimated Time: 45–60 minutes

🎯 Project: Reduce a high-dimensional dataset to its most predictive features and compare model performance before and after selection.

Introduction

Feature selection is the process of selecting a subset of relevant features for model construction. Unlike feature extraction methods like PCA that create new features, feature selection keeps the original features intact, making the model more interpretable. This lesson covers filter methods, wrapper methods, and embedded methods for feature selection, along with practical guidelines for choosing the right approach.

Feature Selection Overview

flowchart TD A[All Features] --> B{Selection Method} B --> C[Filter Methods] C --> C1[Correlation] C --> C2[Chi-Square] C --> C3[Mutual Information] C --> C4[Variance Threshold] B --> D[Wrapper Methods] D --> D1[Forward Selection] D --> D2[Backward Elimination] D --> D3[Recursive Feature Elimination] B --> E[Embedded Methods] E --> E1[LASSO L1] E --> E2[Ridge L2] E --> E3[Tree Importance] E --> E4[ElasticNet] C --> F[Selected Features] D --> F E --> F F --> G[Improved Model] style A fill:#e3f2fd style F fill:#fff9c4 style G fill:#c8e6c9

Why Feature Selection?

graph LR A[Feature Selection Benefits] --> B[Performance] A --> C[Interpretability] A --> D[Training Time] A --> E[Overfitting] B --> B1[Remove Noise] B --> B2[Reduce Complexity] C --> C1[Understand Important Factors] C --> C2[Business Insights] D --> D1[Fewer Computations] D --> D2[Faster Training] E --> E1[Better Generalization] E --> E2[Simpler Models] style A fill:#f0f4c3

Setting Up the Environment

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.datasets import make_classification, load_breast_cancer
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.feature_selection import (
    SelectKBest, f_classif, chi2, mutual_info_classif,
    RFE, RFECV, SelectFromModel, VarianceThreshold
)
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression, Lasso
from sklearn.metrics import accuracy_score, classification_report
import warnings
warnings.filterwarnings('ignore')

# Set style for visualizations
plt.style.use('seaborn-v0_8-darkgrid')
sns.set_palette("husl")

# Set random seed
np.random.seed(42)

print("Feature Selection Methods")
print("=" * 50)

Filter Methods

class FilterMethods:
    """
    Filter methods for feature selection
    """
    
    def __init__(self):
        self.scores = {}
        
    def generate_sample_data(self, n_samples=1000, n_features=20, n_informative=10):
        """
        Generate sample data with informative and noise features
        """
        X, y = make_classification(
            n_samples=n_samples,
            n_features=n_features,
            n_informative=n_informative,
            n_redundant=5,
            n_repeated=0,
            n_classes=2,
            random_state=42,
            shuffle=False
        )
        
        # Create feature names
        feature_names = [f'Feature_{i}' for i in range(n_features)]
        
        # Create DataFrame
        df = pd.DataFrame(X, columns=feature_names)
        df['target'] = y
        
        print(f"Generated dataset: {n_samples} samples, {n_features} features")
        print(f"Informative features: {n_informative}")
        print(f"Redundant features: 5")
        print(f"Noise features: {n_features - n_informative - 5}")
        
        return df, feature_names
    
    def variance_threshold_selection(self, X, threshold=0.1):
        """
        Remove features with low variance
        """
        selector = VarianceThreshold(threshold=threshold)
        X_selected = selector.fit_transform(X)
        
        # Get selected features
        selected_features = X.columns[selector.get_support()]
        removed_features = X.columns[~selector.get_support()]
        
        print(f"\nVariance Threshold (threshold={threshold}):")
        print(f"Selected features: {len(selected_features)}/{len(X.columns)}")
        print(f"Removed low-variance features: {list(removed_features)[:5]}...")
        
        return selected_features, selector
    
    def correlation_selection(self, df, threshold=0.9):
        """
        Remove highly correlated features
        """
        # Calculate correlation matrix
        corr_matrix = df.corr().abs()
        
        # Select upper triangle of correlation matrix
        upper = corr_matrix.where(
            np.triu(np.ones(corr_matrix.shape), k=1).astype(bool)
        )
        
        # Find features with correlation greater than threshold
        to_drop = [column for column in upper.columns 
                  if any(upper[column] > threshold)]
        
        selected_features = [col for col in df.columns if col not in to_drop]
        
        print(f"\nCorrelation-based Selection (threshold={threshold}):")
        print(f"Removed highly correlated features: {to_drop[:5]}...")
        print(f"Selected features: {len(selected_features)}/{len(df.columns)}")
        
        return selected_features, corr_matrix
    
    def univariate_selection(self, X, y, k=10, method='f_classif'):
        """
        Select K best features using statistical tests
        """
        # Choose scoring function
        if method == 'f_classif':
            score_func = f_classif
        elif method == 'chi2':
            # Ensure non-negative values for chi2
            X_positive = X - X.min() + 1
            score_func = chi2
            X = X_positive
        elif method == 'mutual_info':
            score_func = mutual_info_classif
        else:
            raise ValueError(f"Unknown method: {method}")
        
        # Select K best features
        selector = SelectKBest(score_func=score_func, k=k)
        X_selected = selector.fit_transform(X, y)
        
        # Get scores and selected features
        scores = selector.scores_
        selected_features = X.columns[selector.get_support()]
        
        self.scores[method] = pd.DataFrame({
            'feature': X.columns,
            'score': scores
        }).sort_values('score', ascending=False)
        
        print(f"\nUnivariate Selection ({method}, k={k}):")
        print(f"Top features: {list(selected_features[:5])}...")
        
        return selected_features, selector
    
    def visualize_filter_methods(self, df):
        """
        Visualize results from different filter methods
        """
        X = df.drop('target', axis=1)
        y = df['target']
        
        fig, axes = plt.subplots(2, 2, figsize=(14, 10))
        
        # 1. Feature Variance
        variances = X.var()
        axes[0, 0].bar(range(len(variances)), variances.sort_values(ascending=False))
        axes[0, 0].axhline(y=0.1, color='r', linestyle='--', label='Threshold')
        axes[0, 0].set_xlabel('Feature Index')
        axes[0, 0].set_ylabel('Variance')
        axes[0, 0].set_title('Feature Variance')
        axes[0, 0].legend()
        axes[0, 0].grid(True, alpha=0.3)
        
        # 2. Correlation Heatmap
        corr_matrix = X.corr().abs()
        sns.heatmap(corr_matrix, cmap='coolwarm', center=0, 
                   square=True, ax=axes[0, 1], cbar_kws={'label': 'Correlation'})
        axes[0, 1].set_title('Feature Correlation Matrix')
        
        # 3. F-statistic scores
        _, f_selector = self.univariate_selection(X, y, k=10, method='f_classif')
        f_scores = f_selector.scores_
        axes[1, 0].bar(range(len(f_scores)), np.sort(f_scores)[::-1])
        axes[1, 0].set_xlabel('Feature Rank')
        axes[1, 0].set_ylabel('F-statistic Score')
        axes[1, 0].set_title('F-statistic Feature Scores')
        axes[1, 0].grid(True, alpha=0.3)
        
        # 4. Mutual Information scores
        _, mi_selector = self.univariate_selection(X, y, k=10, method='mutual_info')
        mi_scores = mi_selector.scores_
        axes[1, 1].bar(range(len(mi_scores)), np.sort(mi_scores)[::-1])
        axes[1, 1].set_xlabel('Feature Rank')
        axes[1, 1].set_ylabel('Mutual Information')
        axes[1, 1].set_title('Mutual Information Feature Scores')
        axes[1, 1].grid(True, alpha=0.3)
        
        plt.suptitle('Filter Methods for Feature Selection', fontsize=14)
        plt.tight_layout()
        plt.show()

# Demonstrate filter methods
filter_methods = FilterMethods()

# Generate data
df, feature_names = filter_methods.generate_sample_data()

# Apply filter methods
X = df.drop('target', axis=1)
y = df['target']

# Variance threshold
var_features, _ = filter_methods.variance_threshold_selection(X, threshold=0.1)

# Correlation-based selection
corr_features, _ = filter_methods.correlation_selection(X, threshold=0.9)

# Univariate selection
f_features, _ = filter_methods.univariate_selection(X, y, k=10, method='f_classif')
mi_features, _ = filter_methods.univariate_selection(X, y, k=10, method='mutual_info')

# Visualize
filter_methods.visualize_filter_methods(df)

Wrapper Methods

class WrapperMethods:
    """
    Wrapper methods for feature selection
    """
    
    def __init__(self):
        self.selected_features = {}
        
    def forward_selection(self, X, y, estimator, max_features=10):
        """
        Forward feature selection
        """
        n_features = X.shape[1]
        selected = []
        remaining = list(range(n_features))
        scores = []
        
        # Iteratively add features
        for i in range(min(max_features, n_features)):
            best_score = -np.inf
            best_feature = None
            
            for feature in remaining:
                # Try adding this feature
                trial_features = selected + [feature]
                X_subset = X[:, trial_features]
                
                # Evaluate with cross-validation
                cv_scores = cross_val_score(
                    estimator, X_subset, y, cv=3, scoring='accuracy'
                )
                score = cv_scores.mean()
                
                if score > best_score:
                    best_score = score
                    best_feature = feature
            
            if best_feature is not None:
                selected.append(best_feature)
                remaining.remove(best_feature)
                scores.append(best_score)
                
        return selected, scores
    
    def backward_elimination(self, X, y, estimator, min_features=1):
        """
        Backward feature elimination
        """
        n_features = X.shape[1]
        selected = list(range(n_features))
        scores = []
        
        # Iteratively remove features
        while len(selected) > min_features:
            worst_score = np.inf
            worst_feature = None
            
            for feature in selected:
                # Try removing this feature
                trial_features = [f for f in selected if f != feature]
                X_subset = X[:, trial_features]
                
                # Evaluate with cross-validation
                cv_scores = cross_val_score(
                    estimator, X_subset, y, cv=3, scoring='accuracy'
                )
                score = -cv_scores.mean()  # Negative for minimization
                
                if score < worst_score:
                    worst_score = score
                    worst_feature = feature
            
            if worst_feature is not None:
                selected.remove(worst_feature)
                scores.append(-worst_score)  # Convert back to positive
                
        return selected, scores
    
    def recursive_feature_elimination(self, X, y, estimator, n_features_to_select=10):
        """
        Recursive Feature Elimination (RFE)
        """
        # RFE with specified number of features
        rfe = RFE(estimator, n_features_to_select=n_features_to_select)
        rfe.fit(X, y)
        
        # Get selected features
        selected_features = X.columns[rfe.support_]
        feature_ranking = pd.DataFrame({
            'feature': X.columns,
            'ranking': rfe.ranking_
        }).sort_values('ranking')
        
        print(f"\nRFE Selected Features ({n_features_to_select}):")
        print(f"Top features: {list(selected_features[:5])}...")
        
        return selected_features, feature_ranking, rfe
    
    def rfecv_selection(self, X, y, estimator):
        """
        RFE with Cross-Validation to find optimal number of features
        """
        rfecv = RFECV(
            estimator,
            step=1,
            cv=5,
            scoring='accuracy',
            n_jobs=-1
        )
        rfecv.fit(X, y)
        
        # Get results
        optimal_features = X.columns[rfecv.support_]
        n_optimal = rfecv.n_features_
        
        print(f"\nRFECV Optimal Features: {n_optimal}")
        print(f"Selected: {list(optimal_features[:5])}...")
        
        return optimal_features, rfecv
    
    def compare_wrapper_methods(self, X, y):
        """
        Compare different wrapper methods
        """
        # Use simple estimator for speed
        estimator = LogisticRegression(random_state=42, max_iter=100)
        
        # Standardize features
        scaler = StandardScaler()
        X_scaled = scaler.fit_transform(X)
        
        print("\nComparing Wrapper Methods...")
        print("-" * 40)
        
        # Forward Selection (limited for speed)
        print("Running Forward Selection...")
        forward_features, forward_scores = self.forward_selection(
            X_scaled, y, estimator, max_features=10
        )
        
        # RFE
        print("Running RFE...")
        rfe_features, rfe_ranking, rfe = self.recursive_feature_elimination(
            X, y, estimator, n_features_to_select=10
        )
        
        # RFECV
        print("Running RFECV...")
        rfecv_features, rfecv = self.rfecv_selection(X, y, estimator)
        
        # Visualization
        fig, axes = plt.subplots(1, 3, figsize=(15, 5))
        
        # Forward Selection scores
        axes[0].plot(range(1, len(forward_scores) + 1), forward_scores, 'bo-')
        axes[0].set_xlabel('Number of Features')
        axes[0].set_ylabel('CV Accuracy')
        axes[0].set_title('Forward Selection')
        axes[0].grid(True, alpha=0.3)
        
        # RFE Ranking
        axes[1].barh(range(len(rfe_ranking[:15])), 
                    rfe_ranking['ranking'].values[:15])
        axes[1].set_yticks(range(len(rfe_ranking[:15])))
        axes[1].set_yticklabels(rfe_ranking['feature'].values[:15])
        axes[1].set_xlabel('Ranking')
        axes[1].set_title('RFE Feature Ranking (lower is better)')
        axes[1].invert_yaxis()
        
        # RFECV scores
        axes[2].plot(range(1, len(rfecv.cv_results_['mean_test_score']) + 1),
                    rfecv.cv_results_['mean_test_score'], 'go-')
        axes[2].axvline(x=rfecv.n_features_, color='r', linestyle='--',
                       label=f'Optimal: {rfecv.n_features_}')
        axes[2].set_xlabel('Number of Features')
        axes[2].set_ylabel('CV Accuracy')
        axes[2].set_title('RFECV: Optimal Features')
        axes[2].legend()
        axes[2].grid(True, alpha=0.3)
        
        plt.suptitle('Wrapper Methods Comparison', fontsize=14)
        plt.tight_layout()
        plt.show()
        
        return {
            'forward': forward_features,
            'rfe': rfe_features,
            'rfecv': rfecv_features
        }

# Demonstrate wrapper methods
wrapper_methods = WrapperMethods()

print("\n" + "=" * 50)
print("WRAPPER METHODS")
print("=" * 50)

# Compare wrapper methods
wrapper_results = wrapper_methods.compare_wrapper_methods(X, y)

Embedded Methods

flowchart TB A[Training Data] --> B[Model with Built-in Selection] B --> C[L1 Regularization
LASSO] B --> D[L2 Regularization
Ridge] B --> E[Tree-based
Feature Importance] B --> F[ElasticNet
L1 + L2] C --> G[Sparse Coefficients] D --> H[Shrunk Coefficients] E --> I[Importance Scores] F --> J[Balanced Selection] G --> K[Selected Features] I --> K J --> K style A fill:#e3f2fd style K fill:#c8e6c9
class EmbeddedMethods:
    """
    Embedded methods for feature selection
    """
    
    def __init__(self):
        self.models = {}
        
    def lasso_selection(self, X, y, alpha=0.01):
        """
        Feature selection using LASSO (L1 regularization)
        """
        # Standardize features
        scaler = StandardScaler()
        X_scaled = scaler.fit_transform(X)
        
        # Fit LASSO
        lasso = Lasso(alpha=alpha, random_state=42)
        lasso.fit(X_scaled, y)
        
        # Get selected features (non-zero coefficients)
        selected_features = X.columns[lasso.coef_ != 0]
        
        # Create importance DataFrame
        importance = pd.DataFrame({
            'feature': X.columns,
            'coefficient': lasso.coef_,
            'abs_coefficient': np.abs(lasso.coef_)
        }).sort_values('abs_coefficient', ascending=False)
        
        print(f"\nLASSO Selection (alpha={alpha}):")
        print(f"Selected features: {len(selected_features)}/{len(X.columns)}")
        print(f"Top features: {list(selected_features[:5])}...")
        
        return selected_features, importance, lasso
    
    def tree_based_selection(self, X, y, n_estimators=100):
        """
        Feature selection using Random Forest importance
        """
        # Train Random Forest
        rf = RandomForestClassifier(
            n_estimators=n_estimators,
            random_state=42,
            n_jobs=-1
        )
        rf.fit(X, y)
        
        # Get feature importances
        importance = pd.DataFrame({
            'feature': X.columns,
            'importance': rf.feature_importances_
        }).sort_values('importance', ascending=False)
        
        # Select features above mean importance
        threshold = rf.feature_importances_.mean()
        selected_features = X.columns[rf.feature_importances_ > threshold]
        
        print(f"\nTree-based Selection:")
        print(f"Selected features (importance > mean): {len(selected_features)}/{len(X.columns)}")
        print(f"Top features: {list(importance['feature'].head())}...")
        
        return selected_features, importance, rf
    
    def select_from_model(self, X, y, estimator, threshold='mean'):
        """
        Generic selection from any model with feature_importances_ or coef_
        """
        selector = SelectFromModel(estimator, threshold=threshold)
        selector.fit(X, y)
        
        # Get selected features
        selected_features = X.columns[selector.get_support()]
        
        print(f"\nSelectFromModel ({estimator.__class__.__name__}):")
        print(f"Selected features: {len(selected_features)}/{len(X.columns)}")
        
        return selected_features, selector
    
    def compare_embedded_methods(self, X, y):
        """
        Compare different embedded methods
        """
        # Standardize for linear models
        scaler = StandardScaler()
        X_scaled = scaler.fit_transform(X)
        X_scaled_df = pd.DataFrame(X_scaled, columns=X.columns)
        
        # LASSO
        lasso_features, lasso_importance, _ = self.lasso_selection(X, y, alpha=0.01)
        
        # Random Forest
        rf_features, rf_importance, _ = self.tree_based_selection(X, y)
        
        # Logistic Regression with L1
        lr_l1 = LogisticRegression(penalty='l1', solver='liblinear', random_state=42)
        lr_features, _ = self.select_from_model(X_scaled_df, y, lr_l1)
        
        # Visualization
        fig, axes = plt.subplots(2, 2, figsize=(14, 10))
        
        # LASSO coefficients
        top_lasso = lasso_importance.head(15)
        axes[0, 0].barh(range(len(top_lasso)), top_lasso['abs_coefficient'])
        axes[0, 0].set_yticks(range(len(top_lasso)))
        axes[0, 0].set_yticklabels(top_lasso['feature'])
        axes[0, 0].set_xlabel('|Coefficient|')
        axes[0, 0].set_title('LASSO Feature Importance')
        axes[0, 0].invert_yaxis()
        
        # Random Forest importance
        top_rf = rf_importance.head(15)
        axes[0, 1].barh(range(len(top_rf)), top_rf['importance'])
        axes[0, 1].set_yticks(range(len(top_rf)))
        axes[0, 1].set_yticklabels(top_rf['feature'])
        axes[0, 1].set_xlabel('Importance')
        axes[0, 1].set_title('Random Forest Feature Importance')
        axes[0, 1].invert_yaxis()
        
        # Venn diagram of selected features (simplified)
        axes[1, 0].bar(['LASSO', 'Random Forest', 'Logistic L1'],
                      [len(lasso_features), len(rf_features), len(lr_features)])
        axes[1, 0].set_ylabel('Number of Selected Features')
        axes[1, 0].set_title('Features Selected by Each Method')
        axes[1, 0].grid(True, alpha=0.3)
        
        # Feature overlap
        all_methods = set(lasso_features) & set(rf_features)
        axes[1, 1].text(0.5, 0.5, 
                       f'Common features selected by\nLASSO and Random Forest:\n\n' +
                       '\n'.join(list(all_methods)[:10]),
                       ha='center', va='center', fontsize=10)
        axes[1, 1].set_title('Feature Consensus')
        axes[1, 1].axis('off')
        
        plt.suptitle('Embedded Methods Comparison', fontsize=14)
        plt.tight_layout()
        plt.show()
        
        return {
            'lasso': lasso_features,
            'random_forest': rf_features,
            'logistic_l1': lr_features
        }

# Demonstrate embedded methods
embedded_methods = EmbeddedMethods()

print("\n" + "=" * 50)
print("EMBEDDED METHODS")
print("=" * 50)

# Compare embedded methods
embedded_results = embedded_methods.compare_embedded_methods(X, y)

Complete Feature Selection Pipeline

class FeatureSelectionPipeline:
    """
    Complete pipeline for feature selection
    """
    
    def __init__(self):
        self.selected_features = None
        self.selection_report = {}
        
    def evaluate_feature_subset(self, X, y, selected_features, method_name):
        """
        Evaluate model performance with selected features
        """
        # Split data
        X_train, X_test, y_train, y_test = train_test_split(
            X[selected_features], y, test_size=0.3, random_state=42
        )
        
        # Train model
        model = RandomForestClassifier(n_estimators=100, random_state=42)
        model.fit(X_train, y_train)
        
        # Predict and evaluate
        y_pred = model.predict(X_test)
        accuracy = accuracy_score(y_test, y_pred)
        
        # Cross-validation score
        cv_scores = cross_val_score(model, X[selected_features], y, cv=5)
        
        return {
            'method': method_name,
            'n_features': len(selected_features),
            'accuracy': accuracy,
            'cv_mean': cv_scores.mean(),
            'cv_std': cv_scores.std()
        }
    
    def run_complete_pipeline(self, X, y):
        """
        Run complete feature selection pipeline
        """
        results = []
        
        # 1. Baseline - All features
        print("Evaluating baseline (all features)...")
        baseline = self.evaluate_feature_subset(X, y, X.columns, 'All Features')
        results.append(baseline)
        
        # 2. Filter method - Mutual Information
        print("Applying filter method...")
        filter_selector = SelectKBest(mutual_info_classif, k=10)
        filter_selector.fit(X, y)
        filter_features = X.columns[filter_selector.get_support()]
        filter_result = self.evaluate_feature_subset(X, y, filter_features, 'Filter (MI)')
        results.append(filter_result)
        
        # 3. Wrapper method - RFE
        print("Applying wrapper method...")
        estimator = LogisticRegression(random_state=42, max_iter=100)
        rfe = RFE(estimator, n_features_to_select=10)
        rfe.fit(X, y)
        wrapper_features = X.columns[rfe.support_]
        wrapper_result = self.evaluate_feature_subset(X, y, wrapper_features, 'Wrapper (RFE)')
        results.append(wrapper_result)
        
        # 4. Embedded method - Random Forest
        print("Applying embedded method...")
        rf = RandomForestClassifier(n_estimators=100, random_state=42)
        rf.fit(X, y)
        importance_threshold = np.percentile(rf.feature_importances_, 50)
        embedded_features = X.columns[rf.feature_importances_ > importance_threshold]
        embedded_result = self.evaluate_feature_subset(X, y, embedded_features, 'Embedded (RF)')
        results.append(embedded_result)
        
        # Create results DataFrame
        results_df = pd.DataFrame(results)
        
        # Visualization
        fig, axes = plt.subplots(1, 3, figsize=(15, 5))
        
        # Number of features vs accuracy
        axes[0].bar(results_df['method'], results_df['n_features'])
        axes[0].set_ylabel('Number of Features')
        axes[0].set_title('Feature Count by Method')
        axes[0].tick_params(axis='x', rotation=45)
        axes[0].grid(True, alpha=0.3)
        
        # Accuracy comparison
        axes[1].bar(results_df['method'], results_df['accuracy'], color='green', alpha=0.7)
        axes[1].set_ylabel('Test Accuracy')
        axes[1].set_title('Model Performance')
        axes[1].tick_params(axis='x', rotation=45)
        axes[1].grid(True, alpha=0.3)
        
        # CV scores with error bars
        axes[2].bar(results_df['method'], results_df['cv_mean'], 
                   yerr=results_df['cv_std'], capsize=5, color='blue', alpha=0.7)
        axes[2].set_ylabel('CV Accuracy')
        axes[2].set_title('Cross-Validation Performance')
        axes[2].tick_params(axis='x', rotation=45)
        axes[2].grid(True, alpha=0.3)
        
        plt.suptitle('Feature Selection Pipeline Results', fontsize=14)
        plt.tight_layout()
        plt.show()
        
        return results_df

# Run complete pipeline
print("\n" + "=" * 50)
print("COMPLETE FEATURE SELECTION PIPELINE")
print("=" * 50)

pipeline = FeatureSelectionPipeline()
results = pipeline.run_complete_pipeline(X, y)

print("\n" + "=" * 50)
print("PIPELINE RESULTS SUMMARY")
print("=" * 50)
print(results.to_string(index=False))

print("\n" + "=" * 50)
print("RECOMMENDATIONS")
print("=" * 50)
best_method = results.loc[results['accuracy'].idxmax()]
print(f"Best performing method: {best_method['method']}")
print(f"Number of features: {int(best_method['n_features'])}")
print(f"Test accuracy: {best_method['accuracy']:.3f}")
print(f"CV accuracy: {best_method['cv_mean']:.3f} (+/- {best_method['cv_std']:.3f})")

Best Practices

🎯 Feature Selection Guidelines

  • Start Simple: Try filter methods first for initial exploration
  • Consider Scale: Filter methods for many features, wrapper for fewer
  • Domain Knowledge: Combine statistical methods with expertise
  • Avoid Leakage: Do feature selection on training data only
  • Cross-Validation: Always validate selected features
  • Multiple Methods: Try different approaches and compare
  • Interpretability: Balance performance with model simplicity
  • Stability: Check if selected features are consistent across folds

Practice Exercises

Exercise 1: Text Classification Feature Selection

Select features for text classification:

  • Load a text dataset with TF-IDF features
  • Apply chi-square test for feature selection
  • Compare with mutual information
  • Evaluate impact on classification accuracy

Exercise 2: Time Series Feature Selection

Select relevant features for time series prediction:

  • Generate lag features and rolling statistics
  • Use forward selection with time-aware CV
  • Apply LASSO for automatic selection
  • Compare prediction performance

Summary

✅ You've Learned

  • Three main approaches: Filter, Wrapper, and Embedded methods
  • Filter methods: Fast, univariate, good for initial screening
  • Wrapper methods: Consider feature interactions, computationally expensive
  • Embedded methods: Built into model training, balanced approach
  • How to implement various feature selection techniques
  • Evaluation strategies for selected features
  • When to use each method based on your problem

📓 Learning Journal

Keep a learning journal — digital or physical. After this lesson, take a few minutes to write down:

  • Key concepts you learned
  • Techniques that clicked for you
  • Questions or confusion points to revisit
  • Ideas you want to try
  • Your progress and feelings about learning this

✍️ This lesson's prompt: Which selection approach felt most trustworthy to you — filter, wrapper, or embedded — and why might that intuition change if your dataset grew from thousands of rows to millions?

📝 Lesson Summary

🎓 Key Takeaways

  • Feature selection keeps the original features, so models stay interpretable — unlike PCA, which creates new combined features.
  • Filter methods are fast and model-agnostic; wrapper methods are accurate but expensive; embedded methods balance both.
  • Fit selectors on the training data only — selecting features using the test set leaks information and inflates scores.
  • Fewer, well-chosen features reduce overfitting, speed up training, and make results easier to explain.

🎉 What You've Accomplished

You can now take a wide, noisy dataset and trim it down to the handful of features that genuinely drive predictions — improving both performance and interpretability.

❓ Common Questions at This Stage

When should I use feature selection instead of PCA?

Use selection when interpretability matters and you want to keep the original, named features. Use PCA when you only care about predictive signal and can accept transformed, harder-to-read components.

Won't dropping features throw away useful information?

Often it removes noise and redundancy rather than signal — irrelevant or highly correlated features can actively hurt a model. Always confirm with cross-validation that the trimmed set performs as well or better.

How many features should I keep?

There is no fixed number. Plot the cross-validated score against the number of selected features and keep adding until performance stops improving (or starts to drop).

🔭 Looking Ahead

Next you'll move from selecting among existing features to extracting new ones — techniques that transform the data onto new axes rather than choosing a subset of the originals.

✅ Before the Next Lesson

🌟 Encouragement for the Journey

Choosing the right features is often what separates a mediocre model from a great one. You now hold one of the most practical, high-impact skills in all of applied machine learning. Keep going!