Skip to main content

📝 Text Preprocessing: Preparing Text for NLP

📚 What You'll Learn

By the end of this lesson, you will be able to:

  • Clean raw text by lowercasing and removing URLs, emails, punctuation, and special characters
  • Tokenize text at both the word and sentence level
  • Remove stop words and explain the trade-offs of doing so
  • Apply stemming versus lemmatization and choose the right one for a task
  • Build a reusable end-to-end preprocessing pipeline for a downstream NLP task

⏱️ Estimated Time: 45–60 minutes

🎯 Project: Build a text-cleaning pipeline that turns raw documents into normalized tokens ready for vectorization.

Introduction

Text preprocessing is the foundation of Natural Language Processing, transforming raw text into a format suitable for machine learning algorithms. From cleaning and tokenization to stemming and lemmatization, these techniques are essential for extracting meaningful insights from textual data.

Text Preprocessing Pipeline Overview

flowchart TD A[Raw Text] --> B[Text Cleaning] B --> B1[Remove HTML/XML] B --> B2[Handle Special Characters] B --> B3[Remove URLs/Emails] B --> C[Tokenization] C --> C1[Word Tokenization] C --> C2[Sentence Tokenization] C --> D[Normalization] D --> D1[Lowercase Conversion] D --> D2[Expand Contractions] D --> D3[Remove Accents] D --> E[Stop Words Removal] E --> F{Choose Method} F --> G[Stemming] G --> G1[Porter Stemmer] G --> G2[Lancaster Stemmer] F --> H[Lemmatization] H --> H1[With POS Tags] H --> H2[Dictionary Based] G --> I[Processed Text] H --> I I --> J[Ready for ML] style A fill:#f9f,stroke:#333,stroke-width:4px style J fill:#9f9,stroke:#333,stroke-width:4px

Getting Started with NLTK

import nltk
import re
import string
from nltk.tokenize import word_tokenize, sent_tokenize
from nltk.corpus import stopwords
from nltk.stem import PorterStemmer, WordNetLemmatizer
from nltk.tag import pos_tag

# Download required NLTK data
nltk.download('punkt')
nltk.download('stopwords')
nltk.download('wordnet')
nltk.download('averaged_perceptron_tagger')

# Sample text for demonstration
sample_text = """
Hello World! This is a sample text for NLP preprocessing.
We'll clean, tokenize, and normalize this text.
Visit https://example.com for more info or email us at info@example.com
"""

print("Original Text:")
print(sample_text)

Text Cleaning

def clean_text(text):
    """Basic text cleaning operations"""
    
    # Remove URLs
    text = re.sub(r'https?://\S+|www\.\S+', '', text)
    
    # Remove email addresses
    text = re.sub(r'\S+@\S+', '', text)
    
    # Remove HTML tags
    text = re.sub(r'<.*?>', '', text)
    
    # Remove extra whitespace
    text = ' '.join(text.split())
    
    return text

# Apply cleaning
cleaned_text = clean_text(sample_text)
print("Cleaned Text:")
print(cleaned_text)

Tokenization

graph TD A[Text: Hello World] --> B{Tokenization Type} B -->|Word| C[Hello, World] B -->|Sentence| D[Hello World.] B -->|Character| E[H,e,l,l,o, ,W,o,r,l,d] B -->|N-gram| F[Hello World, World]
# Word tokenization
word_tokens = word_tokenize(cleaned_text)
print("Word Tokens:")
print(word_tokens)

# Sentence tokenization
sentences = sent_tokenize(cleaned_text)
print("\nSentences:")
for i, sent in enumerate(sentences, 1):
    print(f"{i}. {sent}")

# Custom tokenization with regex
import re
custom_tokens = re.findall(r'\w+', cleaned_text.lower())
print("\nCustom Tokens (alphanumeric only):")
print(custom_tokens)

Stop Words Removal

# Get English stop words
stop_words = set(stopwords.words('english'))
print(f"Number of stop words: {len(stop_words)}")
print(f"Sample stop words: {list(stop_words)[:10]}")

# Remove stop words
filtered_tokens = [token for token in word_tokens 
                   if token.lower() not in stop_words]

print("\nBefore stop word removal:", len(word_tokens), "tokens")
print("After stop word removal:", len(filtered_tokens), "tokens")
print("Filtered tokens:", filtered_tokens)

Stemming vs Lemmatization

graph LR A[running] --> B{Method} B -->|Stemming| C[run] B -->|Lemmatization| D[run] E[better] --> F{Method} F -->|Stemming| G[better] F -->|Lemmatization| H[good] I[universities] --> J{Method} J -->|Stemming| K[univers] J -->|Lemmatization| L[university] style C fill:#ffcccc style D fill:#ccffcc style G fill:#ffcccc style H fill:#ccffcc style K fill:#ffcccc style L fill:#ccffcc
# Initialize stemmer and lemmatizer
stemmer = PorterStemmer()
lemmatizer = WordNetLemmatizer()

# Example words
test_words = ['running', 'runs', 'ran', 'better', 'good', 
              'universities', 'universal', 'universe']

print("Word -> Stem -> Lemma")
print("-" * 40)
for word in test_words:
    stem = stemmer.stem(word)
    lemma = lemmatizer.lemmatize(word)
    print(f"{word:15} -> {stem:15} -> {lemma:15}")

# Apply to our text
stemmed_tokens = [stemmer.stem(token) for token in filtered_tokens]
lemmatized_tokens = [lemmatizer.lemmatize(token) for token in filtered_tokens]

print("\nOriginal:", filtered_tokens[:5])
print("Stemmed:", stemmed_tokens[:5])
print("Lemmatized:", lemmatized_tokens[:5])

Complete Preprocessing Pipeline

class TextPreprocessor:
    """Complete text preprocessing pipeline"""
    
    def __init__(self):
        self.stemmer = PorterStemmer()
        self.lemmatizer = WordNetLemmatizer()
        self.stop_words = set(stopwords.words('english'))
    
    def preprocess(self, text, method='lemmatize'):
        """
        Complete preprocessing pipeline
        method: 'stem' or 'lemmatize'
        """
        # 1. Clean text
        text = self.clean(text)
        
        # 2. Tokenize
        tokens = word_tokenize(text.lower())
        
        # 3. Remove punctuation
        tokens = [token for token in tokens if token.isalnum()]
        
        # 4. Remove stop words
        tokens = [token for token in tokens 
                 if token not in self.stop_words]
        
        # 5. Apply stemming or lemmatization
        if method == 'stem':
            tokens = [self.stemmer.stem(token) for token in tokens]
        else:
            # Get POS tags for better lemmatization
            pos_tags = pos_tag(tokens)
            tokens = [self.lemmatizer.lemmatize(token, 
                     self.get_wordnet_pos(pos)) 
                     for token, pos in pos_tags]
        
        return tokens
    
    def clean(self, text):
        """Clean text"""
        # Remove URLs
        text = re.sub(r'https?://\S+|www\.\S+', '', text)
        # Remove emails
        text = re.sub(r'\S+@\S+', '', text)
        # Remove HTML
        text = re.sub(r'<.*?>', '', text)
        # Remove special characters
        text = re.sub(r'[^a-zA-Z0-9\s]', ' ', text)
        # Remove extra spaces
        text = ' '.join(text.split())
        return text
    
    def get_wordnet_pos(self, treebank_tag):
        """Convert treebank POS tag to wordnet POS tag"""
        if treebank_tag.startswith('J'):
            return 'a'  # adjective
        elif treebank_tag.startswith('V'):
            return 'v'  # verb
        elif treebank_tag.startswith('N'):
            return 'n'  # noun
        elif treebank_tag.startswith('R'):
            return 'r'  # adverb
        else:
            return 'n'  # default to noun

# Use the pipeline
preprocessor = TextPreprocessor()

# Test text
test_text = """
The quick brown foxes were jumping over the lazy dogs.
They ran quickly through the beautiful forest.
"""

# Process with both methods
stemmed = preprocessor.preprocess(test_text, method='stem')
lemmatized = preprocessor.preprocess(test_text, method='lemmatize')

print("Original:", test_text)
print("\nStemmed:", ' '.join(stemmed))
print("\nLemmatized:", ' '.join(lemmatized))

Task-Specific Preprocessing

graph LR A[Text Data] --> B{Task Type} B -->|Classification| C[Aggressive Cleaning] C --> C1[Remove All Punctuation] C --> C2[Stemming OK] B -->|NER| D[Preserve Structure] D --> D1[Keep Capitalization] D --> D2[Keep Punctuation] B -->|Sentiment| E[Balanced Approach] E --> E1[Keep Emoticons 😊] E --> E2[Preserve Intensifiers] B -->|Topic Modeling| F[Focus on Content] F --> F1[Remove Stop Words] F --> F2[Lemmatization Preferred]

Advanced Techniques

# Handling contractions
contractions = {
    "won't": "will not",
    "can't": "cannot",
    "n't": " not",
    "'re": " are",
    "'ve": " have",
    "'ll": " will",
    "'d": " would",
    "'m": " am"
}

def expand_contractions(text):
    """Expand contractions in text"""
    for contraction, expansion in contractions.items():
        text = text.replace(contraction, expansion)
    return text

# N-gram generation
from nltk.util import ngrams

def generate_ngrams(tokens, n):
    """Generate n-grams from tokens"""
    n_grams = list(ngrams(tokens, n))
    return [' '.join(gram) for gram in n_grams]

tokens = word_tokenize("Machine learning is amazing")
bigrams = generate_ngrams(tokens, 2)
trigrams = generate_ngrams(tokens, 3)

print("Bigrams:", bigrams)
print("Trigrams:", trigrams)

Best Practices

🎯 Key Guidelines

  • Understand your data: Check language, encoding, and format
  • Task-specific approach: Different tasks need different preprocessing
  • Preserve information: Don't over-clean, keep what's relevant
  • Consistency: Same preprocessing for train and test data
  • Document steps: Keep track of preprocessing decisions
  • Validate results: Check output at each step

Practice Exercises

Exercise 1: Build a Custom Preprocessor

Create a preprocessor for social media text that handles:

  • Hashtags and mentions
  • Emoticons and emojis
  • Slang and abbreviations

Exercise 2: Compare Preprocessing Methods

Test different preprocessing combinations and measure their impact on:

  • Text classification accuracy
  • Processing speed
  • Vocabulary size

Summary

✅ You've Learned

  • How to clean text data (remove URLs, emails, special characters)
  • Different tokenization methods (word, sentence, custom)
  • Stop words removal and its impact
  • Difference between stemming and lemmatization
  • Building complete preprocessing pipelines
  • Task-specific preprocessing strategies

📓 Learning Journal

Keep a learning journal — digital or physical. After this lesson, take a few minutes to write down:

  • Key concepts you learned
  • Techniques that clicked for you
  • Questions or confusion points to revisit
  • Ideas you want to try
  • Your progress and feelings about learning this

✍️ This lesson's prompt: Stemming is fast but crude; lemmatization is accurate but slower. For a text project you actually care about, which trade-off would you choose, and what about the task drives that decision?

📝 Lesson Summary

🎓 Key Takeaways

  • Raw text must be normalized — cleaned and tokenized — before any model can use it.
  • Stop-word removal cuts noise, but it can discard meaning for tasks where words like "not" matter.
  • Stemming chops word endings with heuristics; lemmatization maps words to real dictionary base forms using context.
  • Preprocessing choices should be driven by the downstream task, never applied blindly.

🎉 What You've Accomplished

You can now turn messy, real-world text into clean, consistent tokens that are ready for feature extraction — the essential first step of every NLP project.

❓ Common Questions at This Stage

Should I always remove stop words?

No. For tasks like topic modeling, removing them helps. But for sentiment analysis or anything where negation and function words carry meaning, removing them can hurt performance.

Stemming or lemmatization — which should I use?

Use lemmatization when you want accurate, readable base forms and can afford the extra compute. Use stemming when speed matters and rough matching (e.g., for search indexing) is good enough.

Should I lowercase everything?

Usually yes, but case can carry signal — "US" versus "us", or named entities. Consider whether the task benefits from preserving case before you flatten it.

🔭 Looking Ahead

Next you'll turn these cleaned tokens into numbers — bag-of-words, TF-IDF, and word embeddings — the representations that let models actually learn from text.

✅ Before the Next Lesson

🌟 Encouragement for the Journey

Clean text is the quiet foundation of every great NLP model. Master this step and everything downstream — from sentiment to search to chatbots — gets easier. You're building the groundwork well.