📝 Text Preprocessing: Preparing Text for NLP
📚 What You'll Learn
By the end of this lesson, you will be able to:
- Clean raw text by lowercasing and removing URLs, emails, punctuation, and special characters
- Tokenize text at both the word and sentence level
- Remove stop words and explain the trade-offs of doing so
- Apply stemming versus lemmatization and choose the right one for a task
- Build a reusable end-to-end preprocessing pipeline for a downstream NLP task
⏱️ Estimated Time: 45–60 minutes
🎯 Project: Build a text-cleaning pipeline that turns raw documents into normalized tokens ready for vectorization.
Introduction
Text preprocessing is the foundation of Natural Language Processing, transforming raw text into a format suitable for machine learning algorithms. From cleaning and tokenization to stemming and lemmatization, these techniques are essential for extracting meaningful insights from textual data.
Text Preprocessing Pipeline Overview
Getting Started with NLTK
import nltk
import re
import string
from nltk.tokenize import word_tokenize, sent_tokenize
from nltk.corpus import stopwords
from nltk.stem import PorterStemmer, WordNetLemmatizer
from nltk.tag import pos_tag
# Download required NLTK data
nltk.download('punkt')
nltk.download('stopwords')
nltk.download('wordnet')
nltk.download('averaged_perceptron_tagger')
# Sample text for demonstration
sample_text = """
Hello World! This is a sample text for NLP preprocessing.
We'll clean, tokenize, and normalize this text.
Visit https://example.com for more info or email us at info@example.com
"""
print("Original Text:")
print(sample_text)
Text Cleaning
def clean_text(text):
"""Basic text cleaning operations"""
# Remove URLs
text = re.sub(r'https?://\S+|www\.\S+', '', text)
# Remove email addresses
text = re.sub(r'\S+@\S+', '', text)
# Remove HTML tags
text = re.sub(r'<.*?>', '', text)
# Remove extra whitespace
text = ' '.join(text.split())
return text
# Apply cleaning
cleaned_text = clean_text(sample_text)
print("Cleaned Text:")
print(cleaned_text)
Tokenization
# Word tokenization
word_tokens = word_tokenize(cleaned_text)
print("Word Tokens:")
print(word_tokens)
# Sentence tokenization
sentences = sent_tokenize(cleaned_text)
print("\nSentences:")
for i, sent in enumerate(sentences, 1):
print(f"{i}. {sent}")
# Custom tokenization with regex
import re
custom_tokens = re.findall(r'\w+', cleaned_text.lower())
print("\nCustom Tokens (alphanumeric only):")
print(custom_tokens)
Stop Words Removal
# Get English stop words
stop_words = set(stopwords.words('english'))
print(f"Number of stop words: {len(stop_words)}")
print(f"Sample stop words: {list(stop_words)[:10]}")
# Remove stop words
filtered_tokens = [token for token in word_tokens
if token.lower() not in stop_words]
print("\nBefore stop word removal:", len(word_tokens), "tokens")
print("After stop word removal:", len(filtered_tokens), "tokens")
print("Filtered tokens:", filtered_tokens)
Stemming vs Lemmatization
# Initialize stemmer and lemmatizer
stemmer = PorterStemmer()
lemmatizer = WordNetLemmatizer()
# Example words
test_words = ['running', 'runs', 'ran', 'better', 'good',
'universities', 'universal', 'universe']
print("Word -> Stem -> Lemma")
print("-" * 40)
for word in test_words:
stem = stemmer.stem(word)
lemma = lemmatizer.lemmatize(word)
print(f"{word:15} -> {stem:15} -> {lemma:15}")
# Apply to our text
stemmed_tokens = [stemmer.stem(token) for token in filtered_tokens]
lemmatized_tokens = [lemmatizer.lemmatize(token) for token in filtered_tokens]
print("\nOriginal:", filtered_tokens[:5])
print("Stemmed:", stemmed_tokens[:5])
print("Lemmatized:", lemmatized_tokens[:5])
Complete Preprocessing Pipeline
class TextPreprocessor:
"""Complete text preprocessing pipeline"""
def __init__(self):
self.stemmer = PorterStemmer()
self.lemmatizer = WordNetLemmatizer()
self.stop_words = set(stopwords.words('english'))
def preprocess(self, text, method='lemmatize'):
"""
Complete preprocessing pipeline
method: 'stem' or 'lemmatize'
"""
# 1. Clean text
text = self.clean(text)
# 2. Tokenize
tokens = word_tokenize(text.lower())
# 3. Remove punctuation
tokens = [token for token in tokens if token.isalnum()]
# 4. Remove stop words
tokens = [token for token in tokens
if token not in self.stop_words]
# 5. Apply stemming or lemmatization
if method == 'stem':
tokens = [self.stemmer.stem(token) for token in tokens]
else:
# Get POS tags for better lemmatization
pos_tags = pos_tag(tokens)
tokens = [self.lemmatizer.lemmatize(token,
self.get_wordnet_pos(pos))
for token, pos in pos_tags]
return tokens
def clean(self, text):
"""Clean text"""
# Remove URLs
text = re.sub(r'https?://\S+|www\.\S+', '', text)
# Remove emails
text = re.sub(r'\S+@\S+', '', text)
# Remove HTML
text = re.sub(r'<.*?>', '', text)
# Remove special characters
text = re.sub(r'[^a-zA-Z0-9\s]', ' ', text)
# Remove extra spaces
text = ' '.join(text.split())
return text
def get_wordnet_pos(self, treebank_tag):
"""Convert treebank POS tag to wordnet POS tag"""
if treebank_tag.startswith('J'):
return 'a' # adjective
elif treebank_tag.startswith('V'):
return 'v' # verb
elif treebank_tag.startswith('N'):
return 'n' # noun
elif treebank_tag.startswith('R'):
return 'r' # adverb
else:
return 'n' # default to noun
# Use the pipeline
preprocessor = TextPreprocessor()
# Test text
test_text = """
The quick brown foxes were jumping over the lazy dogs.
They ran quickly through the beautiful forest.
"""
# Process with both methods
stemmed = preprocessor.preprocess(test_text, method='stem')
lemmatized = preprocessor.preprocess(test_text, method='lemmatize')
print("Original:", test_text)
print("\nStemmed:", ' '.join(stemmed))
print("\nLemmatized:", ' '.join(lemmatized))
Task-Specific Preprocessing
Advanced Techniques
# Handling contractions
contractions = {
"won't": "will not",
"can't": "cannot",
"n't": " not",
"'re": " are",
"'ve": " have",
"'ll": " will",
"'d": " would",
"'m": " am"
}
def expand_contractions(text):
"""Expand contractions in text"""
for contraction, expansion in contractions.items():
text = text.replace(contraction, expansion)
return text
# N-gram generation
from nltk.util import ngrams
def generate_ngrams(tokens, n):
"""Generate n-grams from tokens"""
n_grams = list(ngrams(tokens, n))
return [' '.join(gram) for gram in n_grams]
tokens = word_tokenize("Machine learning is amazing")
bigrams = generate_ngrams(tokens, 2)
trigrams = generate_ngrams(tokens, 3)
print("Bigrams:", bigrams)
print("Trigrams:", trigrams)
Best Practices
🎯 Key Guidelines
- Understand your data: Check language, encoding, and format
- Task-specific approach: Different tasks need different preprocessing
- Preserve information: Don't over-clean, keep what's relevant
- Consistency: Same preprocessing for train and test data
- Document steps: Keep track of preprocessing decisions
- Validate results: Check output at each step
Practice Exercises
Exercise 1: Build a Custom Preprocessor
Create a preprocessor for social media text that handles:
- Hashtags and mentions
- Emoticons and emojis
- Slang and abbreviations
Exercise 2: Compare Preprocessing Methods
Test different preprocessing combinations and measure their impact on:
- Text classification accuracy
- Processing speed
- Vocabulary size
Summary
✅ You've Learned
- How to clean text data (remove URLs, emails, special characters)
- Different tokenization methods (word, sentence, custom)
- Stop words removal and its impact
- Difference between stemming and lemmatization
- Building complete preprocessing pipelines
- Task-specific preprocessing strategies
📓 Learning Journal
Keep a learning journal — digital or physical. After this lesson, take a few minutes to write down:
- Key concepts you learned
- Techniques that clicked for you
- Questions or confusion points to revisit
- Ideas you want to try
- Your progress and feelings about learning this
✍️ This lesson's prompt: Stemming is fast but crude; lemmatization is accurate but slower. For a text project you actually care about, which trade-off would you choose, and what about the task drives that decision?
📝 Lesson Summary
🎓 Key Takeaways
- Raw text must be normalized — cleaned and tokenized — before any model can use it.
- Stop-word removal cuts noise, but it can discard meaning for tasks where words like "not" matter.
- Stemming chops word endings with heuristics; lemmatization maps words to real dictionary base forms using context.
- Preprocessing choices should be driven by the downstream task, never applied blindly.
🎉 What You've Accomplished
You can now turn messy, real-world text into clean, consistent tokens that are ready for feature extraction — the essential first step of every NLP project.
❓ Common Questions at This Stage
Should I always remove stop words?
No. For tasks like topic modeling, removing them helps. But for sentiment analysis or anything where negation and function words carry meaning, removing them can hurt performance.
Stemming or lemmatization — which should I use?
Use lemmatization when you want accurate, readable base forms and can afford the extra compute. Use stemming when speed matters and rough matching (e.g., for search indexing) is good enough.
Should I lowercase everything?
Usually yes, but case can carry signal — "US" versus "us", or named entities. Consider whether the task benefits from preserving case before you flatten it.
🔭 Looking Ahead
Next you'll turn these cleaned tokens into numbers — bag-of-words, TF-IDF, and word embeddings — the representations that let models actually learn from text.
✅ Before the Next Lesson
- Clean a paragraph of your own text with your pipeline and inspect the output tokens.
- Compare stemming versus lemmatization on the same handful of words and note the differences.
- Write your Learning Journal entry for this lesson.
🌟 Encouragement for the Journey
Clean text is the quiet foundation of every great NLP model. Master this step and everything downstream — from sentiment to search to chatbots — gets easier. You're building the groundwork well.