The Problem
Email spam is one of the oldest problems in computer science, yet it remains incredibly relevant. Every day, over 150 billion spam emails are sent globally. Building an effective spam classifier requires understanding both NLP fundamentals and practical ML engineering.
The Pipeline
Raw Emails → Text Cleaning → TF-IDF Vectorization → Model Training → Evaluation → DeploymentStep 1: Data Preprocessing
The raw email dataset contains noise — HTML tags, special characters, inconsistent formatting. Cleaning is critical:
import pandas as pd
import re
def clean_text(text):
text = re.sub(r'<[^>]+>', '', text) # Remove HTML
text = re.sub(r'[^a-zA-Z\s]', '', text) # Keep only letters
text = text.lower().strip()
return text
df = pd.read_csv('emails.csv')
df['clean_text'] = df['text'].apply(clean_text)Step 2: TF-IDF Vectorization
TF-IDF (Term Frequency - Inverse Document Frequency) converts text into numerical features that capture word importance:
from sklearn.feature_extraction.text import TfidfVectorizer
tfidf = TfidfVectorizer(
max_features=5000,
stop_words='english',
max_df=0.95, # Ignore words appearing in >95% of docs
min_df=5 # Ignore words appearing in <5 docs
)
X_tfidf = tfidf.fit_transform(df['clean_text'])Why TF-IDF over simple word counts?
- Common words like "the" and "is" get downweighted automatically
- Rare, distinctive words (like "lottery" or "winner") get higher scores
- It creates sparse, efficient feature matrices
Step 3: Model Training
I compared two classifiers:
| Model | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|
| Naive Bayes | 96.2% | 0.95 | 0.94 | 0.94 |
| SVM | 97.1% | 0.97 | 0.96 | 0.96 |
from sklearn.naive_bayes import MultinomialNB
from sklearn.svm import LinearSVC
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X_tfidf, df['label'], test_size=0.2, stratify=df['label']
)
# Multinomial Naive Bayes
nb = MultinomialNB(alpha=0.1)
nb.fit(X_train, y_train)
# Linear SVM
svm = LinearSVC(max_iter=1000)
svm.fit(X_train, y_train)Step 4: What Makes Spam "Spammy"?
Looking at the highest-weighted features reveals what the model learned:
Top spam indicators: "free", "winner", "click", "offer", "limited", "act now", "credit", "congratulations"
Top ham indicators: "meeting", "project", "attached", "schedule", "regards", "update"
Lessons Learned
- 1.Data quality > model complexity — cleaning the dataset improved accuracy by 4% without changing the model
- 2.TF-IDF is surprisingly powerful — for text classification, you don't always need word embeddings or transformers
- 3.Naive Bayes is fast and effective — for production spam filtering, speed matters as much as accuracy
- 4.Class imbalance matters — use stratified splits and look at precision/recall, not just accuracy
Full code available on GitHub.
