Machine Learning

How I Built an ML-Powered Email Spam Detector

A deep dive into NLP preprocessing, feature engineering, and probabilistic classification — achieving high-accuracy spam detection from scratch.

June 28, 20266 min read
Suyash Vakhariya
Suyash VakhariyaAI Engineer & Technical Product Manager

The Problem

Email spam is one of the oldest problems in computer science, yet it remains incredibly relevant. Every day, over 150 billion spam emails are sent globally. Building an effective spam classifier requires understanding both NLP fundamentals and practical ML engineering.

The Pipeline

Raw Emails → Text Cleaning → TF-IDF Vectorization → Model Training → Evaluation → Deployment

Step 1: Data Preprocessing

The raw email dataset contains noise — HTML tags, special characters, inconsistent formatting. Cleaning is critical:

python
import pandas as pd
import re

def clean_text(text):
    text = re.sub(r'<[^>]+>', '', text)      # Remove HTML
    text = re.sub(r'[^a-zA-Z\s]', '', text)  # Keep only letters
    text = text.lower().strip()
    return text

df = pd.read_csv('emails.csv')
df['clean_text'] = df['text'].apply(clean_text)

Step 2: TF-IDF Vectorization

TF-IDF (Term Frequency - Inverse Document Frequency) converts text into numerical features that capture word importance:

python
from sklearn.feature_extraction.text import TfidfVectorizer

tfidf = TfidfVectorizer(
    max_features=5000,
    stop_words='english',
    max_df=0.95,    # Ignore words appearing in >95% of docs
    min_df=5         # Ignore words appearing in <5 docs
)
X_tfidf = tfidf.fit_transform(df['clean_text'])

Why TF-IDF over simple word counts?

  • Common words like "the" and "is" get downweighted automatically
  • Rare, distinctive words (like "lottery" or "winner") get higher scores
  • It creates sparse, efficient feature matrices

Step 3: Model Training

I compared two classifiers:

ModelAccuracyPrecisionRecallF1
Naive Bayes96.2%0.950.940.94
SVM97.1%0.970.960.96
python
from sklearn.naive_bayes import MultinomialNB
from sklearn.svm import LinearSVC
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X_tfidf, df['label'], test_size=0.2, stratify=df['label']
)

# Multinomial Naive Bayes
nb = MultinomialNB(alpha=0.1)
nb.fit(X_train, y_train)

# Linear SVM
svm = LinearSVC(max_iter=1000)
svm.fit(X_train, y_train)

Step 4: What Makes Spam "Spammy"?

Looking at the highest-weighted features reveals what the model learned:

Top spam indicators: "free", "winner", "click", "offer", "limited", "act now", "credit", "congratulations"

Top ham indicators: "meeting", "project", "attached", "schedule", "regards", "update"

Lessons Learned

  • 1.Data quality > model complexity — cleaning the dataset improved accuracy by 4% without changing the model
  • 2.TF-IDF is surprisingly powerful — for text classification, you don't always need word embeddings or transformers
  • 3.Naive Bayes is fast and effective — for production spam filtering, speed matters as much as accuracy
  • 4.Class imbalance matters — use stratified splits and look at precision/recall, not just accuracy

Full code available on GitHub.

NLPScikit-learnClassificationPython
Suyash

Suyash Vakhariya

AI Engineer & Technical Product Manager. Building production AI systems.

© 2026 Suyash Vakhariya