AI + Security

Network Intrusion Detection with Machine Learning: A Complete Guide

How to build an ML-based network intrusion detection system (NIDS) using Python — from feature engineering on network traffic data to real-time anomaly classification.

July 25, 202610 min read
Suyash Vakhariya
Suyash VakhariyaAI Engineer & Technical Product Manager

Why ML for Network Security?

Traditional intrusion detection systems (IDS) rely on signature-based detection — they match network traffic against known attack patterns. This approach has a critical flaw: it can't detect zero-day attacks (new, previously unseen threats).

Machine learning offers a solution: anomaly-based detection. By learning what "normal" network traffic looks like, ML models can flag deviations — even attacks that have never been seen before.

The Dataset: NSL-KDD

I used the NSL-KDD dataset, an improved version of the original KDD Cup 1999 dataset. It contains:

  • 125,973 training records and 22,544 test records
  • 41 features per connection record
  • 5 categories: Normal, DoS, Probe, R2L, U2R

Key features include:

FeatureDescription
durationConnection length in seconds
protocol_typeTCP, UDP, or ICMP
serviceNetwork service (HTTP, FTP, etc.)
flagConnection status
src_bytesData bytes from source
dst_bytesData bytes from destination
logged_inLogin status (1/0)

The ML Pipeline

Raw Network Data → Feature Engineering → Normalization → Model Training → Real-Time Classification

Step 1: Feature Engineering

Network data requires careful preprocessing:

python
import pandas as pd
from sklearn.preprocessing import LabelEncoder, StandardScaler

# Encode categorical features
le = LabelEncoder()
for col in ['protocol_type', 'service', 'flag']:
    df[col] = le.fit_transform(df[col])

# Normalize numerical features
scaler = StandardScaler()
numerical_cols = df.select_dtypes(include=['float64', 'int64']).columns
df[numerical_cols] = scaler.fit_transform(df[numerical_cols])

Step 2: Feature Selection

With 41 features, dimensionality reduction is important:

python
from sklearn.feature_selection import SelectKBest, chi2

# Select top 20 most informative features
selector = SelectKBest(chi2, k=20)
X_selected = selector.fit_transform(X, y)

# Most important features:
# src_bytes, dst_bytes, logged_in, count, srv_count,
# same_srv_rate, dst_host_srv_count, dst_host_same_srv_rate

Step 3: Model Comparison

I evaluated multiple classifiers:

ModelAccuracyPrecisionRecallF1Training Time
Random Forest99.2%0.990.990.9912s
Decision Tree98.7%0.980.980.982s
SVM97.4%0.970.960.9645s
KNN97.1%0.970.960.961s*
Logistic Reg92.3%0.910.900.903s

*KNN inference is slow at scale despite fast "training"

Winner: Random Forest — Best accuracy, fast training, and interpretable feature importances.

Step 4: Attack Type Classification

Beyond binary (normal vs. attack), the model classifies attack types:

python
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report

rf = RandomForestClassifier(
    n_estimators=100,
    max_depth=20,
    min_samples_split=5,
    n_jobs=-1,
    random_state=42
)
rf.fit(X_train, y_train)

print(classification_report(y_test, rf.predict(X_test)))

Per-class results:

  • Normal: 99.5% precision
  • DoS: 99.8% precision (easiest to detect — high volume)
  • Probe: 97.2% precision (port scanning, etc.)
  • R2L: 85.3% precision (remote to local — harder, fewer examples)
  • U2R: 78.1% precision (user to root — rarest, most dangerous)

Key Challenges

1. Class Imbalance

U2R attacks are extremely rare (<1% of data). I addressed this with:

  • SMOTE (Synthetic Minority Over-sampling)
  • Class weights in Random Forest
  • Evaluation by precision/recall rather than just accuracy

2. Feature Drift

Network traffic patterns change over time. A model trained on 2024 traffic may not work well in 2026. Continuous retraining is essential.

3. False Positives

In production, false positives are expensive — they cause alert fatigue. I tuned the decision threshold to minimize false positive rate while maintaining >95% detection rate.

Real-World Considerations

For production deployment, you'd need:

  • 1.Packet capture — Tools like tcpdump or Wireshark to collect raw traffic
  • 2.Feature extraction — Convert raw packets to the 41-feature format
  • 3.Streaming inference — Process packets in real-time (not batch)
  • 4.Alert integration — Feed detections into SIEM systems (Splunk, ELK)
  • 5.Model monitoring — Track prediction confidence and retrain on drift

Impact

Network intrusion detection is a $6.2 billion industry (2025). ML-based approaches are increasingly replacing signature-based systems because:

  • They detect zero-day attacks
  • They adapt to new traffic patterns
  • They reduce false positive rates with proper tuning

Source code: GitHub

CybersecurityMLPythonNetwork Security
Suyash

Suyash Vakhariya

AI Engineer & Technical Product Manager. Building production AI systems.

© 2026 Suyash Vakhariya