Why ML for Network Security?
Traditional intrusion detection systems (IDS) rely on signature-based detection — they match network traffic against known attack patterns. This approach has a critical flaw: it can't detect zero-day attacks (new, previously unseen threats).
Machine learning offers a solution: anomaly-based detection. By learning what "normal" network traffic looks like, ML models can flag deviations — even attacks that have never been seen before.
The Dataset: NSL-KDD
I used the NSL-KDD dataset, an improved version of the original KDD Cup 1999 dataset. It contains:
- 125,973 training records and 22,544 test records
- 41 features per connection record
- 5 categories: Normal, DoS, Probe, R2L, U2R
Key features include:
| Feature | Description |
|---|---|
| duration | Connection length in seconds |
| protocol_type | TCP, UDP, or ICMP |
| service | Network service (HTTP, FTP, etc.) |
| flag | Connection status |
| src_bytes | Data bytes from source |
| dst_bytes | Data bytes from destination |
| logged_in | Login status (1/0) |
The ML Pipeline
Raw Network Data → Feature Engineering → Normalization → Model Training → Real-Time ClassificationStep 1: Feature Engineering
Network data requires careful preprocessing:
import pandas as pd
from sklearn.preprocessing import LabelEncoder, StandardScaler
# Encode categorical features
le = LabelEncoder()
for col in ['protocol_type', 'service', 'flag']:
df[col] = le.fit_transform(df[col])
# Normalize numerical features
scaler = StandardScaler()
numerical_cols = df.select_dtypes(include=['float64', 'int64']).columns
df[numerical_cols] = scaler.fit_transform(df[numerical_cols])Step 2: Feature Selection
With 41 features, dimensionality reduction is important:
from sklearn.feature_selection import SelectKBest, chi2
# Select top 20 most informative features
selector = SelectKBest(chi2, k=20)
X_selected = selector.fit_transform(X, y)
# Most important features:
# src_bytes, dst_bytes, logged_in, count, srv_count,
# same_srv_rate, dst_host_srv_count, dst_host_same_srv_rateStep 3: Model Comparison
I evaluated multiple classifiers:
| Model | Accuracy | Precision | Recall | F1 | Training Time |
|---|---|---|---|---|---|
| Random Forest | 99.2% | 0.99 | 0.99 | 0.99 | 12s |
| Decision Tree | 98.7% | 0.98 | 0.98 | 0.98 | 2s |
| SVM | 97.4% | 0.97 | 0.96 | 0.96 | 45s |
| KNN | 97.1% | 0.97 | 0.96 | 0.96 | 1s* |
| Logistic Reg | 92.3% | 0.91 | 0.90 | 0.90 | 3s |
*KNN inference is slow at scale despite fast "training"
Winner: Random Forest — Best accuracy, fast training, and interpretable feature importances.
Step 4: Attack Type Classification
Beyond binary (normal vs. attack), the model classifies attack types:
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report
rf = RandomForestClassifier(
n_estimators=100,
max_depth=20,
min_samples_split=5,
n_jobs=-1,
random_state=42
)
rf.fit(X_train, y_train)
print(classification_report(y_test, rf.predict(X_test)))Per-class results:
- Normal: 99.5% precision
- DoS: 99.8% precision (easiest to detect — high volume)
- Probe: 97.2% precision (port scanning, etc.)
- R2L: 85.3% precision (remote to local — harder, fewer examples)
- U2R: 78.1% precision (user to root — rarest, most dangerous)
Key Challenges
1. Class Imbalance
U2R attacks are extremely rare (<1% of data). I addressed this with:
- SMOTE (Synthetic Minority Over-sampling)
- Class weights in Random Forest
- Evaluation by precision/recall rather than just accuracy
2. Feature Drift
Network traffic patterns change over time. A model trained on 2024 traffic may not work well in 2026. Continuous retraining is essential.
3. False Positives
In production, false positives are expensive — they cause alert fatigue. I tuned the decision threshold to minimize false positive rate while maintaining >95% detection rate.
Real-World Considerations
For production deployment, you'd need:
- 1.Packet capture — Tools like tcpdump or Wireshark to collect raw traffic
- 2.Feature extraction — Convert raw packets to the 41-feature format
- 3.Streaming inference — Process packets in real-time (not batch)
- 4.Alert integration — Feed detections into SIEM systems (Splunk, ELK)
- 5.Model monitoring — Track prediction confidence and retrain on drift
Impact
Network intrusion detection is a $6.2 billion industry (2025). ML-based approaches are increasingly replacing signature-based systems because:
- They detect zero-day attacks
- They adapt to new traffic patterns
- They reduce false positive rates with proper tuning
Source code: GitHub
