Logistic Regression on Network Traffic¶
Logistic Regression is very similar to linear regression, except all of the points can only have $y$-values of $1$ or $0$. This is useful if we want to predict whether something is or isn't part of a particular class. Instead of fitting a line (as in linear regression), logistic regression involves fitting a probability curve.
For example, using our device traffic, let's see whether we can predict a DNS packet is request or response from its length.
First, let's import the data, extract only the DNS packets, and view the first few packets.
# Pandas, Numpy
import numpy as np
import pandas as pd
import logging
logging.getLogger("scapy.runtime").setLevel(logging.ERROR)
# Machine Learning
from sklearn.linear_model import LogisticRegression, LogisticRegressionCV
# Plotting
import matplotlib.pyplot as plt
%matplotlib inline
import sys
#sys.path.insert(1,"/Users/feamster/research/netml/src/")
import netml
from netml.pparser.parser import PCAP
hpcap = PCAP('data/http.pcap', flow_ptks_thres=2, verbose=10)
hpcap.pcap2pandas()
pcap = hpcap.df
'_pcap2pandas()' starts at 2022-11-02 17:17:32 '_pcap2pandas()' ends at 2022-11-02 17:17:39 and takes 0.1196 mins.
Each row in the printed data is a packet and each column is a feature of the packet.
Next let's divide the DNS packets into requests and repsonses, and convert them into points where the $x$-value is the length of the packet and $y$-value is $0$ for requests and $1$ for responses. This will allow us to fit the data to a logistic regression curve.
Let's see how many data points we have.
Next we will convert the DNS response column into a 0/1 value so that it is amenable to logstic regression.
Part 2: Data Packets vs. Acknowledgments¶
In the linear-regression hands-on, we saw that flows split into two populations: flows dominated by full-size data packets and flows dominated by small acknowledgments (ACKs). That suggests a natural classification problem, which came up as an extension in class: given a single TCP packet, is it a data packet or an ACK?
Reuse the per-packet data frame from netml (hpcap.df above, or generate flow features as in the linear-regression hands-on). Define the label yourself, for example by treating a TCP packet with no payload (or a packet length below some threshold) as an ACK, and everything else as a data packet. Then train a logistic regression that predicts the label from packet size, and optionally from the inter-arrival time as a second feature.
- Build the feature matrix (packet size, optionally inter-arrival time) and the 0/1 label.
- Split into training and test sets, and fit a
LogisticRegression. - Report the accuracy on the test set. Is accuracy a reasonable metric here, given how the two classes are balanced?
- Plot the fitted sigmoid over packet size, with the labeled points overlaid, as you did for the DNS example above. Where is the decision boundary, and does it match the threshold you used to define the labels? (If it matches too well, think about whether the problem is circular, and what other feature might make it interesting.)
from sklearn.model_selection import train_test_split
# TODO: select TCP packets from the per-packet data frame
# tcp = pcap[pcap['protocol'] == 'TCP'].copy()
# TODO: define the label: 1 for a data packet, 0 for an ACK
# tcp['is_data'] = (tcp['length'] > ACK_THRESHOLD).astype(int)
# TODO: build features. Start with packet size; optionally add inter-arrival time.
# X = tcp[['length']].values
# y = tcp['is_data'].values
# TODO: train/test split and fit
# X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)
# clf = LogisticRegression().fit(X_train, y_train)
# print("Accuracy: {:.3f}".format(clf.score(X_test, y_test)))
# TODO: plot the sigmoid over packet size, with the labeled points overlaid
# sizes = np.linspace(X[:, 0].min(), X[:, 0].max(), 500).reshape(-1, 1)
# plt.scatter(X[:, 0], y, alpha=0.2, label='packets')
# plt.plot(sizes, clf.predict_proba(sizes)[:, 1], color='red', label='P(data packet)')
# plt.xlabel('Packet size (bytes)')
# plt.ylabel('P(data packet)')
# plt.legend()