Machine Learning Pipeline¶
Now that we have experience preparing data for input to machine learning libraries, the next step will be to train, tune, and test a model. You will perform all three of these steps in this hands-on activity.
The assignment consists of the following steps:
- Load two datasets and prepare their representations and labels for model input.
- Split the data into training and testing.
- Select a model, and identify the parameters to tune.
- Tune the model.
- Evaluate the model's performance.
import logging
logging.getLogger("scapy.runtime").setLevel(logging.ERROR)
from netml.pparser.parser import PCAP
from netml.utils.tool import dump_data, load_data
import pandas as pd
Convert the Packet Capture Into Flows¶
- Load the two packet captures for HTTP requests and Log4j scan,
- convert them into traffic flows,
- generate features from the flow,
- label the traffic,
- normalize your labeled features into a 2D matrix
Evaluating a Machine Learning Model¶
The goal of supervised learning is to train a model that takes examples and predicts labels for these examples that are as close as possible to the actual labels. For instance, in this example above, a model might take features from a traffic trace and predict whether the traffic constitutes regular web traffic or a scan.
How do you measure whether the model is succeeding if you don't know the true labels for new observations? The way to solve this problem is to test the performance of the trained algorithm on additional data that it has never seen, but for which you already know the correct labels.
This requires that you train the algorithm using only a portion of the entire labeled dataset (the training set) and withold the rest of the labeled data (the test set) for testing how well the model generalizes to new information.
To evaluate the model, we will need to split the data into train and test sets.
Split into Training and Test Sets¶
Split your data into a training and test set using scikit-learn. A common split is to train on 80% of your data, while withholding 20% of the data.
Training Your Model¶
Now that you have split your data into training and testing sets, you are ready to train and evaluate a model.
Import a machine learning model of your choice, use your training set to train the model, and use the test set to evaluate it.
Test Your Trained Model¶
You can now evaluate how well your trained model works. There are several valuable ways to visualize your results. You might use techniques such as a confusion matrix, or a receiver operating characteristic (ROC) curve. Below we will gain some experience plotting both of those. This documentation may help you with plotting these results.
Confusion Matrix¶
A confusion matrix is a one way to understand errors of different types. We can see a lot of examples off diagonal, suggesting a fair number of incorrect answers.
Receiver Operating Characteristic¶
Some models can output different classes based on a threshold that is set for the decision.
Area Under the Curve (AUC)¶
From the ROC above, you can also compute a metric called the area under the curve (AUC). Visually, this is the area under the curve that you just plotted. You could see, intuitively, that the "best" performance should yield an AUC of 1, and the worst performance would yield an AUC closer to 0.5.
Scikit learn also has a function for computing AUC. Compute the area under the curve.
Precision–Recall Curve and Average Precision¶
The ROC curve plots the true positive rate against the false positive rate as the decision threshold varies. A related plot is the precision–recall (PR) curve, which plots precision (of the examples the model flagged as positive, what fraction really are positive?) against recall (of the truly positive examples, what fraction did the model flag?) as the threshold varies.
The PR curve is often preferred when the positive class is rare, which is the common case in systems and security problems: attacks, spam, and failures are a tiny fraction of all traffic. With a rare positive class, the false positive rate stays small even when the model raises a large number of false alarms, because the denominator (all negatives) is huge. The ROC curve can therefore look excellent while an operator sees mostly false alarms. Precision, in contrast, is computed only over the examples the model flagged, so the PR curve exposes exactly that problem.
The single-number summary of the PR curve is average precision (AP), which is analogous to AUC for the ROC curve. A random classifier has an AP equal to the fraction of positive examples, not 0.5.
Plot the PR curve for your model using precision_recall_curve and PrecisionRecallDisplay, and compute the average precision with average_precision_score. Compare the shape of the PR curve and the AP value to the ROC curve and AUC that you computed above.
from sklearn.metrics import precision_recall_curve, average_precision_score, PrecisionRecallDisplay
# TODO: get scores for the positive class on the test set
# (e.g., clf.predict_proba(X_test)[:, 1] or clf.decision_function(X_test))
y_score = None
# TODO: compute precision, recall, and thresholds
# precision, recall, thresholds = precision_recall_curve(y_test, y_score)
# TODO: compute average precision
# ap = average_precision_score(y_test, y_score)
# print("Average precision: {:.3f}".format(ap))
# TODO: plot the precision-recall curve
# PrecisionRecallDisplay(precision=precision, recall=recall, average_precision=ap).plot()
# TODO: what fraction of the test examples are positive? How does that compare to the AP
# a random classifier would achieve?
Thought Question¶
Which evaluation model is more appropriate, and when (i.e., under what circumstances)? When might you care more about looking at the confusion matrix (or model accuracy) vs. the ROC, or the area under the curve?
When is the precision–recall curve more informative than the ROC curve? Think about what happens to each curve as the positive class becomes rarer, and about what an operator who has to investigate every alarm actually cares about.