Assignment: Video Quality Inference¶

To this point in the class, you have learned various techniques for leading and analyzing packet captures of various types, generating features from those packet captures, and training and evaluating models using those features.

In this assignment, you will put all of this together, using a network traffic trace to train a model to automatically infer video quality of experience from a labeled traffic trace.

Getting the data¶

Everything you need is in the course data repository on GitHub, noise-courses/data, under video-qoe/. Either clone the whole repository (about 110 MB):

git clone https://github.com/noise-courses/data.git

or fetch just the files you need with the wget lines given in each part below. The four files are:

File Used in Size
netflix.pcap Part 1 22 MB
netflix_dataset.pkl.xz or video_dataset.pkl.xz Part 2 22 MB / 61 MB compressed
netflix_session.pkl Part 3 130 KB

Part 1: Warmup¶

The first part of this assignment builds directly on the hands-on activities but extends them slightly.

Extract Features from the Network Traffic¶

Load the netflix.pcap file, which is a packet trace that includes network traffic:

!wget https://raw.githubusercontent.com/noise-courses/data/main/video-qoe/netflix.pcap
In [ ]:
 

Identifying the Service Type¶

Use the DNS traffic to filter the packet trace for Netflix traffic.

In [ ]:
 

Generate Statistics¶

Generate statistics and features for the Netflix traffic flows. Use the netml library or any other technique that you choose to generate a set of features that you think would be good features for your model.

In [ ]:
 

Write a brief justification for the features that you have chosen.

Inferring Segment downloads¶

In addition to the features that you could generate using the netml library or similar, add to your feature vector a "segment downloads rate" feature, which indicates the number of video segments downloaded for a given time window.

Note: If you are using the netml library, generating features with SAMP style options may be useful, as this option gives you time windows, and you can then simply add the segment download rate to that existing dataframe.

In [ ]:
 

Part 2: Video Quality Inference¶

You will now load the complete video dataset from a previous study to train and test models based on these features to automatically infer the quality of a streaming video flow.

For this part of the assignment, you will need two pickle files. They are hosted in the course data repository on GitHub, noise-courses/data, in the video-qoe/ directory. Download them by running the code below:

!wget https://raw.githubusercontent.com/noise-courses/data/main/video-qoe/netflix_session.pkl
!wget https://raw.githubusercontent.com/noise-courses/data/main/video-qoe/video_dataset.pkl.xz
!xz -d video_dataset.pkl.xz

The dataset is stored xz-compressed in the repository (61 MB compressed, about 265 MB uncompressed), so the last line decompresses it. If xz is not available on your system, you can decompress it in Python instead:

import lzma, shutil
with lzma.open('video_dataset.pkl.xz') as f_in, open('video_dataset.pkl', 'wb') as f_out:
    shutil.copyfileobj(f_in, f_out)

A smaller Netflix-only dataset (netflix_dataset.pkl.xz, 52,279 samples with 251 features) is available in the same directory and may be used instead of the multi-service video_dataset.pkl (204,713 samples with 170 features).

Backup (Google Drive): If GitHub is unavailable, the Netflix-only dataset is also mirrored on Google Drive. This link only works through gdown (pip install gdown); opening it with wget, curl, or a browser returns a virus-scan HTML page instead of the file:

!gdown 'https://drive.google.com/uc?id=1iI1tos08p9FJBI-ADR9b_iFObWRwXANX' -O netflix_dataset.pkl

Load the File¶

Load the video dataset pickle file.

In [ ]:
 

Clean the File¶

  1. The dataset contains video resolutions that are not valid. Remove entries in the dataset that do not contain a valid video resolution. Valid resolutions are 240, 360, 480, 720, 1080.
In [ ]:
 
  1. The file also contains columns that are unnecessary (in fact, unhelpful!) for performing predictions. Identify those columns, and remove them.
In [ ]:
 

Briefly explain why you removed those columns.

Prepare Your Data¶

Prepare your data matrix, determine your features and labels, and perform a train-test split on your data.

In [ ]:
 

Train and Tune Your Model¶

  1. Select a model of your choice.
  2. Train the model using your training data.
In [ ]:
 

Tune Your Model¶

Perform hyperparameter tuning to find optimal parameters for your model.

In [ ]:
 

Evaluate Your Model¶

Evaluate your model accuracy according to the following metrics:

  1. Accuracy
  2. F1 Score
  3. Confusion Matrix
  4. ROC/AUC

Part 3: Predict the Ongoing Resolution of a Real Netflix Session¶

Now that you have your model, it's time to put it in practice!

Use a preprocessed Netflix video session to infer and plot the resolution at 10-second time intervals. The session is netflix_session.pkl (downloaded in Part 2 above; 59 ten-second windows with the same 251 features as netflix_dataset.pkl, plus the true resolution so you can check yourself):

session = pd.read_pickle('netflix_session.pkl')
In [ ]:
 

Acknowledgments¶

Please fill in the following before submitting.

Collaborators: (people you discussed the assignment with, or "none")

Outside sources and code: (documentation, blog posts, repositories, etc., with links, or "none")

AI tools used: (which tool, and roughly how, e.g., "Claude Code for the netml feature extraction and debugging the pickle load", or "none")

Prompts (optional): if you're willing, add a prompts.md next to this notebook with the prompts you used. This is not graded; it helps improve the course.