Machine Learning for Computer Systems
This is a recitable script for the Lecture 23 deck, one row per content slide, written in the first person so it can be read more or less verbatim. The deck’s own ::: {.notes} blocks are shorter cues and are visible in the reveal.js speaker view (press S while presenting); this file is the longer version for teaching the four papers without much prior context. Every URL below was verified this session (arXiv abstract pages fetched on 2026-09-16); all numbers come from the papers’ LaTeX sources. Where a figure ships with the arXiv source but was cut from the published paper, the row says so.
| Slide | Context + Links |
|---|---|
| Roadmap | The point: set expectations. This is the closer, it is long, and it is built around real results from four papers rather than toy examples, so tell them up front what is exam-relevant and what is there for depth. Walk the six items. One: why deployed models decay, the pipeline’s last loop and the kinds of drift networks see. Two: LEAF, Liu and colleagues at CoNEXT 2023, four years of cellular KPIs, why periodic retraining fails, and a detect-explain-mitigate framework. Three: the systems cost of staying accurate, CATO, where the feature representation is a cost knob and the output is a Pareto front. Four: adapting at serving time, AC-DC’s pool of classifiers with a scheduler, and ServeFlow’s cheap-model-first cascade. Five: a maintenance playbook you can apply to your project. Six: summary, with the math in the appendix. Say which parts matter for the exam: the ideas in parts one and five and the LEAF story. The CATO and AC-DC numbers are for depth and for anyone building a serving pipeline. So what: the whole quarter ended at “evaluate on the held-out set.” Today is about what happens after, and the running theme is that every fix has a cost, in retrains, in latency, in memory, and maintenance is managing those budgets. One housekeeping note: the four papers are from this group, so questions about what did not make it into print are fair game. |
| The Pipeline’s Last Loop | The point of this slide: the pipeline we built all quarter stops at “deploy,” and the part that actually costs money in an operational network is everything after that. Walk the two columns. Left side is the four steps you already know: measure, represent, model, deploy. Right side is the loop we have mostly ignored: monitor, diagnose, retrain, redeploy. Draw the arrow from step 8 back to step 4 on the board, because that arrow is the whole lecture. Then ask the room: “What changes in a network between the day you train and the day you deploy?” Let them list things. Firmware upgrades, new base stations, new handsets, a new app that everyone starts using, an adversary who read your paper, the seasons, a pandemic. Every one of those is a reason the model you validated on a held-out set is no longer looking at the same distribution. So what: the loop has to be engineered. A cron job that retrains every Sunday is not a strategy, and one of the first results today, from LEAF, is that calendar retraining can actually make things worse. The three questions in steps 5 to 7 are the questions each paper answers: how do you notice, how do you find out where and why, and what does it cost to fix. Keep a running tally of cost on the board; I will add to it as we go. |
| We Have Seen This Before | The point: drift is not a new topic for this course, it is the thing hiding behind several examples we already discussed. Go down the list and tie each one to a lecture. The COVID traffic shift from the data lecture: residential upstream jumped, campus traffic vanished overnight. That is a distribution shift nobody’s training set anticipated. The security lecture: an intrusion detector trained on last year’s botnet is looking at this year’s evasions, so the adversary is a drift generator. The spurious-feature example from the features and nPrint lectures: a classifier learned that a particular TTL and a particular source subnet meant “malicious,” because all the malware in the capture came from one lab machine. That feature drifts the instant the model leaves the lab. And from the time-series lecture, look-ahead bias and seasonality: a model tuned in the summer is wrong in the fall. The common thread in the box at the bottom is the sentence I want them to remember: the test set is never the deployment set, and time is a distribution shift you cannot shuffle away. Random train-test splits hide this by construction. Ask: “How many of you split your project data randomly rather than by time?” Most hands go up, and that is fine, but today is about what that choice conceals. LEAF, which we get to in a few slides, measures exactly what COVID did to a forecasting model over four years. |
| What Is Concept Drift? | The point: three different things can move underneath a deployed model, and they are detected in different ways, which is why the distinction is worth two minutes. Left column, say it in words rather than symbols. Data drift, also called covariate shift: the inputs change but the rule mapping inputs to outputs does not. Concept drift proper: the same inputs now deserve a different answer. Label drift, or prior shift: the base rate changes, for example attacks become more common. If you want the formal version for the whiteboard: change in P of X, change in P of Y given X, change in P of Y. Right column is the operational consequence. Data drift you can detect on day one without labels, just compare input distributions to the training window. Concept drift is only visible through errors, and errors need ground truth, which in a forecasting task arrives only when the horizon elapses. Label drift changes what “good precision” means, since precision depends on the base rate. Now the LEAF caveat: the paper uses “concept drift” loosely for any degradation in model accuracy over time, and it detects drift from the error stream. LEAF forecasts 180 days ahead (LEAF §2.2), so the error for today’s forecast is not known for six months. Hold that thought, because it comes back on the vignette slide as a design problem. |
| A Taxonomy of Drift | The point: drift comes in shapes, and the shape determines whether periodic retraining has any chance of working. The four panels are real per-eNodeB downlink-volume series from the LEAF dataset. Be honest about provenance: these images ship with the arXiv LaTeX source but were cut from the CoNEXT version, so students will not find them in the published paper. Walk left to right. Sudden: a one-step collapse, think a decommissioned sector or a software change that redefines a counter. Gradual: the level rises slowly while the weekly cycle stays intact, think population growth in a suburb. Incremental: a steady trend, think capacity additions. Recurring: a regime that comes and goes, think a stadium on game days or a campus by semester. Point at the 7-day ripple in every panel: that is seasonality, and it is the reason a model can be right on average and wrong every weekend. Ask the room: “Which of these does retrain-every-N-days handle well?” Gradual and incremental, because recent data really is more representative. Sudden is handled badly if the schedule misses the change by weeks. Recurring is the trap: the “new” data is actually the old regime coming back, so retraining on the last two weeks throws away exactly the history you need. LEAF’s Fig. 2b makes this concrete in a couple of slides. |
| Why Networks Are Especially Drift-Prone | The point: LEAF’s argument for why the drift literature from images and text does not transfer to networks, and why the operator’s 180-day horizon is not negotiable. Run down the bullets. Periodicity: diurnal and weekly cycles in volume, users, and call quality. Gradual evolution: new base stations, capacity upgrades, new handsets. Exogenous shocks: software upgrades, KPI redefinitions, a pandemic. Environment: weather and interference change the radio channel itself, which is a drift source with no analogue in image classification. And the last one is the deep difference: predictions are continuous. It is not one object classified once, it is a whole system forecast every day, 180 days ahead, per base station (LEAF §1 and §2.2). The quote at the bottom is the contrast to say out loud: drift in images is about slowly changing objects; drift in networks is about a system that changes underneath the model while the model keeps running. So what for the operator: why 180 days? Because that is how long it takes to plan, permit, and install infrastructure (LEAF §2.2). A model that is accurate for next week is useless for the capacity decision, and a model that is wrong six months out can justify building towers nobody needs. That asymmetry between over- and under-forecasting comes back in the LEAgram slide. |
| The Setting: Four Years of Cellular KPIs | The point: this is LEAF’s Table 1, and the thing to emphasize is duration, because four years is what makes the rest of the paper possible. Walk the table. Collection period January 1, 2018 to March 28, 2022. One row per eNodeB per day, identified by eNodeB ID and timestamp. 224 KPIs in three groups: resource utilization, network performance, user experience. Two datasets: Fixed, the same 412 eNodeBs throughout, with 699,381 daily logs, and Evolving, up to 898 eNodeBs as the network grew, with 1,084,837 logs. A major US carrier, one large metro area with urban, suburban, and rural sites (LEAF §2.1). Right column: what an eNodeB log is. These are the counters the operator already collects, things like data volume, peak active users, throughput, connection-establishment success, drop rates. So the model is forecasting the network from the network’s own telemetry, no extra instrumentation. Why two datasets? Fixed gives an apples-to-apples view of internal drift: software upgrades, user behavior, COVID. Evolving adds infrastructure drift: brand-new sites with no history. Table 5 later shows the schemes behave differently on the two, so keep the distinction in mind. So what: four years contains a sudden shock, a slow recovery, and a gradual rise peaking in January 2022, and it is long enough to compare retraining policies whose periods are measured in months. Most drift papers cannot do that. |
| Six Forecasting Targets | The point: LEAF Table 2. Six KPIs, two per group, and the single most important row is dispersion, because it predicts which KPIs will be hard to keep accurate. Read the columns: DVol downlink volume and PU peak active UEs from resource utilization; DTP downlink throughput and REst RRC-establishment success from network performance; CDR call-drop rate and GDR RTP gap-duration ratio from user experience. Std over mean on the Evolving dataset: 0.81, 1.76, 0.59, 0.85, 2.48, 8.52. Point at GDR’s 8.52 and CDR’s 2.48 and say: remember these two, they are the bursty ones, and every mitigation scheme struggles on them. The “Data Lost” check under PU flags that some PU data went missing between July 2019 and January 2020, which shows up as a huge error spike on the next slide. Bottom left, the task: regression, all KPIs up to today as features, forecast each target 180 days ahead per eNodeB, one model per KPI. Bottom right, the metric: NRMSE, RMSE divided by the KPI’s range, so a drop-rate model with values under 1 and a volume model with values over 300,000 are comparable. Under 0.1 counts as good (LEAF §2.3, citing the standard regression rule of thumb). The definition is in the appendix if anyone wants it. Ask: “If dispersion is 8.5, what does a 14-day training window even capture?” |
| Drift Is Everywhere, Regardless of Model | The point: LEAF Fig. 1a and 1b. The drift is a property of the data, not of the model you picked. Walk Fig. 1a, downlink volume. X-axis is date, y-axis is daily NRMSE. Four lines, four model families: boosting (CatBoost), bagging (Extra Trees), distance-based (KNeighbors), and an LSTM. All trained once on a 90-day window ending July 1, 2018, forecasting from late December 2018, plotted from mid-March 2019 because of a data gap. Point at April 2020: every line rises together, that is the lockdown, and they stay elevated through October 2020. Then from March 2021 a gradual climb that peaks around January 2022 (LEAF §3.2, §3.3). The inset is a three-week window: the error itself has a weekly rhythm, not just the traffic. Fig. 1b, peak active UEs: same models, and the giant spike in the middle is the mid-2019 data loss, which is a data-quality event masquerading as drift. Two takeaways from the paper. First, because the families behave alike, the authors use CatBoost for the rest of the paper. Second, these are not bad models: for every KPI, at least one model has NRMSE under 0.1 for at least a year, CatBoost on volume averages 0.116 and is under 0.1 on 473 days (LEAF §3.2). They are good models decaying. So what: swapping the algorithm does not fix drift. That is why LEAF is model-agnostic. |
| Drift Patterns Differ by KPI | The point: uncorrelated KPIs drift on different schedules with different textures, so there is no single drift signature a monitor can look for. Four panels, same models and training window as the previous slide, but note the y-axes differ. Fig. 1c, downlink throughput, drifts roughly like volume. Fig. 1d, gap-duration ratio: frequent, short-lived error bursts, and no weekly structure at all. Fig. 7a, RRC establishment, from the appendix. Fig. 7b, call-drop rate: high error early on that actually improves after April 2020, the opposite direction from volume. The paper backs the “different texture” claim with signal processing: an STFT of the NRMSE series finds no weekly pattern for CDR and GDR, but clear weekly patterns for the others (LEAF §3.2). The insets on Fig. 1 show the same thing visually. And the timing differs: volume drifts in April 2020, PU is worst July to November 2019 because of the data loss, CDR gets better when everyone else gets worse. So what for a monitoring system: detection has to be per-model and data-driven. You cannot hard-code “alarm if error rises in the spring.” This is the argument for a nonparametric detector on each model’s own error stream, which is what KSWIN does a few slides from now. Ask: “Why might call-drop rate improve during lockdown?” Fewer people moving, fewer handovers, fewer drops. The model did not get smarter, the network got easier. |
| Drift Persists Regardless of Training Set Size or Period | The point: LEAF Fig. 2. Two knobs people reach for first, more data and fresher data, and neither one removes the drift. Fig. 2a varies training-set size from one week to one year, all ending July 1, 2018, for downlink volume with CatBoost. Every curve drifts in lockstep: same jump after lockdown, same recovery after October 2020, same rise from March 2021 to January 2022. One week is slightly worse; three months and one year have a glitch around June 2020. The practical result is that two weeks performs about like one year and trains 18 times faster, so the paper uses 14-day windows from here on (LEAF §3.3). Write “18x” on the cost tally. Fig. 2b fixes the window at 14 days and slides the period. The finding: models trained on more recent periods do not necessarily do better. Why? After a sudden change, the new distribution can look more like the older past than the recent past (LEAF §3.4). Think of the recurring-drift panel from the taxonomy slide. So what: this is the setup for the retraining result. The industry default is “retrain on the last two weeks every N days.” If recent does not mean accurate, that policy has no guarantee of helping, and Table 3 on the next slide shows it sometimes hurts. Ask: “If you could pick any 14-day window to train on today, which would you pick?” You cannot know without seeing the future, which is the point. |
| Naive Periodic Retraining Is Not Enough | The point: LEAF Table 3. This is the “conventional practice is wrong” slide, so slow down. Setup: CatBoost, 14-day training window, 180-day forecast, retrained every N days on the latest 14 days and evaluated for the next N days. Each cell is the change in average NRMSE relative to a model that was never retrained, so negative is better. Read across the 7-day row: DVol 40.34 percent better, PU 55.36 percent better, DTP 27.21, REst 48.00. Great, but look at the last column: 169 retrains over the dataset. Then CDR in the same row: 47.79 percent worse. Weekly retraining makes call-drop forecasts worse by half. GDR: 90 days is 42.24 percent worse, 180 days is 76.28 percent worse, and 365 days is better than 90 days. Cold-call: “You are the operator. Which row do you pick for CDR?” There is no good row. That is the point. Why does naive retraining fail? Three reasons from LEAF §3.4. It fires on a calendar, but drift is irregular, so it retrains when nothing changed and misses the moment something did. It assumes recent data is best, and the previous slide showed it is not. And it replaces the whole training set, throwing away useful old samples and any fine-grained information about which samples were wrong. So what: retraining is essential, but when and on what data are the design variables. LEAF is an answer to both. |
| The LEAF Framework | The point: LEAF Fig. 3. Three stages, detect, explain, mitigate, and the design constraint that makes it deployable is that it never opens the model. Walk the figure left to right. The in-use model produces predictions; ground truth arrives; the NRMSE stream feeds a detector. When the detector fires, the explainer looks at the model’s inputs and outputs to find which features, where in their range, and how much the error has moved. The mitigator uses that to restructure the training set and retrain. The gray box with the star is the paper’s contribution: explanation and informed mitigation. The detector is deliberately off-the-shelf. Model-agnostic means exactly this: LEAF needs the previous training set, the new data as it arrives, and the model’s outputs. Nothing internal (LEAF §4.1). That matters in practice because the operator’s models here were picked by AutoGluon, an auto-ML pipeline, so there is no single architecture to instrument. So what, tying back to interpretability: this is the same posture as partial-dependence plots from the trees and ensembles lecture, black-box in, explanation out. The twist is that LEAF explains error rather than prediction. The three questions from the pipeline slide map onto the three boxes: monitor is detect, diagnose is explain, retrain is mitigate. Ask the room which box they think is the hard one. Most say detect. The paper says detect is solved; the research is in the other two. |
| Step 1: Detect Drift from the Error Stream | The point: the detector is a two-sample test on a sliding window of error values, and it is the trigger for everything downstream. Right column first. KSWIN, Kolmogorov-Smirnov Windowing: keep a window of recent NRMSE values, compare the newest sub-window against the older part with a KS test, and alarm if the error distribution has changed. Nonparametric, so no assumption about the error distribution, which matters because these series are heavy-tailed and bursty. The paper also tried ADWIN, DDM, HDDM, EDDM, and Page-Hinkley; KSWIN worked best on these series (LEAF §4.1, footnote). Left figure: alarms in red on the NRMSE stream of a downlink-volume CatBoost model trained on 14 days before July 1, 2018. Be honest about provenance: this annotated figure is from the authors’ source files; the published paper describes KSWIN in Appendix B in prose without a figure. There is no ground-truth “drift” label in the dataset, so the authors validate against known events: the mid-2019 data glitch, December 2019, April 2021, and the beginning and end of the quarantine period (LEAF Appendix B). Two subtleties to raise. First, the input is error, which requires ground truth, so with a 180-day horizon the alarm is inherently delayed. Second, a detector alone only tells you when. It says nothing about which data to retrain on, and “triggered retraining” is the baseline that isolates that gap later. The KS statistic is in the appendix. |
| Step 2: Explain Drift with Local Error Approximation | The point: LEA is a partial-dependence plot for error instead of prediction, and the four steps are how you get one for a black-box regressor with 224 features. Walk the steps. One: rank features by permutation importance, shuffle a column and see how much the model degrades. Two: group correlated features and keep the top-ranked feature of each group as its representative. Three: for a representative feature, split samples into N quantile bins. Four: compute NRMSE per bin, giving a vector that says where in feature space the model is wrong (LEAF §4.2). Why grouping? A 224-KPI dataset is full of near-duplicates: bytes, packets, sessions, and establishments all move together. If you mitigated per feature you would double-count the same underlying cause. In the case study, Group 1 has 32 features and its representative is the history of downlink volume itself, which the authors call a sanity check: the best predictor of volume is past volume (LEAF §5). The connection to say out loud: PDP and ALE from the trees lecture plot the model’s output against a feature. LEA plots the model’s error against a feature. Same machinery, different target, and the reason it is useful is that the error vector becomes a sampling weight in the mitigation step. Ask: “Why quantile bins rather than equal-width bins?” So that every bin has enough samples for an NRMSE estimate, even in the tail. |
| LEAplot: Where in Feature Space the Error Lives | The point: LEAF Fig. 4 and Fig. 8. The LEAplot answers “where is the model under-trained?” and the two panels show two different causes. Left, Group 1 representative, the history of downlink volume, 1,000 bins. X-axis is feature value, y-axis is local error. Three curves: training set, full test set, and the “Early 2022” test set that triggered the alarm. Point at the range 0.6e6 to 1.3e6: the Early 2022 error is more than 10 times the training error and nearly 6 times the full-test error. Then point above 1.5e6: the training set simply has no coverage there, so error is high for both test sets (LEAF §5). Right, Group 2 representative, bad-coverage measurements, from the appendix. Above 2e5 the model is poorly fit, again because training never saw that range. Say it plainly: the high-error region of Group 1 is not the same set of samples as the high-error region of Group 2, and that difference is what makes multi-group mitigation possible. Tracing the top 5 percent of errors back to eNodeBs: mostly suburban sites whose mobility patterns changed after the winter break. The third representative feature was RTP gap ratio, essentially VoLTE packet loss, which affects delivered volume. So what: both problems, stale coverage and no coverage, are fixed with data, not architecture. That is why the mitigator resamples rather than retunes. |
| LEAgram: Error Over Time and Feature Value | The point: LEAF Fig. 5. Add time to the LEAplot and keep the sign of the error, and you get the explanation an operator can actually act on. Left panel, before mitigation. Each dot is one eNodeB-day, x is test date, y is the value of the downlink-volume feature, color is signed normalized error, red is overestimation. Point at the block from March 15 to November 1, 2020, at feature values above 1e6: large positive errors. The model kept overestimating demand at high-volume sites because users had moved to home Wi-Fi during lockdown (LEAF §5). Then overestimation returns after October 2021 above 0.6e6. There are also negative errors above 1e6, underestimation, which the paper flags as the user-dissatisfaction case. Right panel, after LEAF: the red blocks largely disappear, and average NRMSE falls 32.68 percent. Ask students to read what remains: mostly the tail, plus some vertical stripes that are data glitches on specific dates. Why sign matters, in the operator’s terms: overestimate and you build capacity you do not need, that is capital expenditure wasted; underestimate and users suffer. Those are different costs, and an unsigned metric like NRMSE hides which one you are paying. That is why the LEAgram switches from NRMSE to normalized error (LEAF §4.2). So what: this is the diagnose step done right. A red block with a date range and a site class is a story an operator can verify: lockdown, suburban sites, Wi-Fi offload. |
| Step 3: Mitigate with Forgetting and Over-Sampling | The point: retrain only on alarm, and choose the retraining set by where the error is rather than just when. Left column. Forget: from the last training set, drop samples that sit in high-error bins but carry low weight, where weights come from the LEA error vector of the newest data. Over-sample: from all data collected so far, including the drifting samples, draw extra samples in proportion to the error vector. Then retrain on the restructured set, and on the next alarm repeat starting from this set, so the process is cumulative (LEAF §4.3). Right column: how aggressive, decided by dispersion. High-dispersion KPIs, PU, CDR, GDR with std over mean above 1: linear forgetting weights and cubic over-sampling weights concentrated on the high-error regions. Low-dispersion KPIs, DVol, DTP, REst: drop samples with over 95 percent error and over-sample with linear weights. Say clearly that the paper states these thresholds are tuned to this dataset and may need adjusting elsewhere. Cold-call: “Why cubic weights for the bursty KPIs?” Because their error regions are sparse and extreme; you want to pile samples into the tail rather than spread them. The key contrast with naive retraining, which the evaluation controls for: the same amount of data per retrain, but chosen to cover the under-trained region instead of “the last 14 days.” Multi-group LEAF repeats the resampling for the top three or five representative features. This is the difference between fixing when and fixing what. |
| Cost vs. Accuracy: The Retraining Trade-Off | The point: LEAF Fig. 6. This is the slide that frames the whole closing question of the course: accuracy per retrain. Axes: x is number of retrains over the dataset, which is cost; y is change in average NRMSE versus a static model, which is accuracy. Bottom-left wins. Schemes: naive retraining every 30 or 90 days (the two best periods from Table 3), triggered retraining, which retrains on the latest data whenever KSWIN fires, and LEAF with 1, 3, or 5 feature groups. Same data volume per retrain across schemes, CatBoost throughout (LEAF §6.1). Fig. 6a, downlink volume: naive-30 needs the most retrains and never beats LEAF; naive-90 is cheapest but sits top-left, least mitigation; triggered lands in the middle; LEAF’s three points cluster bottom-left. Fig. 6b, call-drop rate, the hardest KPI: naive-90 looks competitive but does not reach the best mitigation, and LEAF matches triggered’s error with 30.8 percent fewer retrains. Why triggered is the important baseline: it is what you get if you “add a detector” to the Sunday cron job. It fixes when, not what data. For GDR it makes error worse. LEAF fixes both. So what: an operator could not have known in advance that naive-90 was decent for CDR without running every scheme for four years. A method that is consistently bottom-left is worth more than one that is occasionally best. |
| The Trade-Off for the Other Four KPIs | The point: appendix Fig. 10, the same trade-off plot for PU, DTP, REst, and GDR. Small on a slide; the point is the shape, not the numbers. Same axes as before: retrains on x, change in average NRMSE on y, bottom-left wins. Parentheses give the naive period in days or the number of LEAF feature groups. Point at the GDR panel on the right: triggered retraining is far up the y-axis, 44.56 percent worse than static, and only LEAF sits below zero. That is the clearest picture in the paper of a detector-plus-recent-data policy actively harming a bursty target. The multi-group result from LEAF §6.2: using more feature groups mitigates 0.34 to 2.83 percent more error than single-group LEAF, but the best count differs by KPI. Three groups for DVol, PU, and REst; five for DTP and CDR; one for GDR, the most dispersed KPI, which is the outlier. The paper’s honest guess is that feature importance and group size drive this, and it is not settled. So what: even the winning method has a hyperparameter you tune per target. There is no universal setting, which echoes the “no single drift signature” point from earlier. If time is short, skip this slide and refer to Table 4 on the next two slides, which carries the same information numerically. |
| Results Across Models and KPIs: Tree Ensembles | The point: LEAF Table 4, first half. LEAF is best or near-best on trees, and, more importantly, it is never worse than doing nothing. Explain the cell format: change in average NRMSE versus static, number of retrains in parentheses, Fixed dataset, bold is the best per row. Read one row aloud, CatBoost on GDR: naive-30 makes things 3.37 percent worse with 39 retrains, triggered makes them 44.56 percent worse with 17, LEAF makes them 6.24 percent better with 19. Then scan the LEAF column: every entry is negative. The baselines cannot say that; naive and triggered both go positive on CDR and GDR (LEAF §6.2). Retrain savings versus naive-30: 10.3 to 76.9 percent fewer for CatBoost, 17 to 71.8 percent fewer for ExtraTrees. The paper reports the same pattern for LightGBM and Random Forest in text without a table. One honest footnote from the appendix’s full Table 8, which the slide omits for space: naive retraining every 90 days is actually the best scheme for CatBoost on CDR, 5.39 percent better with 13 retrains, versus LEAF’s 3.63 percent with 9. So on the one KPI where a calendar happens to line up with the drift, the calendar wins on accuracy. That does not undermine the story, but say it if a sharp student asks. So what: “never worse than static” is the property an operator needs before automating retraining. |
| Results Across Models and KPIs: LSTM and KNN | The point: LEAF Table 4, second half. Two teaching points: deep sequence models benefit most from targeted data, and distance-based models are the failure case. LSTM rows first. Naive retraining on recent data often hurts: PU is 37.11 percent worse, DTP 17.08 percent worse. Triggered is mixed. LEAF is best on four of six and wins by wide margins: DTP 37.13 percent better, and the headline, CDR 71.52 percent better with only 11 retrains. The paper’s summary is that LEAF gives the LSTM 7.71 to 50.13 percent less NRMSE than triggered retraining where it wins (LEAF §6.2). The intuition: a sequence model fit to two recent weeks overfits that regime; resampling across the whole history covers the regime that actually recurs. KNeighbors rows: LEAF is best only on CDR, and naive-30 or triggered wins elsewhere. The paper’s explanation: KNN is a lazy regressor that memorizes the training set and predicts from nearest neighbors. Over-sampling one region of feature space shifts the neighbor density and degrades the regions that were fine. Bagging and boosting train learners that absorb new samples more independently, so targeted resampling does not contaminate the good regions. So what: mitigation strategy has to match model family. That is a general lesson beyond LEAF: any data-side fix interacts with the inductive bias of the model, so you validate the fix per model, not once. Ask: “What would you do for KNN instead?” Reweight distances, or retrain on a balanced sample rather than an over-sampled one. |
| LEAF Over Time: Before and After | The point: appendix Fig. 9 and Table 7. Most of the error lives in the tail, so the 95th percentile is the honest metric, and it shows both where LEAF wins big and where it barely moves. Table first: 95th-percentile normalized error, CatBoost, static versus LEAF. DVol 0.29 to 0.19. PU 0.86 to 0.27, that is the headline. DTP 0.17 to 0.13. REst 0.33 to 0.18. CDR 0.24 to 0.23. GDR 0.27 to 0.27, no change. The paper’s explanation is dispersion: on the Fixed dataset the coefficients of variation for PU, CDR, and GDR are 1.34, 1.35, and 2.12, two to four times those of the other KPIs (LEAF Appendix C). Now the three time-series panels, static versus LEAF, blue is LEAF. Fig. 9a, DVol: LEAF stays under 0.125 except for data-error spikes and goes as low as 0.02, so the error is bounded, not just lower on average; average NRMSE down 32.67 percent. Fig. 9b, PU: after the mid-2019 data loss, the static model’s error exceeds 1 for more than half a year; LEAF’s does so for about a week, because the next alarm triggers a retrain. Fig. 9d, REst: sudden changes around January and April 2020 and the gradual 2021 rise are both handled. Say clearly: LEAF does not solve bursty KPIs. Fig. 9f, not shown, still has two GDR spikes around December 2020 and August 2021. What LEAF does is stop making them worse. |
| Does It Hold When the Network Grows? (1/2) | The point: LEAF Table 5, Fixed-dataset half. On the same 412 eNodeBs throughout, LEAF or its multi-group variant is the best scheme on every single KPI. Columns: triggered retraining, single-group LEAF, and LEAF-star, the best multi-group configuration per KPI. Cells are change in average NRMSE versus static, retrains in parentheses, CatBoost throughout. Read down the LEAF-star column: DVol 35.12 percent better with 34 retrains, PU 47.62 with 27, DTP 24.64 with 30, REst 41.27 with 32, CDR 6.22 with 12, GDR 6.24 with 19. Now compare to triggered: it is close on DVol, DTP, and REst, notably worse on PU, and on GDR it is 44.56 percent worse than doing nothing (LEAF §6.3). Point at the retrain counts: LEAF-star is not buying accuracy with more retrains. On PU it uses 27 versus single-group LEAF’s 35 and still does better, because the multi-group resampling picks better data rather than more data. So what: this is the apples-to-apples case, internal drift only, software upgrades, user behavior, COVID. It is the cleanest evidence that choosing what data to retrain on beats choosing only when. Ask: “Why is the CDR row the weakest for everyone?” Dispersion again; the next slide shows a case where the picture changes. |
| Does It Hold When the Network Grows? (2/2) | The point: LEAF Table 5, Evolving-dataset half. New sites with no history are a different kind of drift, and any scheme with a detector benefits from it, while LEAF stays the most consistent. Same columns and format. Compare triggered against the previous slide: PU goes from 35.06 percent better on Fixed to 50.89 percent better here; CDR from 4.21 to 6.79; GDR from 44.56 percent worse to 13.21 percent better, which makes triggered the winner on GDR. The paper’s reading: a timely detector captures newly deployed eNodeBs as they come online, because “new site appears” is exactly the sudden change a KS test catches, and for a site with no history the most recent data really is the right data (LEAF §6.3). LEAF-star still takes five of six rows: DVol 32.80, PU 51.72, DTP 22.58, REst 48.01, CDR 7.15 percent better. Its edge narrows but holds. So what for practitioners, the sentence to write down: a detector alone buys you a lot when the drift is “new entities”; targeted resampling buys you more when the drift is “old entities behaving differently.” Real networks have both, which is why you want both. Ask: “Why does LEAF’s advantage shrink here?” A new site has no error history to explain, so the explainer has less to work with, and single-group LEAF on CDR and GDR shows it. |
| When the World Changed Overnight | The point: the deck’s dated real-world hook, drawn from the paper itself rather than a news story. Say so: a search in September 2026 did not turn up a verifiable public post-mortem of a production network model failing from drift, so I am using the best-documented case I have. Tell it as a story. April 2020, a major US cellular network. Every downlink-volume forecasting model in the four-year dataset, boosting, bagging, k-NN, LSTM, jumps in error as lockdowns begin, stays elevated through October 2020, then recovers (LEAF Fig. 1a). The LEAgram says why: at high-volume sites the models kept overestimating demand because users had moved to home Wi-Fi (LEAF Fig. 5a, §5). An operator trusting those forecasts would have provisioned capacity for traffic that never came. The same detector also caught the mid-2019 data loss and the gradual drift that peaked in January 2022. Then the thought question, and give it real time. The forecast horizon is 180 days. The first day you can measure the April 2020 error is October 2020. What should a monitoring system have watched instead? Let them work toward it: the inputs. The KPIs themselves shifted in April 2020 and that was visible immediately, no labels needed. Also the prediction distribution: if the model’s forecasts for six months out suddenly look nothing like last month’s, that is a signal. And agreement between the incumbent model and a freshly trained canary. So what: this is the bridge to the playbook. With delayed labels you need label-free monitors too. |
| Accuracy Is Not the Only Objective | The point: CATO Fig. 1, and the other half of maintenance. Even a perfectly accurate model is useless if the pipeline that feeds it drops packets. Walk the figure: packet capture, connection tracking and reassembly, feature extraction, then inference. Say the surprising part: the model is the last and often the cheapest stage. The bullets give the constraints from CATO §1 and §2: real-time traffic analysis needs sub-second reactions and hundreds of gigabits per second without loss, and models built without measuring serving cost, quoting the paper, “often turn out to be unusable in practice.” Credit Traffic Refinery, Bronzino and colleagues 2021, as the paper that made this case first. It profiled the cost of feature classes, packet counters, timing, TCP counters, and let the operator trade accuracy for cost. But the exploration was by hand. CATO automates the search. Connect to LEAF explicitly: LEAF’s cost was retrains; CATO’s cost is serving. Same operator, two budgets. And they interact: every retrain that changes the feature set changes the pipeline cost, which I will come back to on the wall-clock slide. Ask the room: “In your project, did anyone measure how long feature extraction took?” Usually nobody. That is the gap this section is about. |
| Representation Is a Cost Knob | The point: CATO Fig. 2. Two knobs, which features and how many packets to wait for, and they interact non-linearly, which is why you cannot optimize them separately. Setup from CATO §2.2: IoT device classification on the Sivanathan dataset, six candidate features, packet depth from 1 to 50, and the authors trained, compiled, and measured every one of the 2 to the 6 times 50 equals 3,200 configurations. That took 5 days. With 25 candidate features it would take over 7,000 years. Fig. 2a: x is packet depth, y is F1, three of the 64 feature sets shown as F-A, F-B, F-C. F-A is best within the first 10 packets, then the ranking flips; F-B and F-C improve with depth while F-A gets worse. Fig. 2b: x is depth, y is execution time. For a given set, cost rises with depth, but waiting 50 packets for the cheap set F-B costs less than waiting 30 for F-C. The bullet about prior work: papers pick 10, 50, or “all” packets with little justification. Now you see why that leaves performance on the table. So what: this is a search problem with an expensive black-box objective, and that is exactly the setting for Bayesian optimization from the hyperparameter-tuning lecture. Also note that you cannot predict a set’s cost from its members: parsing is shared, and a mean needs a sum you already computed. That is the argument for measuring, which comes up twice more. |
| CATO: A Multi-Objective Optimizer Plus a Profiler | The point: CATO Fig. 3. Two components: an optimizer that proposes feature representations and a profiler that builds and measures the real pipeline for each one. Optimizer, left column. Search space is every subset of the candidate features crossed with a connection depth up to N, encoded as one binary indicator per feature plus one depth variable (CATO §3.1). Two objectives, minimize cost and maximize performance, so the output is a Pareto front, not a single point. Why a front? Because the latency or throughput budget is rarely known in advance and changes with traffic; a front lets the operator move the operating point later without re-running the search (CATO §3.1). Two tailoring tricks: drop features with zero mutual information, and inject priors, MI-based over features and a linearly decaying prior over depth, so the search leans toward informative features and few packets (CATO §3.3). Profiler, right column. Conditional compilation in Rust on top of Retina produces a binary containing only the operations the chosen features need, so the measurement matches a hand-written pipeline. It trains a fresh model per candidate and measures latency, zero-loss throughput, or CPU time on real or replayed traffic. The same measurement guides the search and validates the result (CATO §3.4). Implementation facts for questions: HyperMapper with a random-forest surrogate, pi-BO style priors, Beta(1,2) depth prior, three random initial samples, damping 0.4, 67 candidate features in 1,600 lines of Rust (CATO §4). |
| Use Cases and What the Optimizer Found | The point: CATO Table 2 plus the headline result. Three tasks of different shapes, and on the reproducible one the gain is minutes to a tenth of a second. Walk the table. app-class: classification on live campus traffic, decision tree, labels from the TLS server name, seven classes: Netflix, Twitch, Zoom, Teams, Facebook, Twitter, other. iot-class: the UNSW Sivanathan dataset, 28 device types, random forest, offline traces. startup-inf: video startup delay on the YouTube dataset from Bronzino and colleagues, 4,287 connections, a deep neural network, regression (CATO §5.1, Appendix B). Baselines: all features, top-10 by recursive feature elimination, top-10 by mutual information, each at 10, 50, or all packets. CATO runs 50 iterations with a maximum depth of 50. The headline on iot-class: RFE with 10 packets reaches F1 0.970 at 7.9 seconds of latency; CATO finds a different feature set at 3 packets with F1 0.979 and 0.1 seconds (CATO §5.2). Latency reductions: 11 to 79 times versus 10-packet baselines, 817 to 2,000 times versus 50, over 3,600 times versus waiting for the whole connection. On app-class, zero-loss throughput up 1.6 to 3.7 times versus end-of-connection, after exploring 50 of 2 to the 67 times 50 configurations. Why is latency dominated by depth? Because end-to-end latency includes waiting for packets to arrive, and inter-arrival times dwarf compute. That is the same insight ServeFlow builds on. |
| Pareto Fronts in Practice: Latency | The point: CATO Fig. 5a and 5b. Every point on CATO’s front dominates the baselines, and the x-axis is on a log scale, so the wins are orders of magnitude. Fig. 5a, iot-class. X is end-to-end inference latency, log scale; y is F1. Gray points are the 50 candidates CATO explored; the front is the non-dominated subset. The baselines, all-features, RFE-10, MI-10 at depths 10, 50, and all, sit to the right. Ask students to find “all features, all packets”: far right, and not even the most accurate. Latency drops from minutes to under 0.1 seconds (CATO §5.2). Fig. 5b, startup-inf. Y is now RMSE, lower is better. CATO infers video startup delay in under one second, a 2.2 to 2,900 times speedup depending on the baseline, with lower error. This one matters because it shows the framework is not tied to classification or to tree models; the DNN is trained in TensorFlow (CATO §4). The bullet to emphasize: the wins come from waiting for fewer packets, not from a faster model. The model was never the bottleneck. So what, back to operating points: pick any point on the front and you have a validated pipeline. The operator does not have to re-run anything to move from “most accurate” to “fastest acceptable.” That is the practical advantage of a front over a single optimized configuration. |
| Pareto Fronts in Practice: Live Traffic | The point: CATO Fig. 5c and 5d, on live campus traffic. This is where the operator’s real question, “can I run this on the box I have,” gets answered. Fig. 5c, app-class latency. Two baselines edge CATO on F1: MI-10 at 0.963 and RFE-50 at 0.962 versus CATO’s 0.960. But CATO’s point sits at 0.54 seconds, 2.6 times and 19 times faster respectively (CATO §5.2). A difference of 0.003 in F1 for a 19-fold latency cut is the kind of trade an operator makes without thinking. Fig. 5d, zero-loss throughput on a single core. How it is measured: start at the full traffic rate and lower the NIC’s flow-sampling rate until no packet is dropped for 30 seconds, repeat, average (CATO Appendix D). Single core on purpose, so no configuration saturates the campus ingress and the methods can be told apart; Retina scales per core. The result to read aloud: trading F1 from 0.96 to 0.93 buys 37 percent more throughput. Versus end-of-connection baselines, 1.6 to 3.7 times the throughput; versus 50-packet, 1.3 to 2.7 times. So what: that 37 percent for 0.03 is an operating-point decision, and it is exactly why a front is the right output. Mention the ethics footnote briefly: only aggregate flow statistics and the TLS server name were collected, no payloads, no client IPs, on a tap with a copy of traffic, in partnership with campus networking and security (CATO Appendix H). |
| How Good Is the Search? | The point: CATO Fig. 7 and 8. The optimizer gets within 2 percent of the true front while sampling under 1.6 percent of the space, and the priors are worth a 2.76 times speedup over plain BO. The trick that makes this evaluable: the six-feature mini candidate set (duration, source load, source packet count, source bytes sum, source bytes mean, source inter-arrival mean) is small enough to measure exhaustively, all 3,200 configurations, so there is a ground-truth front (CATO §5.3). Fig. 7: estimated fronts after 50 iterations for CATO, simulated annealing, random search, and iterate-all-features, against the true front. The metric is the hypervolume indicator, area under the front normalized so both objectives count equally. CATO 0.98, annealing 0.88, random 0.86, iterate-all 0.77. Be fair: CATO is not dominant everywhere; around F1 of 0.3 the others do fine. But restrict to F1 at least 0.8, the region an operator actually uses, and it is 0.95 versus 0.39, 0.39, and 0, no solutions found. Fig. 8: convergence over 1,500 iterations, mean and standard error of 20 runs. CATO passes 0.99 HVI at 87 iterations; CATO without preprocessing and priors at 240, the 2.76 times; annealing at 1,295; random at 1,469. So what: the priors are where the domain knowledge goes, MI over features and “fewer packets is cheaper” over depth, and the user supplies neither by hand. |
| Why Measure Instead of Estimate the Cost? | The point: CATO Fig. 9, the profiler ablation. Every cheap proxy for cost lands farther from the true front, and none of them can validate the operating point you pick. Setup: keep the optimizer, swap the profiler’s measurements for heuristics, run 50 iterations on the six-feature set, then measure the true cost and performance of every sampled point afterward to compare fronts (CATO §5.4). Four variants. Naive cost: sum of each feature’s cost in isolation. It ignores shared parsing, so it overestimates. Model-inference cost: time only the model, ignoring capture and extraction; it underestimates. Packet depth as cost: no deployable number at all. Naive performance: sum of each feature’s mutual information; ignores interactions. CATO with real measurements comes closest to the true front. The subtle argument in the box: even if a heuristic found the same front, it could not tell you the actual throughput of the point you chose, so you would still have to build and measure it. CATO folds that validation into the search. The failure mode the paper names is deploying an unrealizable model, one that overestimates throughput or underestimates latency, and then drops packets in production. Wall-clock cost, from appendix Table 5: 9.5 hours for app-class with 67 features and throughput as the cost metric, 2 hours for the six-feature IoT case, and most of it is the profiler measuring cost, 546.7 seconds per iteration for live throughput. Ask: “When is a proxy good enough?” When you are pruning, not when you are deploying. |
| Packet Depth: How Long to Wait | The point: CATO Table 3. “How many packets” is a learnable quantity, it is usually small, and you must give the optimizer a ceiling. Walk the table, iot-class with all 67 candidate features. Left block is the Pareto-optimal solution with the highest F1 for each maximum depth N; right block is the one with the lowest execution time. Cap N at 3: best F1 0.959 at 3 packets. Cap at 5: 0.983 at 4. Cap at 10: 0.994 at 7 packets, 2.04 microseconds. Cap at 25, 50, or 100: CATO still picks 7 to 10 packets, F1 around 0.99. So a generous ceiling does not hurt, the optimizer finds the cheap depth anyway (CATO §5.5). Bottom row, N unbounded: the best-F1 solution wanders out to 42,000 packets at 28.2 microseconds, and even the “lowest time” solution is at 52,000 packets. With no ceiling the depth variable can take any value up to the longest flow in the training set, and the search fails to converge to a cheap point. Right block also shows the floor: one packet gets F1 0.31 to 0.52 for a quarter microsecond. So the front spans from “one packet, coin-flip accuracy” to “seven packets, 0.99.” So what: bound the search space. That is a general lesson for any BO you run: the prior over depth helps, but a hard cap is what keeps the optimizer honest. Compare to the 10-or-50 folklore from the motivation slide: the right answer here was 7. |
| Sensitivity and Wall-Clock Cost | The point: CATO Fig. 10 and appendix Table 5. Skip live unless someone asks about hyperparameters, but the wall-clock table is worth one sentence. Fig. 10a: the damping coefficient delta on the feature priors, from 0 to 1, HVI after 50 iterations on the six-feature set. Delta of 1 means uniform priors, and that is worst at 0.93. Less damping converges faster early. Delta of 0.4 is best at 50 iterations, and 0 also does well (CATO §5.5). Fig. 10b: number of BO initialization samples. Little sensitivity; one point is empirically best, but the authors chose three by convention. Table 5: app-class with 67 features and zero-loss throughput as cost takes 9.5 hours total; iot-class with 6 features and processing time takes 2 hours. Per iteration, the BO sample itself is 55.5 seconds on the big space and 1.4 on the small one, pipeline generation about 50 seconds, measuring performance about 30, and measuring cost 546.7 seconds for live throughput because you have to run against real traffic long enough to see zero loss. So what: the optimizer is cheap; the expense is measuring each candidate pipeline, which is exactly the validation you want. Nine and a half hours is a one-time cost per model version, but note that it recurs whenever a retrain changes the feature set. Put that on the cost tally next to LEAF’s retrains: maintenance is retrain cost plus re-profiling cost. |
| CATO vs. Traffic Refinery | The point: CATO Fig. 6. Be fair to Traffic Refinery, which established cost-aware feature engineering for network ML; the difference is systematic coverage of the space. The figure: F1 versus pipeline execution time on iot-class. Red points are Traffic Refinery’s feature classes, packet counters PC, packet timing PT, TCP counters TC, at depths 10, 50, and all. The authors simulated Traffic Refinery by aggregating those classes from CATO’s own feature set and measuring with CATO’s profiler, since execution time is the cost metric both share; Traffic Refinery also profiles state and storage, which CATO does not (CATO §5.2, Appendix F). Point at PC-10: packet counters at 10 packets happen to be an excellent cheap point. CATO finds a similar-F1 alternative at a modest 344 nanoseconds higher execution time, so it does not strictly beat it. Everywhere else CATO clusters closer to the front. Adding timing and TCP state raises F1 but costs much more, and finding that combination by hand is trial and error because it is not clear which classes work at which depth. Right column, what changed: manual exploration over coarse classes with per-feature cost profiles, versus automated search over individual features and depth with end-to-end cost. So what: this is a nice example of research lineage from the same group. Traffic Refinery asked the right question and gave the operator a dial; CATO turns the dial automatically and shows the whole front. |
| AC-DC: The Environment Drifts Too | The point: AC-DC Table 1 and Fig. 1. LEAF handled data changing under a fixed model; AC-DC handles the system changing under a fixed accuracy goal. Both are maintenance, both answer with adaptation. Walk Table 1. Flow-statistics classifier, a Gaussian mixture model from Bernaille and colleagues on the sizes and inter-arrival times of the first four packets: F1 0.391, time-to-decision 0.023 seconds, 0.503 GB. Packet-capture classifier, the Deep Packet 1D CNN on the first 1,500 bytes of the first three packets: F1 0.977, 19.689 seconds, 82.463 GB. That is roughly 855 times slower and 163 times more memory (AC-DC §2.2). Time-to-decision, TTD, is preprocessing plus inference; Fig. 1 shows the span. The bottleneck is preprocessing, 76.3 percent of TTD for the GMM and 74.7 percent for the CNN; model execution never exceeds 30 percent. Right column: what drifts here is the operating environment. Traffic rate and free memory change through the day. A classifier that fits at 500 flows per second may exhaust memory at 7,000. Ask: “Why does the memory bill grow with rate?” Because when TTD exceeds the time to fill a batch, instances pile up and run concurrently. AC-DC’s idea: a pool of classifiers with different feature requirements and a scheduler that switches. One caveat to state: the “ensemble” is not bagging of algorithms. Every member is the same LightGBM on a different subset of nPrint header fields, so members differ in cost, not inductive bias (AC-DC §3). |
| AC-DC: A Pool of Classifiers and an Adaptive Scheduler | The point: AC-DC Fig. 2 and §3. Three stages: build a pool offline, measure it offline, schedule online. Stage one. Start from 37 header fields from nPrint, drop payload, IPs, and ports, keep 32 from the first three packets. Exhaustive search over feature subsets is the sum over k of 33 choose k, impractical. The heuristic rests on two measured correlations: permutation importance correlates with marginal F1 gain at 0.753, and the total number of bits in a subset correlates with TTD and memory above 0.99. So rank features by importance per bit, and for each size from 1 to 9 take the 10 best combinations: 90 classifiers (AC-DC §3.2, §5.2). Sizes above 9 were pruned because features ranked 10th and below add cost without accuracy. Stage two. For every classifier and batch size, the number of flows collected before one inference call, measure F1, TTD, and per-instance memory. Concurrent instances equal TTD divided by the time to fill a batch, rounded up, and total memory is that count times per-instance memory; formula in the appendix. Memory is measured by shrinking a Linux cgroup limit by binary search until the classifier misbehaves. Stage three. Inputs: current rate and free memory. Drop combinations that would exceed memory, then pick the highest F1-to-TTD ratio, optionally above a minimum performance requirement, MPR. Re-evaluate as conditions change (AC-DC Algorithm 1). Cold-call: “Why not always pick the most accurate classifier?” At 15,000 flows per second its TTD means dozens of concurrent instances, and memory explodes. |
| AC-DC Results: Throughput and Accuracy | The point: AC-DC Fig. 3 and Table 3. With enough memory, AC-DC keeps up with the input rate like the cheap classifier while landing much closer to the expensive one on accuracy. Fig. 3: x is traffic rate from 100 to 15,000 flows per second, y is classification throughput, defined as rate divided by TTD. The GMM is the upper bound since it keeps up everywhere. AC-DC, with and without an MPR of 0.85, matches that upper bound. The packet-capture LightGBM and CNN fall behind; at 15,000 flows per second AC-DC beats them by 176.95 and 48.98 times (AC-DC §5.2). How? With unlimited memory the scheduler picks the best F1-to-TTD classifier at batch size 1 and spawns as many instances as the rate requires. Table 3: average F1 0.855 for AC-DC with MPR 0.85, versus 0.394 for the GMM, 0.944 for packet-capture LightGBM, 0.975 for the CNN. That is 117 percent above flow statistics and 0.089 and 0.12 below the packet-capture models. Without an MPR AC-DC still reaches 0.80, 103 percent above the GMM. Average TTD 0.235 seconds versus 96 and 30 seconds for the packet-capture models. So what: the accuracy gap is the price of using a handful of header fields, and the MPR is the knob that lets the operator set that price. Ask: “What is the MPR really?” A floor on the operating point, exactly the constraint CATO chose not to hard-code. |
| AC-DC: The Dataset Behind the Results | The point: AC-DC Table 2. Know the benchmark before trusting the numbers, and notice that it quietly contains a drift story of its own. Walk the table. Video streaming, 9,465 flows: Netflix 4,104, YouTube 2,702, Amazon 1,509, Twitch 1,150, collected June 2018 from the Bronzino QoE work. Video conferencing, 6,511 flows: Teams 3,886, Meet 1,313, Zoom 1,312, collected May 2020, the MacMillan lockdown measurements. Social media, 3,610 flows: Facebook 1,477, Twitter 1,260, Instagram 873, collected February 2022. Ten classes, all TLS-encrypted, so only headers, sizes, and timing are usable; roughly 20,000 flows with a 50/50 train-test split (AC-DC §2.2, §5.1). Two things to say. First, class imbalance: Netflix has nearly five times Instagram’s flows, which is why the results table reports average, maximum, and minimum F1, not just one number. Second, the collection dates: three snapshots four years apart. The paper trains and tests within the same pooled collection, so the results do not measure drift across years, but the dataset would let you. So what: this is a curated, cleaned dataset, DNS-resolved to service IPs and split by five-tuple. Live campus traffic, as in CATO, is messier. Ask: “What would happen to the 2018 streaming flows if you tested on 2022 traffic?” Probably exactly what LEAF measured. |
| AC-DC Results: Memory and the Balance Point | The point: AC-DC Fig. 4 and Fig. 5. Same shape as CATO and LEAF: the operator picks a point, here via the MPR, and the system holds it as conditions move. Fig. 4: minimum memory needed to sustain full throughput, rates from 500 to 7,000 flows per second. Without an MPR, AC-DC needs less memory than both the GMM and the packet-capture classifiers across the range, because the pool was built to minimize bits per unit of importance. With MPR 0.85 it rises slightly above the GMM starting around 3,500 flows per second, since higher accuracy means more feature bits per instance. Even so, at 7,000 flows per second it needs 118.8 and 126 times less memory than the nPrintML-LightGBM and Deep Packet CNN classifiers (AC-DC §5.3). Remind them how memory is measured: shrink a cgroup limit by binary search until the classifier fails or its TTD grows. Fig. 5, the summary picture: F1 versus TTD on a log axis, at 100 percent throughput and minimum memory, across rates. Three clusters. Packet-capture models top-right, accurate but slow; GMM bottom-left, fast but poor; AC-DC in between, two orders of magnitude left of the packet-capture models and two to three times the GMM’s accuracy, and the MPR moves it up. Ask: “Why are the packet-capture clusters wide horizontally?” Their TTD grows with rate because batch equals rate. AC-DC’s does not, because the scheduler changes batch size and classifier as rate changes. |
| AC-DC Adapts as Conditions Change | The point: AC-DC Fig. 6 and Fig. 7, the “is the controller doing what we said” plots. Fig. 6: fix the rate at 15,000 flows per second and step memory up. As memory becomes available, AC-DC lowers the batch size, which shortens TTD but spawns more concurrent instances, so memory use climbs to meet what is offered. With an MPR the chosen batches are larger, because the higher-accuracy classifiers cost more memory per instance and the scheduler compensates by running fewer of them (AC-DC §5.4). Fig. 7: fix memory at 10 GB and raise the rate from 100 to 15,000 flows per second. AC-DC raises the batch size as rate grows, capping the number of concurrent instances so it stays inside the memory limit. Again larger batches with an MPR. Throughput stays at 100 percent of the input rate in both experiments. The bullet to say out loud: the objective is a ratio, so the scheduler spends resources on shorter decisions when it can and on fewer instances when it must. Tie back to the maintenance loop: this is monitoring and adaptation on a timescale of seconds, versus LEAF’s weeks to months. The three ingredients are the same: a signal to watch, a pool of alternatives measured in advance, and a rule for switching. Limitations to mention, from AC-DC §6: header-only features, one model family, and no treatment of data drift itself, which the authors leave to future work. So AC-DC and LEAF are complementary, not competing. |
| Most Flows Are Easy; Waiting Is the Cost | The point: ServeFlow Tables 1 and 2. The same insight as CATO, waiting dominates, but now turned into a serving architecture rather than a feature search. Table 1, left: weighted F1 versus packets for three tasks and three models. Service recognition with LightGBM: 0.921 on one packet, 0.967 on five, 0.973 on ten. Device identification: 0.913, 0.918, 0.919, almost flat. Features are nPrint bit representations of raw headers, so recall the nPrint lecture: one packet’s header bits already carry a lot. Table 2, right: median feature computation and inference time in milliseconds. Inference takes 0.05 to 7 milliseconds depending on model and depth. Contrast with flow-collection time from the paper’s Fig. 3: waiting for the second packet takes tens of milliseconds at the median and up to 10 to the 3 through 10 to the 6 milliseconds at the tail (ServeFlow §2.2). The abstract’s phrasing: waiting is six to eight orders of magnitude longer than inference. Second insight: for the same task, inference time across models differs by 1.8 to 141.3 times; the extreme is a decision tree on one packet versus a CNN on ten. And the most accurate one-packet model, LightGBM, is not the slowest. So what: for most flows, nine extra packets buy 0.05 of F1 at a huge latency cost. Spend the wait only on the flows that need it. Ask: “Which flows are easy?” Ones whose first packet has distinctive TCP options. |
| ServeFlow: Cheap Model First, Expensive Model Only When Needed | The point: ServeFlow Fig. 4 and Fig. 2. A fast-slow cascade where the assignment rule is the fast model’s own uncertainty. Fig. 4, walk the arrows. Every flow’s first packet becomes an nPrint vector and goes to Queue 1 for the fastest model, a decision tree on packet one. Uncertain predictions go to the fast model, LightGBM on the same packet. Still-uncertain flows go to Queue 3, while features from packets two through N accumulate in Queue 2; the slow model, LightGBM on ten packets, joins them by flow ID. Arrow width is the share of flows, narrowing at each stage; unmatched entries time out (ServeFlow §3.1, §4.1). The tiers come from the F1-versus-latency Pareto front over models and depths. Fig. 2: F1 versus normalized latency as more flows are escalated. ServeFlow tracks the oracle, which knows which predictions are wrong, far better than random escalation; the star is the chosen operating point. Surprise from §3.2: even the oracle should not escalate everything, since the slow model sometimes breaks correct fast predictions; ideal portions are about 17 and 7 percent. Numbers: 76.3 percent of flows answered in under 16 milliseconds, 40.5 times lower median latency, F1 0.964 versus 0.973 for best-effort, over 48.5 thousand new flows per second on 16 cores (ServeFlow §5). Now connect the four papers: LEAF is accuracy versus retrains, CATO accuracy versus serving cost, AC-DC a scheduler holding an accuracy floor, ServeFlow accuracy versus latency per flow. Maintenance is choosing and re-choosing points on those curves. |
| A Maintenance Playbook | The point: turn the four papers into a checklist a student could apply to their own project on Monday. What to monitor, left column. Input distributions: per-feature KS or population-stability tests against the training window; catches data drift the day it happens, no labels needed. Prediction distributions: a classifier that suddenly says one class 90 percent of the time is telling you something. Delayed labels and errors: the NRMSE stream once ground truth arrives, the only direct signal for concept drift; this is LEAF’s detector. Serving metrics: latency, drop rate, zero-loss throughput, because a model that falls behind is wrong too; this is CATO’s point. The asymmetry: input drift is cheap and label-free, but only error proves the model wrong, so you need both. When to retrain: on a trigger, not a calendar, and with data chosen by where the error is. Right column. Keep the Pareto front, not one model. Keep a pool and a switching rule for load and memory, with every member measured in advance. Re-measure cost after retraining, because a new feature set changes the pipeline. Bound the search: cap packet depth, cap retrains per month. Humans in the loop: explanations exist so an operator can confirm a cause; keep the old model and roll back; match mitigation to model family, since resampling helped trees and LSTMs and hurt k-NN. Make it concrete: for your project, what would you monitor, what is the label delay, what does a retrain cost? Most have no monitoring; that is the gap. |
| Summary (1/2) | The point: begin the close. The pipeline from lecture one was measure, represent, model, deploy; today added monitor, diagnose, retrain, and the budgets that constrain them. Read the three bullets with their numbers. Deployed models decay: in four years of cellular KPIs, every model family drifted, regardless of training size or period, through sudden change like COVID, gradual change like the 2021 rise, and recurring change like the weekly cycle. Periodic retraining is not a solution: it is either costly, 169 retrains for weekly, or harmful, 47.79 percent worse on call-drop rate, because it ignores when, where, and why drift occurs. LEAF: detect from the error stream with KSWIN, explain with Local Error Approximation as LEAplot and LEAgram, mitigate by forgetting and over-sampling; best or near-best across models and KPIs with 10 to 77 percent fewer retrains, and never worse than static. Pause on “never worse than static.” That is the property that lets an operator automate retraining without a human checking every result. None of the baselines have it. So what: the first half of the lecture was about the data moving under the model. The second half, next slide, is about the system moving under the model, and about what accuracy costs to serve. |
| Summary (2/2) | The point: finish the close with the systems half, then the one sentence to leave them with. Accuracy has a serving cost: CATO searches features and packet depth jointly with multi-objective Bayesian optimization and a real profiler, and finds pipelines with up to 3,600 times lower latency and 3.7 times higher throughput at equal or better accuracy. The environment drifts too: AC-DC keeps a pool of classifiers with different feature costs and a scheduler that switches classifier and batch size as traffic rate and memory change; F1 117 percent above flow statistics, throughput up to 177 times above packet-capture models, 118.8 times less memory. ServeFlow: cheap model first, expensive model only for uncertain flows; 76.3 percent of flows answered in under 16 milliseconds. Last bullet: maintenance is choosing points on curves, and re-choosing them as the network changes. LEAF’s curve is accuracy versus retrains, CATO’s is accuracy versus serving cost, ServeFlow’s is accuracy versus latency, and AC-DC is the controller that holds a point as conditions move. Invite questions about what did not make it into the papers, since all four are from this group: which thresholds were hand-tuned, what the operator actually did with the LEAgrams, what happened when the pipeline moved to another network. Then thank them; this is the last lecture. |
| Normalized RMSE and Its Drift | The point: the two definitions every LEAF number depends on, for reference rather than for reciting. Top equation: NRMSE for model M on date t and target KPI k. Numerator is the usual RMSE over the N-sub-t samples on that date, prediction minus truth, squared, averaged, square-rooted. Denominator is the range of the KPI on that date, max minus min. Say why: call-drop rate lives under 1, downlink volume lives above 300,000, and without normalization you could not compare drift across KPIs or put them on one figure (LEAF §2.3). The definition is given in words in the published paper and as a commented-out equation in the source. Under 0.1 is the paper’s “good” threshold, citing the standard regression guidance. Bottom equation, Equation 1 in the paper: delta average NRMSE compares a mitigated model M-one against the static model M-zero, as the percentage change in the time-averaged NRMSE. Negative means the mitigation helped. Every table in the LEAF section reports this. Two things worth saying. First, the average over time hides the tail, which is why the paper also reports the 95th percentile. Second, the paper checked R-squared, MAE, MAPE, explained variance, Pearson correlation, MSE, and median absolute error and reports the same phenomena for all of them (LEAF §2.3 footnote), so the story is not an artifact of the metric. Ask: “What would range normalization do on a day with one outlier eNodeB?” Shrink everyone’s error. That is a known weakness of max-min scaling. |
| The Kolmogorov-Smirnov Test Behind KSWIN | The point: the statistic behind LEAF’s detector, and why nonparametric matters for error streams. KSWIN keeps a window of the most recent n NRMSE values and compares the newest r of them against the rest. Build the empirical CDF of each side; the KS statistic D is the largest vertical gap between the two CDFs. Drift is flagged when D exceeds the critical value for the chosen significance level, and the threshold on the slide is the standard two-sample approximation: square root of minus one-half log alpha times n plus r over n times r. Be honest about sourcing: LEAF describes KSWIN in Appendix B in prose and cites Raab and colleagues 2020 and Togbe and colleagues 2021; the formula on the slide is the textbook KS approximation, not copied from the paper. Why this test. NRMSE series are heavy-tailed and bursty, so a test that assumes Gaussian errors would fire constantly or never. KS makes no distributional assumption; it asks whether the recent errors look drawn from the same distribution as the older ones. The paper tried ADWIN, DDM, HDDM, EDDM, and Page-Hinkley too and found KSWIN best on these series (LEAF §4.1). Knobs: window size n sets how much history you compare against, r sets how quickly a change becomes visible, alpha sets the false-alarm rate. Ask: “What happens to the false-alarm rate if you run this daily on six KPIs?” Multiple-testing inflation, which is why the alarms are cross-checked against known events. |
| CATO’s Feature Priors | The point: the one formula that encodes CATO’s domain knowledge, and the reason the optimizer converges 2.76 times faster than plain BO. Top equation: the prior probability that feature f belongs to a Pareto-optimal representation is one minus delta, times the feature’s mutual information with the target divided by the maximum MI over all candidates, plus delta over two. Walk the extremes: delta of 0 gives a prior proportional to normalized MI, so the most informative feature is always included; delta of 1 gives a uniform prior of one-half for every feature. CATO uses 0.4, tuned in Fig. 10a (CATO §3.3, §5.5). The damping exists precisely so the top-MI feature is not forced into every candidate. The depth prior: a probability mass function that decays linearly with connection depth, a Beta with alpha 1 and beta 2 in the implementation, so the optimizer leans toward shallow depths without being told the optimal one (CATO §4). The search space, for completeness: the power set of the candidate features crossed with depth, so the feature count plus one dimensions, one binary indicator per feature plus the depth variable. Objective: minimize cost and negative performance jointly; output is the set of non-dominated points. Say the important caveat from the paper: despite the word “prior,” the user supplies no prior knowledge. MI is computed from the labeled data and the depth prior is a fixed shape. That is what makes CATO usable by someone who is not an expert in the traffic. |
| AC-DC’s Memory Model | The point: the two equations that let AC-DC’s scheduler predict, before launching anything, whether a classifier will fit in memory at the current rate. First equation: the number of concurrently running instances is TTD at batch size B, divided by the time to fill one batch, which is B over the rate R, rounded up. In words: if a classifier takes longer to finish a batch than it takes for the next batch to arrive, batches pile up and instances overlap. Second equation: total memory is that instance count times the measured per-instance memory (AC-DC §3.3, Equations 1 and 2). Walk the worked example from the paper: 5 GB available, 1,500 flows per second, TTD 1.5 seconds, batch 500 flows, 1.5 GB per instance. Time to fill a batch is 500 over 1,500, one third of a second; 1.5 divided by one third is 4.5 instances. The paper reports 6.75 GB, which is 4.5 times 1.5 without applying the ceiling; with the ceiling it is 5 instances and 7.5 GB. Either way it exceeds 5 GB, so the scheduler filters that combination out. The lever this exposes: raising the batch size lowers the instance count at the cost of a longer wait per decision. That is exactly what Figures 6 and 7 showed the scheduler doing. So what: a capacity model, the back-of-envelope calculation you should be able to do for any serving system. Ask: “What does it ignore?” Feature-extraction time, which AC-DC excludes from TTD, and CPU contention among instances. |
LEAF: Navigating Concept Drift in Cellular Networks. Liu, Bronzino, Schmitt, Bhagoji, Feamster, Garcia Crespo, Coyle, Ward. https://arxiv.org/abs/2109.03011
CATO: End-to-End Optimization of ML-Based Traffic Analysis Pipelines. Wan, Liu, Bronzino, Feamster, Durumeric. https://arxiv.org/abs/2402.06099
AC-DC: Adaptive Ensemble Classification for Network Traffic Identification. Jiang, Liu, Naama, Bronzino, Schmitt, Feamster. https://arxiv.org/abs/2302.11718
ServeFlow: A Fast-Slow Model Architecture for Network Traffic Analysis. Liu, Shaowang, Wan, Chae, Marques, Krishnan, Feamster. https://arxiv.org/abs/2402.03694