← Personal projects Machine learning Personal project · 2026

Finding incidents the metrics can't see

Anomaly detection over a microservice platform's metrics and distributed traces. Metrics alone caught a third of the labelled failures on the first day I evaluated; adding trace latency caught all 38 across three days — and the write-up is honest about what that number does and doesn't prove.

  • Python
  • pandas
  • Parquet
  • Streamlit
38 / 38
Labelled failures detected
3 / 9
Metrics-only baseline
597k rows · 133 KPIs
One day of metrics
Rolling MAD z-score
Detector

Architecture

  1. 01Raw archivesmetrics + trace logs
  2. 02Time series1-min entity × KPI
  3. 03Anomaliesrolling MAD z-score
  4. 04Incident windowsmerge adjacent signals
  5. 05Hybrid matchmetrics + trace symptoms
  6. 06DashboardStreamlit, evidence view

The problem

A platform of a few dozen containers, databases and services emits a metric for everything. One day of the dataset I used — the public AIOps Challenge 2020 data — is roughly 597,000 metric rows across 58 entities and 133 distinct KPIs, plus distributed-trace logs. Somewhere in that day, engineers injected faults: CPU exhaustion, network delay, packet loss, database connection limits, a database going away.

The task is the one an on-call engineer has at 3 a.m. Something is wrong. Which thing, and since when?

Why a robust z-score

The first detector was an ordinary rolling z-score, and it has a well-known flaw: the anomaly you’re trying to detect inflates the standard deviation you’re measuring it against. A big enough spike hides itself.

So the detector uses the median and the median absolute deviation instead, over a 15-minute rolling window:

z = 0.6745 × (x − median) / MAD

Neither statistic moves much when a handful of points go wild, which is exactly the property you want. It’s a few lines of pandas and it was a clear improvement over the mean-and-variance version on the same series.

Anomalous points on the same entity that sit close together in time are merged into a single incident window, with the contributing KPIs ranked by peak z-score. That ranking is the root-cause hint: not “something is wrong with db_007” but “these three KPIs moved first, and this far.”

The baseline was bad, and that was the useful part

Evaluated against the labelled failures for one day, watching one KPI per entity, the metrics-only detector found three of nine. When it did fire, it fired fast — a median of two minutes after the fault began — but it missed two thirds of them.

Looking at which ones it missed was more useful than the number. They were almost all network faults. And that makes sense once you say it out loud: a container with 300 ms of injected network delay has perfectly normal CPU and memory. Nothing is wrong inside it. The damage shows up in the services that call it.

Adding traces

Distributed traces record how long each call between services took. So I built the same one-minute series from trace data — latency and error rate per service — and ran the same robust detector over them.

The hybrid rule is simple: a failure is detected if either the metrics for the affected entity or the trace latency of the surrounding services goes anomalous during the fault window.

On the first day that took detection from the baseline to 14 of 14. The five network faults that metrics missed entirely were caught by traces, with peak z-scores between 12 and 37. Across the three days in the dataset that have labelled failures in the window, the hybrid detector found 38 of 38.

What that number doesn’t prove

This is the part I’d want an interviewer to ask about, so here it is first.

Recall is not precision. On the day with 14 real failures, the pipeline produced 388 incident windows. A detector that fires 388 times and catches all 14 is not something you could page a human with.

Long windows match everything. One database incident window ran for 354 minutes. Any fault on that entity during those six hours counts as “detected,” including one logged as starting two hours before the fault it was matched to. The evaluation as written rewards windows for being long.

So the accurate summary is: traces carry signal that metrics don’t, and that’s worth building on. Whether this is a good detector is a question the current evaluation can’t answer.

What I’d do differently

Score precision, and cap the window. Count incident windows that overlap no labelled fault, and require that a detection start within a few minutes of the fault rather than merely overlap it. I’d expect the headline number to drop, and the new one would mean something.

Use the topology. The trace data already encodes which service calls which. A network fault on one container should light up its callers and not its siblings; using the call graph to localize the fault would be a real root-cause step rather than a ranked list.

Timestamps bit me once. The metric archive and the failure labels were eight hours apart — one in UTC, one in local time. Until that was aligned, nothing matched and the detector looked broken. It’s the first thing I check now.

Contact

Want the longer version?

Happy to walk through any of this in detail — the parts that broke are usually the interesting bit.