Detection Metrics Without Vanity Numbers
Use detection metrics to answer operational and security questions with defined populations, denominators, time windows, and uncertainty instead of celebrating activity counts.
A metric is an argument made with measurements
A metric combines measurements according to a definition so you can reason about a decision. “1,200 alerts” is a count. “The proportion of investigation-ready alerts among all delivered alerts this week” is closer to a metric because it names a relationship, population, and interval.
Start with the question. Are you deciding whether a source is healthy, an analytic is precise enough for paging, a portfolio covers priority behavior, or analysts receive enough context? Each question needs different evidence.
A dashboard cannot repair an undefined concept. Write the numerator, denominator, inclusion rules, time basis, owner, and intended interpretation before choosing a chart.
Denominators make comparisons honest
A denominator states the population against which a count is interpreted. Alert volume per thousand identities, observed source coverage per expected asset, and useful detections per validated scenario answer different questions from their raw numerators.
Keep populations stable or describe the change. If tenant count doubles, a doubled alert count may mean the rate stayed constant. If logging disappears from quiet systems, an apparent precision improvement may be survivorship bias.
Segment where causes differ: source, tenant, platform, analytic version, entity class, or analyst queue. Overall averages can hide one failing population behind another healthy one.
Quality metrics depend on trustworthy labels
Precision-like measures require a definition of a useful or correct alert. Analyst dispositions are tempting labels, but they often mix truth, authorization, scope, and workflow. A closed alert may be benign, unresolved, duplicated, or deprioritized.
Define label categories and audit a sample for consistency. Preserve “unknown” rather than forcing every case into true or false. Where incidents are rare, confidence intervals and small sample sizes matter more than a neat percentage.
Measure evidence completeness and time to a defensible decision alongside disposition. These can reveal improvements even when reliable ground truth is limited.
Operational and security outcomes should not be collapsed
Operational metrics describe whether the service runs: freshness, execution success, latency, cost, delivery, and source health. Security-effectiveness metrics describe whether it supports useful decisions: behavior coverage, evidence quality, investigative value, and contribution to risk reduction.
A fast, reliable analytic can detect the wrong thing. A valuable analytic can be operationally unusable if it arrives after the decision window. Report both dimensions so one cannot conceal the other.
Portfolio measures should connect to prioritized threats and assumptions, not reward inventory size. Ten redundant rules do not create ten times the coverage.
Metrics should change a decision or disappear
Assign every metric an owner, review cadence, data lineage, and response when a threshold changes. If nobody can name an action or interpretation that follows, the measure is probably reporting activity rather than governing a service.
Guard against Goodhart’s law: when a measure becomes a target, people and systems optimize the number rather than the underlying outcome. Alert closure speed can improve by shallow closures; rule count can improve by fragmentation. Pair measures and review examples.
Connect operational thresholds to detection service levels, and retire metrics when their question, population, or decision no longer exists.
Annotate changes that affect comparability: source onboarding, parser revisions, threshold changes, staffing, queue policy, and incident spikes. A trend line without these events invites causal stories that the number cannot support. Review a small set of underlying examples whenever a metric moves materially; measurement should lead you back to evidence, not replace it.
Frequently asked questions
Why are alert counts poor detection metrics?
An alert count mixes environment activity, collection health, analytic behavior, suppression, and routing. Without a denominator and decision context, a rising or falling count cannot tell you whether detection quality improved.