Ask most teams how mature their detection rules are and you'll get a shrug, or a coverage percentage pulled from a framework mapping that says nothing about whether the rules actually work. Coverage tells you a technique is theoretically watched. It doesn't tell you whether the rule watching it will survive contact with a Tuesday.

The maturity framework I use grades rules against how they behave over time, not whether a box is ticked. Three tiers, each one a real jump in engineering effort:

T1 - Fixed-condition logic

A T1 rule fires on a static condition: a known-bad hash, a specific command-line string, a hardcoded threshold. It's reliable in the narrow sense that it does exactly what it was written to do, every time. The problem is everything outside that narrow condition. Rename the file, adjust the argument order, or wait for the threshold to drift and the rule goes quiet without ever technically failing.

Most detection content sold as "coverage" lives here. It's not wrong to have T1 rules - some conditions genuinely are static - but a ruleset that's entirely T1 is a ruleset that ages out the moment an adversary changes one string.

T2 - Enriched and context-aware

A T2 rule pulls in maintained context at evaluation time: asset criticality, user role, known-good baselines, threat intel that's actually refreshed rather than imported once and forgotten. The logic is still largely deterministic, but the inputs are alive, which means the rule's behaviour changes as the environment changes without anyone having to rewrite it.

This is where automation starts to pay off, because the enrichment is doing the work a human would otherwise do by hand at triage - and doing it before the analyst ever sees the alert.

The jump from T1 to T2 is rarely about writing cleverer logic. It's almost always about whether there's a maintained data source to enrich against in the first place.

T3 - Adaptive

A T3 rule detects deviation from a baseline rather than matching a fixed pattern. It doesn't need to know the exact string an adversary will use next, because it isn't looking for a string - it's looking for behaviour that doesn't fit what's normal for that user, host, or service. This is the tier that survives an adversary changing their tooling, because the rule was never keyed to the tooling.

T3 is also the most expensive tier to build and the easiest to build badly. A baseline that's noisy, too short a window, or built on bad historical data produces a rule that's technically adaptive and practically useless. Getting here is worth it for high-value detections; it's overkill for a rule that only ever needs to catch one specific, unchanging thing.

Grading what you've already got

The point of the framework isn't to push every rule to T3. It's to know, rule by rule, which tier you're actually operating at versus which tier you assumed - and to make a deliberate call about which rules are worth the investment to move up. In practice that usually means:

  • A small number of high-value rules pushed to T2 or T3, because the cost of a miss is high enough to justify the build effort.
  • A larger set of T1 rules left exactly as they are, because the condition they watch genuinely is static and enrichment would be effort spent for nothing.
  • A pile of rules nobody has looked at since they were written, which turn out on review to be silently broken, not just immature.

That last category is usually the biggest one, and it's the reason a structured review finds more than a quick read-through does - it's not about opinion, it's about checking each rule against the same tier definitions every time.

Want your ruleset graded against this framework?

The Detection Rule Maturity Review assesses every rule in scope, tiers it, and shows the reasoning behind the grade.

See the services