A practical framework for working out where your detection quality actually sits, and what to do about it.
If you've ever sat in a post-incident review and heard "our detection didn't fire" - or worse, "it fired but the analyst couldn't tell if it mattered" - you have already met the problem I want to talk about.
The way most SOCs measure detection quality is through counting. How many rules do we have? What percentage of the MITRE ATT&CK matrix do we cover? What's our alert volume trending like this quarter? These numbers get reported to the board and they feel like they mean something. They mostly don't. What they measure is activity, not capability.
What actually matters is whether your detections would survive contact with a real attacker in your specific environment. That comes down to something more fundamental than rule count: how those rules use context, and whether they can adapt to what your organisation looks like.
Here's the framework I built to make that visible.
Every rule in your SOC, regardless of platform, sits at one of three tiers. The tiers are not about query length, joins, or how sophisticated the parsing looks. They're about one question: does the rule get any smarter as your organisation learns more about itself?
Fixed logical conditions on raw data. Same input, same verdict, every time. Cheap to write, reliable to maintain, catches the obvious.
Enriched with maintained information about your environment - identity data, threat intel, watchlists, a CMDB. Entities get treated differently.
Detects deviation from learned normal. Baselines, peer groups, statistical rarity, UEBA. The verdict changes as the model learns.
A T1 rule applies fixed logical conditions to raw data. It looks at one type of event, checks it against criteria usually baked into the rule itself, and produces a verdict. The verdict is the same every time you run it against the same input. It doesn't matter how many attacks your SOC has already seen, how well you know your own users, or what your business does. The rule doesn't care.
A rule flags any successful sign-in from outside the UK by an admin account. It geo-locates the IP, checks the account against an admin group baked into the query, fires if both are true. A perfectly valid rule. Also T1: it treats every admin the same, every non-UK country the same, and it never learns anything from what it has already seen.
Most rules in most SOCs are T1, and that isn't in itself a bad thing. But if T1 is where your coverage ends, you have a specific and predictable set of problems: high noise, low precision, and detections an attacker can defeat with one small adjustment to their playbook.
A T2 rule enriches the raw event with maintained information about your specific environment. It joins to something you keep updated: identity data (who works where, who reports to whom, who has what privileges), a threat intelligence feed, a watchlist you curate, a configuration management database.
The moment you make that join, the rule starts treating entities differently based on what your organisation knows about them. That same non-UK sign-in looks different when the rule knows this admin has been travelling on business for a week, or that this user is a contractor whose contract ends tomorrow, or that this account has been under quiet investigation for a fortnight.
You've moved from "does this activity match a bad pattern" to "does this activity match a bad pattern given what I know about the person doing it".
That's a real leap, and not many SOCs make it consistently - it requires the joined data to be trustworthy and current, which is an operational investment most SOCs underweight.
A T3 rule detects deviation from learned normal behaviour. It isn't just enriched with what you know, it adapts based on what it has seen. Baselines, peer-group comparisons, statistical rarity, machine-learning models, UEBA - all T3 mechanisms.
Run the exact same event through the detection twice. Would you get a different verdict depending on what the detection had learned between the two runs? If yes, it's T3.
T3 is where you catch what atomic and contextual rules can't. The insider going bad slowly. The compromised account being used carefully. The attacker who has done their homework and is trying to look normal.
It also fails in ways T1 and T2 don't. Bad baselines produce bad detections. New employees look anomalous until the model catches up. Deploy a T3 rule before the baseline has matured and you will produce noise nobody can triage - which is why freshly-deployed T3 rules should be treated as provisional until they've been through proper flight testing.
Detection maturity is not about how complex your queries are, how many lines of KQL a rule takes up, or how much parsing it does. It's about whether your detections use organisational context, and whether they learn from what they see.
I've reviewed queries two hundred lines long - dense with parsing, joins and dynamic logic - that were still T1. Sophisticated in construction, unsophisticated in thinking. And I've seen twenty-line T2 rules with one join that produced a genuinely intelligent verdict reflecting how the business actually worked.
If you don't know, that's your first exercise. Sample twenty rules and read them. No join to identity data, no watchlist enrichment, no threat intel correlation, no baselining - it's T1. In most SOCs I've looked at, the answer sits between eighty and ninety-five percent.
T2 requires maintained enrichment data. Who owns it? Is your identity information current, or does it lag HR by weeks? Is threat intel adding signal or noise? Where T2 exists, it's usually a handful of rules built by one engineer who cared enough. That's fragile. It should be structural.
T3 isn't the goal for every rule - some things should be detected atomically because they should never happen at all. Ask where behavioural detection materially improves your position (usually authentication anomalies, exfiltration patterns, insider activity). No T3 anywhere is a coverage gap, and a signal.
Not an audit. Ten rules, chosen at random. Work each one down this tree - the first "no" is your answer.
Does this rule join any maintained enrichment data?
Does that enrichment change how the rule treats different entities?
Does the rule adapt based on what it has learned over time?
Would it produce different verdicts on identical input as time passes?
The distribution you find is your starting position. It will almost certainly be more T1-heavy than you expect. That is not a criticism of your team - it's the natural gravity of how SOCs get built. T1 is easier, cheaper, faster to ship. The question is where you want to be, and what the plan to get there looks like.
I run this framework across live detection catalogues as part of what I do at Detection Engine. Happy to help if working through it against your own rules would be useful. But the framework itself is free to use - I'd rather you tried it, argued with it, and told me where it breaks than took my word for any of it.