If you're running detection content across more than one Sentinel workspace, you've probably already got an exclusions problem, even if it doesn't feel like one yet. Maybe you're early enough into the rollout that the noise hasn't started. Give it time. The rule logic you want to ship is generic, built to go from one repo out to many customers. What it generates is anything but generic. The noise is bespoke, per customer, every time.
Take software deployment as the obvious example. One tenant's IT team distributes software from an internal UNC share. Another uses SCCM. A third pushes everything through Intune. A rule that fires on a file-association registry write from a UNC path is correct in all three environments. Useful in precisely one of them, and only once you've taught it what normal actually looks like there.
The wrong way to solve this is embedding tenant-specific exclusions straight into the rule body. The better way is picking from four patterns - saved functions, watchlists, and two flavours of CI/CD-injected exceptions - and knowing which one fits which job.
The problem, restated
Here's the hunt query I'll use to walk through this. It's looking for file-association hijack, registry writes to a \shell\open\command key that hand an attacker code execution every time a user opens a file of a given type.
DeviceRegistryEvents
| where ActionType == @"RegistryValueSet"
| where RegistryKey contains @"\SOFTWARE\Classes\"
| where RegistryKey endswith @"\shell\open\command"
| where InitiatingProcessAccountName != @"system"
The technique's real and the query's sound. The trouble starts the moment you run it across a multi-tenant estate. In one tenant you're seeing hundreds of hits a day off a Lenovo System Update rollout. In another it's a wave of TeamViewer Host installs from an internal software share. In a third it's a Visual Studio update spraying VSIX registrations across every developer workstation in the place.
None of that's malicious. All of it needs tuning out, and each one needs tuning out differently. You can't do that with a one-to-many rule deployment, not without introducing blind spots somewhere or excluding on string content that ends up exposing customer data to a query that shouldn't be seeing it.
The naive response is to keep bolting on | where not (...) clauses until the noise stops. Do that for a few months and you've built yourself a two-hundred-line monster of a query, eighty percent of it exclusions, and nobody on the team can tell you which line protects against which false positive or which customer it was even added for. Worse than that, Customer A's exclusions are running against Customer B's data every single time the rule fires. Best case that's wasted compute. Worst case it's a data-segregation problem you'd rather not be explaining to anyone.
Four patterns solve this properly. They're not mutually exclusive - most mature setups end up running a mix - but each one has its own strengths and its own limits.
Saved functions
A saved function is a named, parameterised piece of KQL sitting inside a Log Analytics workspace. Save it once and you can call it from any query in that workspace as if it were a built-in operator. If you've ever used the ASIM parsers, you've been calling saved functions without realising it - _ASim_ProcessEvent is a function, not a table.
For per-tenant exclusions, define a function whose body holds the allowlist logic, then call it from your hunt or analytics rule. The hunt content stays tenant-agnostic. Only the function body changes, workspace by workspace.
In each tenant's workspace, save a function with the alias IsTrustedSoftwareSource and parameters folderPath:string, valueData:string. Here's the body for a customer running two internal software shares:
let trustedUncPaths = dynamic([
@"\\srv-util01\software installs\",
@"\\srv-util02\data\"
]);
let isTrustedSource =
folderPath has_any (trustedUncPaths)
or folderPath startswith @"c:\program files\"
or folderPath startswith @"c:\program files (x86)\";
let isTrustedTarget =
valueData matches regex @'^"?[A-Z]:\\Program Files( \(x86\))?\\';
isTrustedSource and isTrustedTarget
And the hunt becomes:
DeviceRegistryEvents
| where ActionType == @"RegistryValueSet"
| where RegistryKey contains @"\SOFTWARE\Classes\"
| where RegistryKey endswith @"\shell\open\command"
| where InitiatingProcessAccountName != @"system"
| where not (IsTrustedSoftwareSource(InitiatingProcessFolderPath, RegistryValueData))
The hunt now holds zero tenant-specific data. Every tenant's workspace gets the exact same hunt, each one just calls its own version of the function. Tenant adds a new software share? Update one function in one workspace and every hunt that calls it picks up the change next time it runs.
- The query stays readable.
- You can test the exclusion logic on its own, independently of whatever rule's calling it.
- Updates take effect immediately. No deployment pipeline needed.
- Functions can hold proper logic - regex matching, conditional branches, multi-parameter checks - in a way watchlists simply can't.
- You can't share a function natively across workspaces, so in a Lighthouse setup you're defining it separately in every single one.
- Cross-workspace calls via
workspace("xxx").FunctionNameexist, but they're fiddly, and they don't always behave inside analytics rule contexts. - Because the function lives inside the workspace, version control is somebody else's problem. There's no native git-style history unless you've built one yourself through infrastructure-as-code.
Use functions when the exclusion logic's complex enough to deserve its own home, when it's reused across a good number of rules, or when it needs to take parameters and hand back a computed result rather than just matching against a list.
Watchlists
A watchlist is a Sentinel-native feature: a small reference table you upload as a CSV, edit through the portal, and join against from KQL via _GetWatchlist. They're built for exactly this kind of "list of known-good things" job, and they're genuinely handy if you're a detection engineer working inside a SOC with just the one tenant to worry about.
Same example again, this time as a watchlist called TrustedSoftwarePaths with a single SearchKey column:
let trusted = toscalar(
_GetWatchlist("TrustedSoftwarePaths")
| summarize make_list(SearchKey)
);
DeviceRegistryEvents
| where ActionType == @"RegistryValueSet"
| where RegistryKey contains @"\SOFTWARE\Classes\"
| where RegistryKey endswith @"\shell\open\command"
| where InitiatingProcessAccountName != @"system"
| where not (InitiatingProcessFolderPath has_any (trusted))
The strengths here are mostly about the humans involved, not the tech. The weaknesses are what make them a poor fit for multi-tenant work, and in some setups they just won't cut it:
- There's a CSV editor in the portal, so non-engineers can maintain the list without writing KQL.
- They suit long, flat lists that change frequently - IP allowlists, known service accounts, approved installer hashes.
- They support multi-column lookups, so each entry can carry metadata (owner, justification, expiry) that the rule output can surface.
- Watchlists are size-limited, and lookups against large ones get slow.
- They're tenant-scoped just like functions, so multi-tenant deployments still need a synchronisation story.
- They're built for "is this thing in the list" matching, full stop. Bring in regex, conditional branches, or relationships between fields and you're straight back in function territory - and multi-field matching just brings back the original ungainly-query problem anyway.
Use watchlists when what you've got is a list of values rather than logic, when it changes often enough that someone outside engineering needs to keep it updated, or when each entry needs its own metadata attached.
Patterns 3 and 4 - CI/CD-injected exception blocks
The third and fourth patterns are doing the same job at heart: splicing a customer-specific KQL fragment into a base rule at deployment time. Both sit outside the product itself, which means you're building this yourself into your deployment pipeline. Don't go looking for either of these in Microsoft's documentation. You won't find much.
They differ enough in the implementation that I'd treat them as two genuine alternatives rather than two flavours of the same idea. Which one you end up with comes down to a handful of distinct engineering choices, or, just as often, whatever you've inherited when you land in a role and get stuck in.
Marker-fenced exception blocks
This approach uses KQL line comments as machine-readable bookends:
/// Start of exception
/// End of exception
Those are completely inert at query time - KQL just ignores them - but a deployment script can go and find them. The script has two jobs: strip every region between the markers, looping until nothing's left so prior exceptions get properly cleared, then append fresh markers wrapping the fragment pulled straight from the customer's exception file, verbatim.
Invisible to KQL, fully visible to your pipeline. Find the markers, wipe out whatever was sat there before, drop the new customer-specific exception in. Clean slate on every deployment, so you're not building up exclusion zombies.
That's genuinely the whole mechanism. No parsing, no AST, no template substitution, nothing wrapping the fragment in | where not(...) for you. The customer-specific KQL goes in byte-for-byte between the markers, and those markers exist for one reason only: re-running the workflow with a new exception replaces the prior block instead of stacking on top of it.
Everything else is orchestration. Read the base rule JSON out of the repo, load the customer's workspace details from a config file, pull service principal credentials from a secret store, find the existing rule by display name, and PUT the modified rule body back into that customer's analytics rule.
Here's a worked example, the exception file for one customer:
| where not (InitiatingProcessFolderPath has_any (
@"\\srv-util01\software installs\",
@"\\srv-util02\data\"
))
| where not (
InitiatingProcessFileName in~ ("teamviewer_host_setup.exe", "teamviewer_host_setup_x64.exe")
and InitiatingProcessVersionInfoCompanyName =~ "TeamViewer"
)
After deployment, that customer's rule is the base query with the block appended between the markers. A different customer gets a different block, or none at all. Same source rule, tenant-tailored result every time.
- Everything lives in git, so you get version control, PR review, blame history and rollback without having to build any of it yourself.
- The exception is data but the deployment is code, so changes go through the same review process as anything else.
- Because the fragment gets appended literally rather than templated, you keep the full expressiveness of KQL. Joins, lets, regex - anything you can write in a base rule, you can write in an exception.
- The fragment gets concatenated raw onto whatever the base query ended with. If either end doesn't terminate cleanly, you've got invalid KQL.
- There's no syntactic check before the PUT, so a malformed exception only shows itself when Azure rejects the deployment - or worse, accepts it and the rule quietly fails at evaluation time instead.
- Convention's doing all the heavy lifting here. Exception files always begin with
| where ..., but nothing actually enforces that. The repo's treated as source of truth for everything except the exception block, so any manual portal edit gets silently overwritten on the next deployment.
Tuning overlay with a workflow dialog
A different take on the same idea, with a fair few engineering choices going the other way. The most visible difference is how the exception actually gets into the file in the first place.
Instead of leaning on engineers to raise pull requests, this ships a GitHub Actions workflow with a manual dispatch dialog. An engineer - or in principle an analyst with the right permissions - heads to Actions, picks "Apply Tuning and Deploy", and fills in three fields: the customer from a dropdown, the detection ID, and the KQL lines to add.
From there the workflow finds the tuning file for that customer and detection under Tuning/<Customer>/<detection-id>.kql, creates it if it's not there yet, appends the new lines if it is, commits the file back to main so the change lands in version control, merges the tuning into the detection query in memory, and deploys the merged rule using OIDC authentication.
The splice mechanism is similar in spirit to the marker-fenced approach, but uses a single separator line rather than paired markers:
//==========TUNING BELOW==========/
The base query sits above the separator, customer-specific tuning sits below it. On re-deployment the script replaces everything below the line rather than duplicating it, so you get the same clean result run after run.
The differences from Pattern 3 reflect different bets about how exceptions should reach the pipeline:
Pattern 3 assumes a git-native workflow: edit the file, raise a PR, get it reviewed, merge, deploy. Pattern 4 puts a form in front of the Actions UI instead. Same end state, different friction - the dialog makes it accessible to someone who's never opened a YAML file in their life, while the PR gives you more rigorous review by default. Neither's universally better. It comes down to whether your team sits closer to engineering or to operations.
Pattern 3 typically pulls service principal credentials from a secret store at runtime. Pattern 4 uses OIDC federated credentials issued by GitHub and trusted by Azure, with each customer mapped to a named GitHub Environment holding that customer's tenant, client and subscription IDs. No long-lived credentials sat in the repo or in secrets anywhere.
Pattern 3 replaces the whole exception block on each apply, so the file always equals the current set of exceptions. Pattern 4 appends instead, which builds you a chronological history of every suppression added - you can see in the git log exactly when each line arrived - but the file keeps growing and needs the odd tidy-up.
Pattern 4 runs one job per customer with an explicit condition, evaluated at queue time so anything that doesn't match gets skipped entirely. Selecting Customer A can't touch Customer B's workspace or credentials, even if the workflow gets modified badly down the line. That's a sharper form of isolation than one job branching internally off a parameter.
The strengths overlap heavily with Pattern 3 - git-backed, version-controlled, full KQL expressiveness, tenant-scoped - plus a three-minute round trip from "this is firing repeatedly on legitimate behaviour" to "exception merged and deployed", and it's accessible to people who've never raised a PR in their life. OIDC also wipes out a whole class of credential-leak concerns in one go.
So do the weaknesses, plus a few specific to the design:
- Same no-syntactic-check problem as before. Type a malformed line into the dialog and you've got a broken rule, and you won't know until deployment time.
- The dialog bypasses PR review by design - that's rather the point - but it also means a junior analyst can deploy a broken exception and nobody notices until the next on-call rotation picks it up. An automated KQL validation step, or a dry-run toggle that merges but skips the PUT, would go a long way here.
- The append-only model builds up zombie exceptions over time. Suppressions added for a one-off rollout that never happened again just sit there, never removed, tucked away inside each tenant individually.
Picking between them
A useful mental model:
A list of values that changes weekly, edited by humans who don't write KQL, with metadata attached
that's a watchlistA piece of logic reused across many rules that benefits from parameterisation
that's a functionPer-customer-per-rule, version-controlled, reviewed via PR, in a git-native team
that's Pattern 3Per-customer-per-rule, accessible to people who don't raise PRs, with OIDC credential isolation
that's Pattern 4In practice most mature setups run a mix of all four. Watchlists pick up the things that change too often to put through PR review and don't need the full expressiveness of KQL. Functions sit underneath, holding the logic that's genuinely reused across a good number of rules. One of the two pipeline patterns handles the long tail - the per-customer-per-rule tuning that's too specific for functions and too logic-heavy for watchlists.
The pattern you start with matters less than recognising early that you need one.
Embedding tenant-specific exclusions inside the rule body is the trap, and it's an easy one to fall into when all you're trying to do is stop the noise. Make that choice and ship it, though, and every rule update turns into an exercise in not breaking whatever's been layered on top. The cognitive overhead just compounds with every new tenant you bring on.
Pick a pattern - or more honestly, pick a mix - and commit to it before the noise ends up making the decision for you.
Running detection content across multiple tenants?
Detection Engineering Sprints cover the tuning layer as well as the rules - exclusions built as structure, not bolted onto the side of the query.
See the servicesA longer version of this post first appeared on Medium.