Alert Triage Automation and False Positive Rates in Co-Managed SOC Models
Automation cuts false positives by half while catching nearly all real threats in production SOCs.

A SOC that catches more threats than it did five years ago is not, by itself, a SOC that's easier to run. The average organization now receives 2,992 alerts in a given period, and 63% of them go unaddressed. That's what happens when detection coverage improves faster than the ability to sort through what it produces. It's what happens when detection coverage improves faster than the ability to sort through what it produces.
The mismatch has a structural cause, and it's not mysterious. Most security stacks now run a double-digit number of tools, each one generating its own alert stream, in its own format, with its own severity scale. A security tool calls something "critical."" A firewall calls a related event "medium." A cloud security tool flags the same underlying activity under a category that doesn't map cleanly to either one. Layer enough of these consoles on top of each other and alert fatigue results by design, not by accident.
Detection rules make the problem worse in a subtler way. Most are built for coverage, not precision. That's a reasonable choice on its own: a rule that's tuned to catch everything suspicious will, by definition, flag a lot of things that turn out to be harmless. The rule is doing what it was designed to do. It's just doing exactly what it was told to do, which is cast a wide net and let someone downstream sort the catch. The question that actually matters is who, or what, is doing that sorting, and how well.
What the research shows about false positive rates
If five people in this industry are asked what the false positive rate is, the honest answer is: it depends. The real range runs from 46% to 83%, and the width of that range says as much about the state of SOC maturity as any single number could.
The floor comes from Microsoft and Omdia's State of the SOC report, which puts the baseline at 46%. Almost half of everything an analyst looks at generates zero security value. That's the floor, in a well-resourced environment, with modern tooling.
The middle of the range comes from Devo's SOC Performance Report, which found false positive rates running up to 53%, alongside a separate finding that 70% of SOCs say they struggle to manage alert volume at all. Those two numbers reinforce each other. A team drowning in noise doesn't have the bandwidth to tune the rules producing that noise, so the noise persists.
The Devo SOC Performance Report found that, at the high end, some environments report false positive rates as high as 80%. And in the most extreme documented cases, the number stops looking like a percentage and starts looking like a rounding error in the other direction. In one documented case, an oil refinery's intrusion detection system generated nearly 27,000 alerts, of which only 76 were legitimate threats, a false positive rate exceeding 99.7%. Run the math on that: an analyst's job on a system tuned that badly isn't hunting threats, it's wading through static hoping to hear a signal underneath it.
How many alerts genuinely need a human analyst, and its implications for triage design
Most estimates put the share of alerts that actually need a human analyst's judgment somewhere between 10% and 30%. Everything else is noise that a properly tuned system should catch before a person ever sees it.
One commonly cited benchmark frames this as an escalation rate target, suggesting SOCs keep the share of alerts escalated to deeper investigation under 20%. Expel's 2026 Annual Threat Report puts a sharper point on it: when 90% or more of investigated alerts close as benign, the true positive rate sits below 10%. That 10% figure is the benchmark leading managed SOC providers are trying to hit. It's the benchmark leading managed SOC providers are trying to hit.
Putting those two data points together makes the design implication obvious. The goal of a well-run triage process is not to review alerts faster. It's to make sure the 70% to 90% that qualify as noise never reach an analyst's queue in the first place. Speed on the wrong workload is still the wrong workload, just processed with more urgency.
What automation delivers in real SOC environments, not vendor claims
Vendor pitch decks are full of impressive numbers. The ones to trust come from production deployments, not lab conditions, and they come paired with a second number that vendors sometimes leave out.
One system deployed in a live SOC cut analyst-facing alerts by 61% over a six-month period. That's a meaningful drop on its own. But the number that actually validates it is the false-negative rate, which held at 1.36% throughout. Those two figures have to travel together, and here's why: a 61% reduction in alert volume paired with a high miss rate would just mean the system got good at hiding real threats instead of catching them. A 1.36% false-negative rate means the reduction came from suppressing genuine noise. That's the test any alert-reduction system has to pass, and it's the test to ask about before believing any vendor's volume-reduction claim.
Multi-agent systems built on large language models show a similar pattern, with accuracy gains that come alongside, not instead of, precision gains. One collaborative system of LLM agents cut its false-positive rate from 24.9% down to 14.2%, a meaningful drop, while simultaneously raising its actionable-decision F1 score from 0.66 to 0.78. That result came from production traces spanning more than ten security scenarios.
Then there's the speed dimension, where automation isn't just better than manual review, it's operating in a category humans can't compete in at all. An automated pipeline classified more than 12,000 SSH connection attempts with a mean time to block of 0.86 seconds. No analyst, however sharp, is reviewing a connection attempt and issuing a block decision in under a second. That kind of response time only exists because a machine is making the call.
How co-managed SOC models divide the work between automation and human analysts
A co-managed SOC is neither fully outsourced nor fully run in-house. It combines an external provider's 24/7 monitoring, its automation infrastructure, and its analyst capacity with an internal team's knowledge of how the business actually works. Neither side does the whole job alone, and that's the point.
The external layer typically handles the volume work: routine monitoring, first-pass triage, executing SOAR playbooks, auto-closing patterns already known to be benign, enriching alerts with threat intelligence, and making the initial call on whether something needs to escalate further.
The internal layer keeps what an outside provider structurally cannot have: knowledge of which assets actually matter to the business, what the organization's risk tolerance looks like, and the judgment needed for high-complexity investigations. Final decisions on significant response actions, the kind that could take down a production system or lock out a business unit, tend to stay in-house for the same reason.
An outside provider with no business context will escalate things that anyone inside the company would immediately wave off as noise, because they don't know that the flagged server gets rebooted ever... An outside provider with no business context will escalate things that anyone inside the company would immediately wave off as noise, because they don't know that the flagged server gets rebooted every Tuesday for routine maintenance. An internal team without automation infrastructure behind it will simply drown, the same way the average SOC drowns in its 2,992 alerts. The division of labor is the product. Get it wrong, and you've just built a slower, more expensive version of either extreme.
How to tell whether a co-managed SOC is reducing false positives or just moving the noise elsewhere
There's one question that cuts through most of the marketing language around co-managed SOC services: is the provider actually suppressing noise before it reaches the internal team, or are they just forwarding escalations that a well-tuned system should have filtered out automatically? Everything else is detail.
Escalation rate is the first place to look. If more than 20% of alerts are getting escalated for human review, the automation layer is not pulling its weight, and the escalation rate is what shows it. Best-in-class targets are below that 20% ceiling, consistent with estimates that only 10% to 30% of alerts genuinely need a human analyst's judgment.
False-negative visibility is the second signal, and it's the one that separates real performance from marketing copy. A provider should be able to state their miss rate, not just their alert-reduction percentage. A vendor who can quote a substantial drop in alert volume but goes quiet when asked about false negatives hasn't proven they're suppressing noise. They've proven they can hide it, which is a different and much worse thing.
Tuning cadence is the third signal worth pressing on. Ask how often detection rules get updated based on analyst feedback, and ask for specifics, not a general assurance that tuning happens. A provider running the same rule set for months without adjustment is, functionally, running the same rules that drove false positive rates into the 46% to 83% range across the industry in the first place. Static rules produce static noise. The whole value of a co-managed model rests on a feedback loop that actually closes, where what analysts flag as noise this month changes what the system filters next month. Without that loop, the noise hasn't been reduced. It's just been handed to someone else to file.
Sources
- What Percentage of SOC Alerts Are False Positives (And How Many Actually Need a Human)?
- Can SOC Alert Triage Be Automated to Cut Analyst Load?
- Why False Positives Are Still Killing Security Teams | OP Innovate
- What Is Alert Fatigue? Causes, Impact & How to Reduce It
- How do you reduce false positives in SOC operations?
- Can SOC Alert Triage Be Automated to Cut Analyst Load?


