
An AI SOC puts an AI SOC analyst in front of the alert queue. It closes, escalates and explains alerts on its own. That changes what a SOC has to measure. Alert volume and mean time to close still matter, but they say nothing about whether the verdicts were right.
Why an AI SOC needs its own scorecard
The gap between adoption and proof is wide. The SANS 2026 SOC Survey found that 79% of SOCs use AI or ML tools. Only 36% have integrated them into a defined SOC workflow. A year earlier, the SANS 2025 SOC Survey found that 69% of SOCs still report their metrics manually or mostly manually.
Analysts at Gartner expect 70% of large SOCs to pilot AI agents for Tier 1 and Tier 2 work by 2028. They also expect only 15% to achieve measurable improvement without structured evaluation, as reported by Help Net Security. This guide covers structured evaluation.
The guide covers five things: ground truth, the core metrics, sampling, the move from shadow mode to production, and the numbers worth reporting upward.
Build ground truth before you measure
Every accuracy metric compares a verdict to a correct answer. If you have no agreed correct answer, you have no metric. Ground truth comes first.
Label alerts from your own environment
Vendor benchmarks run on vendor data. Your alert mix, log sources and detection rules differ. The NIST AI RMF Measure playbook asks you to document where results stop holding outside the conditions a system was built for. It also asks for false positive and false negative rates measured against ground truth within the deployment context.
For an AI SOC, that means a labeled set of your own recent alerts. Take a few weeks of closed alerts across every alert class you plan to hand over. Include the rare classes, since those are where verdicts fail quietly.
Settle disagreements between analysts
Two experienced analysts will sometimes label the same alert differently. Have two people label each alert in the set independently. Send disagreements to a senior reviewer, and record the reason for the final label.
The disagreement rate is useful on its own. If your analysts agree on 90% of labels in a class, an AI cannot be scored with more precision than that.
Keep the set current
Environments change. New log sources, new detection rules and new attacker behavior all shift the alert mix. The NIST playbook calls for monitoring AI systems in production, because they may meet new issues as the environment evolves. Refresh the labeled set on a fixed schedule, and after any major detection or logging change.
The five metrics that matter
These five metrics cover correctness, trust and cost. Measure each one per alert class, then roll up.
| Metric | Question it answers | How to compute it | Warning sign |
|---|---|---|---|
| Missed true positive rate | How often does the AI close a real threat as benign? | Real threats closed as benign, divided by all real threats in the sample | Any confirmed miss in a high-severity class |
| Escalation precision | How much of what the AI escalates is real? | Confirmed threats divided by all AI escalations | Precision close to your pre-AI queue |
| Evidence completeness | Can an analyst verify each verdict from what it cites? | Share of sampled verdicts where every cited artifact exists and supports the conclusion | Verdicts that cite summaries instead of raw events |
| Analyst minutes per true positive | Is the team finding real threats faster? | Analyst time spent on the queue divided by true positives confirmed | Time drops while missed true positives rise |
| Override rate by direction | Do analysts trust the AI too much or too little? | Share of verdicts analysts reverse, split into benign-to-malicious and malicious-to-benign | Near-zero overrides with no seeded tests |
Missed true positive rate
This is the metric that matters most. An AI SOC that escalates too much wastes time. An AI SOC that closes a real intrusion as benign loses it.
Overall accuracy hides this failure. A 2026 survey of alert-triage research cites a field measurement where roughly 0.01% of daily alerts linked to true attacks. At that base rate, a system that labels every alert benign scores 99.99% accuracy. Report the missed true positive rate directly, and never let accuracy stand in for it.
Escalation precision
Precision tells you whether the queue analysts see is cleaner than before. Compare it to the precision of your pre-AI queue for the same alert classes. If the two numbers are close, the AI is sorting alerts without adding judgment.
Evidence completeness
A verdict is only checkable if it points to the raw events behind it. Sample verdicts weekly. For each one, confirm that every cited log line, process tree or network flow exists and supports the conclusion.
This metric also protects against confident wrong answers. A verdict with a plausible narrative and missing evidence should count as a failure, even when the label happens to be correct.
Analyst minutes per true positive
Time saved is the number vendors lead with. Tie it to outcomes. Microsoft researchers published a randomized controlled trial of their Phishing Triage Agent as a preprint. It used true positives per analyst minute as its productivity metric.
The trial ran 167 external analysts through a 25-email queue built from real user-reported phishing. Analysts working with the agent found up to 6.5 times as many true positives per analyst minute. Their verdict accuracy, measured by F1 score, improved 77% over the control group.
The trial also split the gain by source. About 83% of the productivity gain came from queue prioritization, and 17% came from the agent's verdicts. Run the same split on your data. It tells you which part of the AI SOC is doing the work.
Override rate by direction
Override rate shows how analysts treat AI verdicts. A near-zero rate can mean the AI is right. It can also mean analysts have stopped checking.
The Microsoft trial tested for this directly. Researchers planted synthetic false positives and false negatives in the agent's output to simulate 80% accuracy. They then compared how often analysts confirmed real agent findings against planted wrong ones. The analysts were no more likely to confirm the wrong verdicts, so the researchers concluded they were not rubber-stamping.
You can run the same test in a training or shadow queue. Seed a small number of known-wrong verdicts each month, and track how many analysts catch.
How many closed alerts to sample
Nobody reviews every alert an AI SOC closes. That makes sampling the only way to estimate how many real threats hide inside benign closures.
The statistical rule of three gives a quick bound. Say a random sample of n alerts contains zero misses. The 95% upper bound on the true rate of misses is then about 3 divided by n.
What the rule of three means in practice
A sample of 300 benign closures with zero misses bounds the share of hidden threats at about 1%. A sample of 3,000 bounds it at about 0.1%. Zero misses in a small sample proves less than it appears to.
Take a class that closes 5,000 alerts a month as benign. If 1% of those closures hid real threats, that would be 50 missed threats a month. For high-severity classes, size the sample so the bound falls below the rate you can accept.
Stratify the sample
Easy alert classes inflate averages. A SOC with thousands of low-risk closures and a few hundred identity alerts can report a strong overall number and still miss identity attacks. Sample each alert class separately, and weight the sample toward the classes with the highest impact.
From shadow mode to production
Measurement changes as the AI SOC takes on more of the queue. Plan three stages and decide the exit criteria for each before you start.
Stage 1: shadow mode
The AI produces verdicts, and analysts work the queue as they do today. Nobody acts on the AI output. Compare AI verdicts to analyst verdicts and to your labeled set.
Shadow mode gives you the cleanest missed true positive rate you will get, since analysts review everything. Run it long enough to cover your rare alert classes.
Stage 2: assisted triage
Analysts see AI verdicts and evidence before they decide. Track override rate by direction and run the seeded-verdict test. Track analyst minutes per true positive against your shadow-mode baseline.
Stage 3: autonomous closure, class by class
Grant autonomous closure one alert class at a time. Grant it only where the missed true positive bound meets the threshold you set. Keep sampling that class after the change. The NIST playbook's point about production monitoring applies here: a class that passed in March can drift by June.
When detection content changes, re-baseline that class. Detection quality drives alert quality, so the discipline in Sigma rules done right affects AI SOC metrics too. Replaying a detector against months of historical data before release shows how the alert mix will shift.
What to report to leadership
Leadership needs a short monthly view that shows risk and cost together. Five numbers are enough:
- Missed true positive upper bound, per high-severity alert class
- Escalation precision, compared to the pre-AI baseline
- Analyst minutes per true positive
- Override rate by direction, plus the seeded-verdict catch rate
- Share of alert volume under autonomous closure, by class
Leave "alerts closed by AI" out of the headline. It measures activity. A rise in closures with a rising miss bound is a warning, even though the volume chart looks good.
Collect these numbers from the case management system, not from spreadsheets. With most SOCs still reporting by hand, automated collection is what makes a monthly cadence last. It also keeps analysts off reporting work, which matters for sustainable on-call design.
Where AI SOC numbers mislead
Four patterns produce good-looking numbers that do not hold up:
- Benchmarks on someone else's data. Demo results do not transfer to your alert mix. Test on your labeled set.
- Ground truth from the AI's own closures. If the AI's benign verdicts become labels, its miss rate will read as zero. Labels must come from analysts.
- Averages across alert classes. High-volume, low-risk classes hide failures in rare ones. Report per class.
- Time savings without accuracy. Faster triage with more misses is a worse SOC. Always pair speed with the missed true positive rate.
How we approach measurement
We built Soc0 to attach a confidence score and a full evidence trail to every verdict. Each conclusion points back to raw evidence, and autonomy is set per action type. That design supports the evidence completeness and class-by-class autonomy checks above. The metrics in this guide apply to any AI SOC, including ours.
See Port0 on your own data.
Bring your noisiest alert queue. Watch Soc0 investigate it live.






