SOC Automation: What to Automate First, Stage by Stage

A six-stage order for SOC automation, from evidence gathering to contained response, with the research behind each stage.

The Port0 owl on the third of a rising chain of snowy floating islands, with faded islands still ahead along a dashed path

Automate SOC work in the order you can prove it. Evidence gathering first, then noise filtering, then queue ordering, then correlation, then verdicts on alert classes you can already describe in writing, then contained response. Each step earns the next. Skipping to autonomous response is where automation programs stall, because nothing upstream produces the evidence a response decision needs.

Tools are bought, workflows are missing

The 2026 SANS SOC Survey, published June 15, 2026 from 444 qualified responses, found 79% of SOCs use AI or ML tools. Only 36% have integrated those tools into a defined SOC workflow.

That 43-point gap explains a lot of disappointment. The capability is bought. The decision it is supposed to make has never been written down.

So start from the decision, then pick the product. Find a decision your team already makes the same way every time. Then check whether you can describe it in enough detail for a machine to verify your reasoning.

The six stages at a glance

StageWhat it doesRisk if wrongReversible
1. Evidence gatheringAttaches context to every alertNone, no verdict changesYes
2. FilteringSuppresses demonstrable noise before the queueMissed alert, caught by samplingYes, with a sample control group
3. Queue orderingRanks what an analyst opens firstDelayed handlingYes
4. CorrelationGroups alerts into one incident narrativeSplit or merged incidentsYes
5. VerdictsCloses alert classes with written proceduresMissed intrusionOnly if closures are logged and reopenable
6. ResponseExecutes containment actionsBusiness disruptionDepends on the action

Stage 1: evidence gathering on every alert

Gather context for every alert before anyone opens it. Resolve the asset, the identity, the recent authentication history, the process lineage, the destination reputation, and prior alerts on the same entities.

This stage has the best risk profile in the whole program. It changes no verdict and closes nothing. It only removes the twenty minutes an analyst spends pulling records from four consoles.

Automate it first because it is reversible, auditable, and useful even if every later stage fails. It also produces the labelled evidence that later stages need.

Stage 2: filtering before the queue

Once evidence is attached, suppress alerts that are demonstrably noise. This is pre-queue filtering, and it is the most studied part of the pipeline.

A 2026 survey of AI-driven alert screening catalogues the production results. Ban et al. (2023) cut false positive count by 53% in enterprise logs. An earlier system from the same group reported 99.598% recall at a 0.001% false positive rate on a highly critical alert class.

Both results come from production enterprise logs, which is rare in this field. Filtering is safe to automate early because a suppressed alert can still be logged, sampled, and reviewed.

Sample what you suppress

Keep 1% to 5% of suppressed alerts flowing to a human queue as a control group. If suppression is wrong, the sample tells you within a sprint instead of after an incident.

Stage 3: queue ordering

Ranking carries less risk than closing. Nothing is discarded, and the ordering only changes what an analyst sees first.

The same survey reports Aminanto et al. (2020) reaching 0.93 AUC with a 31% reduction in false positive rate. The evaluation used day-forward-chaining, so the model never trained on future data. That evaluation design matters more than the score. Models that look strong on random splits degrade in production, a failure the survey attributes to temporal drift documented by Pendlebury et al. (2019).

Ask any vendor how their model was split for evaluation. Chronological or nothing.

Stage 4: correlation into incidents

Correlation groups related alerts into one narrative: one identity, one host chain, one session. It cuts duplicate investigations, which is often the largest single time saving in a SOC.

Research results here look excellent, then travel poorly. KAIROS (Cheng et al., 2024) reports 0.99 AUC on DARPA Transparent Computing provenance graphs. The survey notes that only a minority of correlation papers report latency, memory overhead, or analyst-time savings under realistic enterprise loads.

Treat published correlation accuracy as an upper bound. Measure duplicate investigation rate in your own environment before and after.

Stage 5: verdicts on alert classes you can write down

Automated closure is the first stage where a mistake means a missed intrusion. Gate it behind one test. Can you write the triage procedure for this alert class, step by step? Include the evidence that would change your mind.

If you can, automation has something to check. If you cannot, you are asking a model to invent a standard you never set.

Start with two or three high-volume classes where verdicts are mechanical. Impossible travel with a known VPN egress. Antivirus quarantine events already contained. Failed authentication bursts from a known scanner range.

Run it in shadow mode first

Let the system produce a verdict without acting on it. Compare against analyst decisions for four weeks. Promote only the classes where disagreement is rare, and where every disagreement you inspect turns out to favour the machine.

Stage 6: response, with the blast radius set

Response automation splits cleanly by reversibility. Blocking an indicator on a perimeter device is cheap to undo. Isolating a domain controller or disabling an executive's account is not.

Automate the reversible half. Put approval gates on the rest. High-impact actions such as host isolation or account disablement should wait for a human yes. An AI analyst that investigates and acts inside set guardrails holds those actions until you decide otherwise.

What stays with analysts

Three things resist automation for structural reasons. A better model does not fix any of them.

Base rates. Field measurement cited in the 2026 survey found roughly 0.01% of daily alerts linked to true attacks. At that base rate, a small error rate still produces a large number of wrong calls. Novel, low-volume, high-impact alerts also give a model the least data to learn from.

Unquantified failure modes. The survey notes that hallucination in deployed generative SOC agents remains unquantified. It also finds that adversarial evasion is seldom reported in production alert-filtering studies. Absence of published failure data is not evidence of safety.

Judgement about consequences. Whether to pull an executive's laptop during a board meeting is a business decision. Identity remains the most common route in. Expel's 2026 report, covering 2025 incident data, found 47.7% of identity incidents ended with attackers gaining account access using stolen credentials. Those cases turn on context a SOC often does not hold.

What has to be true before stage 1

Four preconditions, and each of them is boring.

  1. Query access to data where it lives. If enrichment requires a ticket to another team, automation stops at the ticket.
  2. An audit trail per conclusion. Every automated verdict points back to the raw records behind it. Without that, nobody can review a wrong call, and trust never builds.
  3. Written triage procedures for the classes you intend to close. Detection content you can read, version, and test. Sigma rules that survive production are a reasonable model for the discipline involved. Tooling that lets you replay a detector against months of historical data before release makes the standard testable.
  4. A rollback path for every automated action.

How to tell whether it worked

Mean time to respond moves for reasons that have nothing to do with automation, including a quiet month. Track these instead, per alert class:

  • Analyst touch rate. Share of alerts in the class that a human opens. This is the number automation is supposed to move.
  • Time to first evidence. Minutes between alert creation and complete context being attached.
  • Reopen rate after automated closure. Any sustained rise here stops promotion of new classes immediately.
  • Duplicate investigation rate. Investigations that turn out to belong to an incident already being worked.
  • Coverage of written procedures. Share of alert volume belonging to classes with a documented triage standard.

Report the reopen rate next to the touch rate, always. The first number without the second rewards closing alerts rather than resolving them.

A 90-day sequence

Weeks 1 to 4. Pick the two highest-volume alert classes. Automate evidence gathering for both. Measure time to first evidence before and after.

Weeks 5 to 8. Add pre-queue filtering with a 5% sample flowing to humans. Run verdict generation in shadow mode on the same two classes.

Weeks 9 to 12. Promote verdicts on whichever class showed the lower disagreement rate. Add reversible response actions only. Write the approval policy for everything else before anyone asks for it.

At the end of the quarter you will know your real automation ceiling, measured on your own alerts rather than on a vendor benchmark. That number is worth more than any pilot result.

See Port0 on your own data.

Bring your noisiest alert queue. Watch Soc0 investigate it live.

Book a Demo

Never Miss an Insight

Subscribe to get the latest posts delivered to your inbox.