
An agentic SOC is a security operations model where AI agents plan and run investigations on their own. The agent decides what evidence to pull, pulls it, weighs what it finds, and returns a scored verdict. Nobody writes the steps in advance.
That last sentence carries the whole definition. Automation in the SOC has existed for a decade. What changed is who chooses the next step.
What an agentic SOC actually is
A useful test: give the system an alert type it has never seen, with no matching playbook. A scripted system fails or falls through to a human queue. An agentic system reasons about the alert, picks tools, and produces an answer you can audit.
The category is young and crowded. Gartner moved AI SOC Agents from Innovation Trigger to the Peak of Inflated Expectations in a single year. It also warns buyers about "AI washing" in the market. So the definition matters more than usual right now.
Agentic SOC compared to SOAR and AI-assisted triage
Three things get sold under similar language. They behave differently under load.
SOAR runs the playbook you wrote
SOAR executes predefined sequences when conditions match. Enrich an IP, query the EDR, open a ticket, notify a channel. It is fast, deterministic, and fully auditable.
Splunk SOAR, Tines, and Torq are good at this. Tines and Torq in particular made playbook authoring fast enough that a SOC can maintain real coverage without a dedicated automation engineer.
The limit is coverage, and it's structural. Every branch is a branch someone authored. Detection content changes weekly and playbooks don't keep pace. Teams end up maintaining hundreds of playbooks that each handle one alert shape.
Copilots answer questions you ask
A copilot summarizes an alert, translates plain language into a query, or drafts an incident note. The analyst still owns the loop. They decide what to ask and when to stop.
Microsoft Security Copilot and CrowdStrike Charlotte AI are strong here. Both sit close to their own telemetry, so their answers are grounded in data the vendor already holds.
Copilots reduce typing and context switching. They don't reduce the number of alerts a human opens, because a human still opens each one to ask the question.
Agents own the loop
An agent receives the alert and runs the investigation to a conclusion. It forms a hypothesis, gathers evidence across identity, endpoint, network, and cloud, discards what does not fit, and scores the result.
The output is a verdict with the evidence attached, ready for a human to accept or overturn.
What an agent needs before it can investigate anything
Failed agentic pilots usually fail on plumbing, not reasoning. Three prerequisites carry the weight.
Access to data where it already lives
An agent that can only see what a SIEM ingested is limited to what the SIEM ingested. Investigations cross systems that nobody pays to centralize: identity logs, SaaS audit trails, cloud control plane events, raw network telemetry.
Re-ingesting all of it to make it reachable is expensive and slow. Federated query across the tools a team already runs avoids the second copy and keeps the agent's view current.
Tool permissions and an agent identity
An agent that can query is a research tool. An agent that can isolate a host, disable an account, or block a domain is an actor in your environment. It needs the controls an actor gets.
That means a distinct identity, scoped permissions per action, and logging that attributes every action to that identity. Shared service accounts destroy attribution the moment something goes wrong.
An evidence trail per verdict
A verdict without its evidence is an opinion. Every conclusion needs the queries that produced it, the raw results, and the reasoning that connected them.
This is also your only defense against silent drift. When an agent starts closing a class of true positives, the trail is how you find out why.
Where agentic SOCs break
The failure modes are specific, and several of them are new.
The evidence is attacker-controlled
An agent investigating an alert reads logs, file paths, process command lines, email bodies, and user agent strings. An attacker writes some of that content.
NIST calls this agent hijacking. The definition: "an attacker inserts malicious instructions into data that may be ingested by an AI agent, causing it to take unintended, harmful actions". In NIST CAISI testing published in January 2025, baseline attacks succeeded 11% of the time against one frontier model. Adaptive red team attacks against the same model succeeded 81% of the time.
The SOC case is harder than the general case. A SOC agent's job is reading hostile input. That happens on every alert.
Three practical controls. Treat all ingested telemetry as untrusted data rather than instructions. Separate the reasoning context from the action authority. Require a second signal before any destructive action.
Confident wrong verdicts at scale
A human analyst who misreads an alert misreads one alert. An agent that misreads an alert class misreads every instance of it, quietly, until someone checks.
Research on multi-agent triage shows the direction of travel is good. The CORTEX architecture, published in September 2025, uses separate behavior analysis, evidence gathering, and reasoning agents. Its authors report that it "substantially reduces false positives and improves investigation quality over state-of-the-art single-agent LLMs."
Better than a single agent is still not the same as verified. You need sampling. Pull a random set of auto-closed alerts every week and have a human re-run them. Track the disagreement rate over time, not once at deployment.
Autonomy boundaries nobody wrote down
Ask any vendor exactly which actions the agent takes without approval. The honest answers cluster into three tiers: read-only enrichment, reversible containment such as session revocation, and irreversible action such as deleting a resource.
Write the tier boundaries into policy before deployment. Tie them to blast radius and reversibility, not to the agent's confidence score. A confident agent operating on poisoned evidence is confident for the wrong reason.
Evaluation against a demo instead of your data
Vendor demos run on vendor alerts. Your environment has a decade of naming conventions, legacy agents, and one application that generates a large share of your noise.
Gartner's forecast quantifies the gap. By 2028, 70% of large SOCs will pilot AI agents for Tier 1 and Tier 2 work. But only 15% will see measurable improvement without structured evaluation.
Run any pilot in shadow mode against your own historical alerts, with known outcomes, before the agent touches a live queue.
Questions worth asking a vendor
Six that separate working systems from demos:
- Which specific actions run without human approval, and which need it?
- What does the agent see when the data is not in the SIEM?
- Show us a verdict we would disagree with, and the full evidence trail behind it.
- How does the system behave when telemetry contains instructions aimed at the agent?
- What is the measured disagreement rate against human analysts on our alerts?
- Who holds the agent's credentials, and what is the permission scope per action?
The OWASP Agentic Security Initiative published "Agentic AI - Threats and Mitigations" in February 2025 as a threat-model reference for these systems. It is a reasonable baseline to hold a vendor against.
Where autonomy should stop
The goal is a SOC where humans keep the judgment calls. Analyst attrition is driven by interrupt load and repetitive triage. Sustainable on-call design matters more than raw headcount reduction.
Autonomy earns its place on the repetitive share: the enrichment, the correlation, the clear false positives that consume the shift. Judgment calls with business context, novel attacker behavior, and anything irreversible stay with a human.
We built Soc0 for exactly that split: an AI analyst that investigates every alert and acts inside guardrails you set. The agent does the gathering and the scoring. You decide what it's allowed to do with the answer.
If a vendor tells you that boundary does not exist, that is the answer to the autonomy question.
See Port0 on your own data.
Bring your noisiest alert queue. Watch Soc0 investigate it live.






