Content
Services
Industries
Talk to an expert

SOC automation: a CISO checklist for SOAR and playbooks

· 7 min read · Network Secure

Automating the SOC is not about buying a SOAR platform and switching on every playbook at once. It is about deciding, deliberately, which decisions the machine can make on its own, which need a human, and who is accountable when something goes wrong. This checklist structures that decision for the CISO: what to automate first, how to define autonomy levels, how to agree containment with the business and how to prove the automation is working.

Why automate and where to start

The core argument is time. NIST SP 800-61 Rev. 3, published in April 2025, reorganized incident response recommendations around the functions of the NIST CSF 2.0 and treats response as part of risk management rather than an isolated event. Faster detection and containment reduce impact, and automation helps with that. The real gain, however, comes from taking repetitive work off analysts so they can spend time on what requires judgment, as we explain in Traditional SOC vs. modern SOC.

The practical rule for choosing the first target: start with what is frequent, standardized and low-risk if it goes wrong. Leave for later what is rare, ambiguous or capable of stopping operations.

  1. Alert enrichment. Look up IP, domain and hash reputation, asset owner and criticality, user history and the related MITRE ATT&CK technique. It is read-only: it changes nothing in the environment.
  2. Triage. Group duplicate alerts, discard patterns already confirmed as false positives and adjust priority according to asset criticality. It requires documented rules and periodic review.
  3. Case creation and routing. Open the case in the ticketing system with evidence already attached and send it to the right queue with the right deadline. It removes the informal handoff between whoever detects and whoever acts.
  4. Low-risk containment. Reversible actions with limited reach: block an indicator already confirmed as malicious, quarantine a phishing email, force a password reset on a standard account, isolate a user workstation through EDR.

SOAR playbooks and autonomy levels

A playbook is the documented sequence of steps for a type of case: what to collect, what to check, when to escalate and what to execute. Before it becomes automation, it must exist on paper and work manually. CISA has published incident and vulnerability response playbooks for US federal agencies that serve as a structural model, and the open CACAO standard from OASIS defines a format for describing security playbooks so they can be shared and executed by different tools.

Every action in a playbook should have an explicit autonomy level. A simple four-step scale:

  1. Notify only. The automation detects, enriches and alerts. No action in the environment.
  2. Suggest. The automation proposes the action, with rationale and evidence. The analyst decides and executes.
  3. Execute with approval. The action is ready; an authorized person approves with one click and the platform executes and logs it.
  4. Execute. The platform acts on its own and reports afterwards. Reserved for pre-approved, reversible and tested actions.

A new playbook should start on the lower steps and move up only after an observation period in which its suggestions proved correct. The same action can have different levels depending on the target: isolating an employee's workstation can be automatic; isolating a payments server, never without approval.

Pre-approving containment with the business

The worst time to discuss whether the SOC may disable an executive's account is during the incident. That conversation must happen beforehand, with the business units, and be written down.

  • Action matrix by asset. For each asset class (workstations, servers, standard accounts, privileged accounts, critical systems), which actions are pre-approved and at which autonomy level.
  • Exception list. Assets and accounts that never receive automatic action, such as sensitive production systems and service accounts that support critical processes. For privileged accounts, see also privileged access management.
  • Owners and approvers. Who approves each type of action, with backups and after-hours contacts.
  • Communication criteria. Who is informed after an automatic action and through which channel.

Containment has a cost too. Isolating a server by mistake stops operations. Pre-approval exists to balance the risk of the attack against the risk of the response itself, and that decision belongs to the business, not to security alone.

Testing and rollback

Faulty automation fails at scale. Before moving a playbook up a level, confirm:

  • Testing in a controlled environment. Run the playbook against past real cases or in a lab before production.
  • Documented and tested rollback. Every containment action needs a known way back: unblock, re-enable, release from isolation.
  • Safety limits. A cap on actions per time window (for example, never isolating more than a defined number of machines without human approval) so a rule error does not become an incident.
  • Review on every change. Updates to tools, APIs and detection rules can silently break a playbook. Include playbooks in change management.

Metrics that show whether automation works

  • MTTR. Time from detection to containment. Compare cases handled with and without automation, by incident type.
  • Percentage of automated cases. How many cases were fully or partially resolved by a playbook. Not a goal in itself: automating the wrong case only speeds up the mistake.
  • False positives. The rate of false positives reaching analysts and, above all, automatic actions executed against something legitimate. The latter should be tracked case by case.
  • Reversals. How many automatic actions needed rollback, and why.

Governance and audit trail

Every automatic action must answer the same questions as a human one: what was done, when, on which asset, based on which evidence, by which playbook and version, and who approved it. In practice:

  • Immutable logging. Execution logs stored out of reach of those who operate the platform.
  • Playbook versioning. Change history with author, reviewer and date.
  • Least-privilege service accounts. The SOAR platform has access to firewall, EDR and directory; it is a target itself and needs protected, monitored credentials.
  • Periodic review. A committee or formal owner reviews autonomy levels, exceptions and metrics, aligned with the CSF 2.0 Govern function.

Common mistakes

  • Automating a process that does not exist. If the manual playbook is unclear, automation only locks the confusion in.
  • Starting with containment. Skipping enrichment and triage and going straight to automatic blocking is the fastest way to lose the business's trust.
  • Informal pre-approval. "We agreed by email" protects no one on the day of the crisis.
  • No playbook owner. A playbook without an owner goes stale and fails exactly when it is needed.

In practice: automation with human validation

In the Network Secure SOC, the Open-XDR platform handles triage, Threat Intelligence enrichment and automatic containment, while analysts validate, investigate what is critical and tune the rules. Learn more about the SOC and MDR service and how it connects to your incident response plan.

Frequently asked questions

What is SOAR?

SOAR (Security Orchestration, Automation and Response) is the category of tool that integrates security solutions and runs playbooks. Its value depends on the playbooks and governance, not just the tool.

Does automation replace SOC analysts?

No. It absorbs repetitive, standardized tasks. Investigating ambiguous cases, making decisions with business impact and improving detections still require people.

Which containment actions can be fully automatic?

Those that are reversible, limited in reach, pre-approved by the business and tested, such as quarantining a malicious email or isolating a user workstation. Actions on critical systems and privileged accounts should require human approval.

How do I know the automation is working?

Track MTTR by case type, false positives and the automatic actions that needed rollback.

Sources consulted: NIST SP 800-61 Rev. 3, Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile (April 2025); NIST Cybersecurity Framework 2.0 (CSWP 29), Govern function; CISA, Federal Government Cybersecurity Incident and Vulnerability Response Playbooks; OASIS, CACAO Security Playbooks Version 2.0; MITRE ATT&CK and MITRE D3FEND. The autonomy-level scale is a practical proposal of this article, not a standard's classification. This article is educational and does not replace an assessment of your environment and your response processes.

Official references

Read next