
Imagine that someone breaks into an office building.
The organisation has spent a lot of time and money trying to prevent this from happening:
- Doors are locked
- Security guards are present
- CCTV is installed
- Alarms are fitted
- Staff have ID cards
- Sensitive areas are restricted
But eventually, somehow, an attacker gets inside.
At that point, the organisation needs to know:
What has happened? What has the attacker accessed? How do we contain the problem? How do we remove the attacker? And how do we get back to normal?
This is where Incident Response comes in.
Incident response is the organised process used to identify, investigate, contain, eradicate and recover from security incidents.
Security controls try to prevent incidents. Incident response deals with them when they happen.
Incident response is therefore not a single technology or tool. It is a combination of people, processes and technology.
What Is a Security Incident?
Before discussing incident response, we need to understand what an incident actually is.
A security incident is an event that could compromise the confidentiality, integrity or availability of information or systems.
Examples include:
- Malware infection
- Ransomware
- Phishing
- Account compromise
- Data theft
- Unauthorised access
- Denial-of-service attacks
- Insider misuse
- Lost or stolen devices
- Accidental disclosure of sensitive information
The organisation now needs to respond.
An Event Isn’t Necessarily an Incident
This is an important distinction.
Organisations generate enormous numbers of security events – depending on the size of the organisation, this can number in the millions every day
Most of these are perfectly normal – Just staff and other services going about their every-day business.
An incident is generally a sequence of unusual events, which in isolation are probably nothing, but when chained together become something that requires investigation or action because it represents, or could represent, a security problem.
This is why technologies such as SIEM, EDR and EUBA are useful: they help identify the events that deserve attention.
Incident Response vs Incident Management
The terms incident response and incident management are sometimes used interchangeably.
They are closely related, but it is useful to distinguish them.
Incident response focuses on the technical and operational actions needed to deal with a security incident, whereas incident management is broader and can include:
- Business impact
- Communications
- Escalation
- Management decisions
- Legal considerations
- Regulatory requirements
- Customer communications
- Public relations
- Recovery
For simplicity, this post will use incident response to cover the overall security response process.
Why Is Incident Response Important?
No organisation can guarantee that it will never suffer a security incident – Even organisations with excellent security controls can be attacked.
The objective is therefore not simply to prevent every attack. Rather it is to
“Prevent attacks where possible, detect them quickly when they occur, minimise their impact and recover effectively.”
The faster and more effectively this process happens, the less damage an incident may cause.
The Incident Response Lifecycle
A typical incident response lifecycle can be represented as follows:

The process is cyclic – Lessons learned from one incident should improve the organisation’s ability to deal with the next one.
1. Preparation
Good incident response starts before an incident occurs – Preparation involves making sure the organisation is ready to respond.
This can include:
- Incident response plans
- Response procedures
- Contact lists
- Security tooling
- Logging
- Backups
- Staff training
- Incident playbooks
- Communication procedures
- Tabletop exercises
If preparation is poor, responding to a major incident can become chaotic.
An organisation should have a documented incident response plan.
It should explain things such as:
- Who is responsible for what?
- Who declares an incident?
- Who needs to be contacted?
- Who has authority to isolate systems?
- How is evidence preserved?
- How are senior management informed?
- How are customers informed?
- How is recovery authorised?
The plan should be practical rather than simply existing as a document nobody has read.
Larger organisations may have a dedicated Incident Response Team.
This might include Key member from the security, IT, Management and HR teams, but the exact team depends on the organisation and the incident.
Everyone involved should understand their role.
For example:
- Security Team – Investigates the technical incident.
- IT – Helps isolate, rebuild and recover systems.
- Management – Makes business decisions and sets priorities.
- Legal – Provides advice on legal and regulatory requirements.
- HR – May become involved in insider incidents.
- Communications / PR – Handles internal and external communications where necessary.
One way of ensuring each team member knows what their roles and objectives are is to use Incident Response Playbooks
A playbook is a predefined set of actions for dealing with a particular type of incident.
For example:
RANSOMWARE PLAYBOOK
1. Confirm incident
2. Isolate affected systems
3. Preserve evidence
4. Identify scope
5. Disable compromised accounts
6. Contain malware
7. Investigate
8. Eradicate
9. Restore systems
10. Review incident
A playbook reduces the need to invent procedures while under pressure.
Organisations may have playbooks for:
- Ransomware
- Phishing
- Malware
- Account compromise
- Data breach
- Lost device
- Insider threat
- DDoS
- Cloud compromise
- Privilege escalation
Different incidents require different responses.
2. Detection
The next stage is identifying that something has happened and determining whether it is actually an incident or not.
Detection may come from:
- SIEM
- EDR
- IDS/IPS
- Firewall
- EUBA
- Antivirus
- Users
- Security researchers
- External organisations
- Threat intelligence
Remember that alerts are not automatically incidents – A security team may receive thousands of alerts.
An analyst needs to determine – “Is this actually malicious?“
This initial process is often called triage.
3. Incident Triage
Triage means quickly assessing an alert to determine:
- What happened?
- Is it real?
- How serious is it?
- What systems are affected?
- Is the attack still happening?
- Does it require escalation?
Not every incident has the same impact, so organisations need to determine the incident severity.
An organisation may classify incidents as:
- LOW – Limited impact
- MEDIUM – Significant impact
- HIGH – Major business impact
- CRITICAL – Severe / widespread impact
The severity may depend on:
- Number of systems affected
- Sensitivity of information
- Business disruption
- Financial impact
- Regulatory implications
- Whether the attacker still has access
4. Containment
Once an incident is declared, the organisation needs to contain it – The objective is to stop the attack from spreading or causing additional damage.
Short-Term Containment
Immediate containment might involve:
- Isolating a device
- Blocking an IP address
- Disabling an account
- Revoking sessions
- Blocking malicious domains
- Disconnecting a server
- Stopping a malicious process
Long-Term Containment
It may be that the organisation may need more extensive containment – The objective is to maintain essential business operations while preventing the incident from spreading.
This is one reason network segmentation is so valuable.
Imagine ransomware reaches a workstation. Without segmentation the ransomware could easily reach across the entire network in minutes, affecting other workstations, servers, databases, backup systems, and even cloud services.
With segmentation however, the ransomware would be limited just to the impacted network segment
Segmentation can also limit an attacker’s ability to move laterally.
Limit account access
Sometimes the most important containment action is disabling or restricting an account.
This can be particularly important when dealing with stolen credentials.
Preserve evidence
During an incident, the organisation may need to preserve evidence.
Evidence might include:
- System logs
- Network traffic
- Memory
- Disk images
- Malware samples
- Emails
- Authentication records
- File timestamps
- Cloud activity
One of the challenges of incident response is that the obvious action isn’t always the best one.
Suppose malware is found on a computer. The instinct might be to simply wipe the computer and rebuild it.
But doing that immediately could destroy valuable evidence – The correct response depends on the incident and organisational procedures.
Sometimes containment must happen immediately; sometimes evidence should be collected first.
This is why incident response needs planning.
5. Eradication
Once the attack has been contained, the organisation needs to remove the cause of the incident.
This is eradication. The objective is to ensure that the attacker can no longer maintain access.
Simply deleting malware may not be enough.
The attacker may have:
- Created additional accounts
- Installed backdoors
- Added scheduled tasks
- Stolen credentials
- Modified configurations
- Created persistence mechanisms
- Compromised other systems
The organisation needs to understand the full scope of the compromise, and ideally identify the root cause.
Incident response should attempt to identify how the incident happened.
The root cause might therefore be related to:
- Phishing
- Vulnerable software
- Weak credentials
- Misconfiguration
- Excessive privileges
- Poor access controls
Sometimes there are several contributing factors to the root cause of an incident.
Incident response should identify all the weaknesses that allowed the incident to occur and spread.
6. Recovery
After the attacker has been removed, systems need to be returned to normal operation.
This is the recovery stage.
Recovery should be a carefully controlled process.
Good backups can make a huge difference during incidents such as ransomware – But backups can be attacked too.
Modern attackers understand the value of backups.
They may attempt to:
- Delete backups
- Encrypt backups
- Steal backup credentials
- Disable backup systems
Therefore, backups should themselves be protected and before they are used to restore data – they should be verified that they have not been compromised.
Recovery Doesn’t Mean “Everything Is Fine”
A system may appear to be working again while an attacker still has access.
Recovery should therefore include validation that the environment is clean.
Verification
Before returning systems to normal operation, organisations should verify:
- Malware has been removed
- Vulnerabilities have been fixed
- Credentials have been changed
- Persistence has been removed
- Security controls are functioning
- Logging is working
- Systems are correctly configured
7. Lessons Learned
One of the most important parts of incident response happens after the immediate crisis is over.
The organisation should ask itself – “What can we learn from this?“
For example:
- What happened?
- Why did it happen?
- How did it happen?
- What worked?
- What failed?
- What needs to be changed?
This is often called a lessons-learned review or post-incident review.
The incident should make the organisation stronger.
Incident Response and other controls
SIEM
A SIEM can provide important information during incident response. The security team can use the SIEM to build a timeline of events.
EDR
EDR can be particularly useful when investigating compromised endpoints.
It can help answer questions such as:
- What process started?
- What files were created?
- What commands were executed?
- What network connections occurred?
- What account was used?
IDS/IPS
IDS and IPS can provide evidence of network-based attacks. This information can help reconstruct what happened.
EUBA
EUBA can help identify unusual behaviour. This can be particularly useful for account compromise and insider-threat investigations.
SOAR
SOAR can automate parts of the response process. Automation can reduce the time between detection and response.
The Importance of Speed
Time matters enormously during an incident. Early detection and containment can interrupt the progression of the attack.
Security teams often measure how long it takes to identify an incident – This may be referred to as Mean Time to Detect (MTTD).
The shorter this period, the sooner the organisation can respond.
Another useful measurement is Mean Time to Respond (MTTR) – This looks at how long it takes to take meaningful action after detection.
Organisations may use these metrics to evaluate their incident response capabilities.
Communication During an Incident
Incident response isn’t purely technical – Communication is extremely important.
During a serious incident, organisations may need to communicate with:
- Employees
- Management
- Customers
- Suppliers
- Regulators
- Law enforcement
- Insurance providers
- External specialists
Poor communication can make an already serious incident worse.
Incident Response Exercises
Organisations shouldn’t wait for a real incident to test their response.
They can conduct tabletop exercises (TTXs)
The participants talk through their response without actually attacking systems.
More advanced organisations may conduct technical exercises which can expose weaknesses that aren’t obvious from reading a plan.
The NCSC provides a series of free TTX exercises organisations can use to test various incident scenarios, and how team members should respond – This is called Exercise in a box
Incident Response and Backups
Incident response should be closely connected to backup and recovery procedures.
A backup that has never been tested isn’t a particularly reassuring recovery strategy, so organisations should periodically test whether backups can actually be restored.
This is part of resilience.
Incident Response and Business Continuity
Incident response focuses on dealing with the security incident. Business continuity focuses on keeping critical business operations functioning.
Whilst they overlap, a mature organisation should consider both.
Incident Response and Disaster Recovery
Disaster recovery focuses on restoring systems and services after a disruptive event.
An incident may trigger a disaster recovery processes.
The two disciplines work together but have different primary objectives.
In Summary
Incident response is the organised process of dealing with security incidents when they occur.
It typically involves:
- Preparation – Make sure people, processes and technology are ready before an incident occurs.
- Detection – Identify suspicious activity and determine whether it is a genuine incident.
- Triage – assess its severity.
- Containment – Stop the incident from spreading or causing further damage.
- Evidence Preservation – Collect and protect information needed to understand what happened.
- Eradication – Remove the attacker’s malware, persistence, access and other causes of the compromise.
- Recovery – Restore systems and services to normal operation and verify that they are secure.
- Lessons Learned – Understand what happened, why it happened and what can be improved.
Incident response works closely with many of the other security controls covered in this series:
- Firewalls help prevent unwanted network connections.
- Access Controls restrict who can access resources.
- MFA makes stolen credentials harder to use.
- Network Segmentation limits lateral movement.
- Encryption protects information.
- Hardening removes unnecessary attack opportunities.
- Patch Management fixes known vulnerabilities.
- Logging and Auditing provide evidence of what happened.
- IDS/IPS detect and potentially block attacks.
- EDR provides visibility into endpoint activity.
- EUBA identifies unusual behaviour.
- SIEM brings security events together.
- SOAR can automate parts of the response.
- Honeypots can provide early warning and attacker intelligence.
The key lesson is that incident response isn’t an admission that security has failed – It is an acknowledgement that security incidents can still happen even when good controls are in place.
The real measure of an organisation’s security maturity isn’t simply whether it has been attacked, it’s how quickly it can detect the attack, how effectively it can contain it, how well it can recover, and whether it learns from what happened.
You can’t guarantee that an incident will never happen. You can make sure you’re ready when it does.