Also: incident management, ITIL incident, service desk, major incident, problem management
Incident management is the service-management practice of restoring a disrupted IT service to normal operation as quickly as possible, through a defined process of logging, categorising, prioritising, resolving and closing each incident, separate from the later work of finding and fixing its root cause.
Our take. The single most useful sentence in service management is that an incident is not a problem. The incident is the outage; the problem is why it keeps happening. Teams that try to fix the cause during the outage restore service slowly, and teams that never look for the cause afterwards restore the same service every week. Two practices, two queues, two kinds of person.
Term
Means
Not to be confused with
Incident
An unplanned interruption or reduction in quality of a service
A request for something new, which is a service request
Major incident
An incident with high impact needing a coordinated response and communication
Any incident a senior person is upset about
Problem
The underlying cause of one or more incidents
The incident itself
Workaround
A way to restore service without fixing the cause
A fix
Known error
A problem with a documented cause and workaround, not yet fixed
A closed problem
Service level
The agreed target for how quickly incidents of each priority are resolved
A promise about when the cause will be fixed
Something is broken and people cannot work. The incident practice exists to get them working again, and everything about it follows from that goal. Log it, so it exists and can be tracked. Categorise it, so it reaches people who know that system. Prioritise it by impact and urgency, so the payment system outage outranks the broken mouse. Diagnose enough to restore, which may mean a restart, a failover or a workaround rather than a fix. Restore the service. Close the incident with the user's agreement. The root cause, if it is not obvious and fixed along the way, goes to a different practice with a different pace, because finding out why under the pressure of an outage produces bad answers and slow restorations.
This is the vocabulary that the ITIL Foundation certificate examines and that every large organisation's service desk speaks, and it is the source of endless confusion because the words are ordinary and the meanings are specific. An incident is not a problem. A workaround is not a fix. A service request is not an incident. The platform certification for the ticketing system most enterprises run tests the same distinctions inside one product. Outside service management the word incident belongs to security, where incident response means a coordinated reaction to a compromise; the two practices share a name, a blameless review and little else.
In practice
Every Monday morning the finance system is unavailable for twenty minutes and the service desk restores it by restarting a service. Each Monday's incident is logged, prioritised, restored inside its target and closed, and the metrics look fine. Nobody has opened a problem. A service manager notices the pattern, raises a problem record, and the investigation, done calmly on a Wednesday afternoon rather than during the outage, finds a weekend batch job that exhausts a connection pool. The fix takes a day. The incidents stop. The incident practice had worked perfectly every week and could never have stopped them, because stopping them was not its job.
Every Monday: an incident, restored inside target, closed. Metrics green.
One Wednesday: a problem record, a calm investigation, a batch job found.
Security incident response reacts to a compromise: containing an attacker, preserving evidence, notifying regulators. Service-management incident management restores a broken service. Same word, different practices; a ransomware outage involves both at once and needs both teams.
Change control on a project governs alterations to a baseline. In service management, change enablement governs alterations to live systems, and a large share of incidents are caused by changes, which is why the two practices are linked in every framework.
A major incident review borrows the retrospective's blameless stance for a single event. A retrospective examines a team's ordinary working over a period. Both feed improvement; only one is triggered by something breaking.
Key takeaways
→Restore service first; find the cause later, in a separate practice at a calmer pace.
→Incident, problem, workaround, known error and service request are precise terms, and the exams test the precision.
→Perfect incident metrics can coexist with the same outage every week; that is problem management's queue.
Certifications that test this
Vendor exams whose syllabus covers this concept — facts, cost and a preparation path on each page.
A rotating selection from the course directory, drawn from the subcategories where this concept is taught rather than picked for it. Details, price and the provider link are on the course page.
With the growing threat of climate change due to the excessive release of carbon emissions, many nations are looking to…
Udemy
FAQ
What is the difference between an incident and a problem?
An incident is the interruption; a problem is its cause. Incidents are restored quickly, often with a workaround; problems are investigated deliberately and fixed. Keeping the two in separate queues with separate priorities is the framework's central design decision.
What makes an incident a major incident?
Impact and urgency high enough to need coordination across teams and proactive communication to the business, as defined by the organisation's own thresholds rather than by who is complaining. Major incidents have a manager, a bridge, regular updates and a review afterwards.