ITIL problem management vs incident management is one of the most searched distinctions in ITSM — and one of the most commonly confused in practice. Many service desks treat every recurring fault as a new incident, burning agent time on symptoms instead of causes. This post explains exactly how the two practices differ, where they connect, and how to run both in a way that actually stops the cycle of repeat tickets.
Why the Confusion Exists in the First Place
Incident management and problem management both deal with things going wrong. That surface similarity is what causes teams to collapse them into a single workflow — usually the incident queue.
The ITIL v4 framework, maintained by AXELOS, defines them as separate practices with different goals, owners, timelines and outputs. Incident management is about speed: restore service as fast as possible. Problem management is about understanding: find out why the incident happened and prevent it from recurring.
When organisations skip the separation, they get:
- Agents closing the same ticket type week after week
- No formal root cause analysis ever completed
- Known errors that live in someone's head rather than a knowledge base
- SLA pressure that always crowds out investigation time
The fix is not more staff. It is a cleaner process boundary between the two practices.
Incident Management: Restore First, Explain Later

An incident is any unplanned interruption to a service or reduction in service quality. The goal of incident management is to return the affected user or system to normal operation as quickly as possible — explanation is secondary.
What incident management covers
- Logging, categorising and prioritising the fault
- Matching against known errors or previous incidents
- Applying a workaround or fix to restore service
- Communicating status to the affected user
- Closing the ticket once service is confirmed restored
What incident management does not cover
- Investigating why the fault occurred in the first place
- Preventing the same fault from recurring
- Updating configuration baselines or change records
- Producing a post-incident review (that belongs to problem management)
Speed is the defining metric. IT service desk metrics like mean time to restore (MTTR) and first contact resolution (FCR) live here. The incident record closes when the user is back to work — not when the root cause is understood.
Problem Management: Find the Cause, Remove It Permanently

A problem is the underlying cause of one or more incidents. Problem management exists to identify that cause, document it, and either remove it permanently or produce a known error record so future incidents can be resolved faster until a permanent fix is in place.
The three phases of problem management
Reactive problem management starts when a pattern of incidents triggers a formal problem record. The team investigates, identifies the root cause and either implements a permanent fix or raises a change request to do so.
Proactive problem management starts before incidents occur. Teams analyse trend data, event logs and capacity reports to spot weaknesses before they surface as user-facing faults. This is harder to sustain but delivers the highest long-term value.
Known error management sits between the two. Once a root cause is identified but a permanent fix is not yet deployed, the team documents a workaround in a known error record. Agents can apply that workaround immediately the next time the related incident appears, cutting resolution time dramatically.
Outputs that matter
- A problem record with root cause documented
- A known error record with a tested workaround
- A change request if infrastructure or configuration must change
- An updated knowledge base article for the service desk
TIKTING links problem records directly to the incident tickets that triggered them and to any change requests raised as a result, giving teams a single thread to follow from symptom to resolution.
How the Two Practices Connect in Practice

Incident management feeds problem management — but the handoff is where most teams fail. Common failure points include:
- No formal trigger for when an incident should generate a problem record
- Problem records opened but never assigned an owner or deadline
- Root cause analysis started but abandoned when the next major incident hits
- Known errors documented once and never reviewed again
Building a working handoff
Define a clear trigger rule. Most teams use a combination of: the same category of incident occurring more than a set number of times in a rolling period, or any major incident that caused significant business impact. Both conditions should automatically prompt the service desk tool to suggest or create a problem record.
Assign a problem owner who is not the same person resolving day-to-day incidents. Problem investigation requires uninterrupted focus that the incident queue rarely allows.
Set a review cadence. Problem records without activity after a defined period should escalate automatically. A problem that sits open for months is a known risk with no one accountable for it.
Connect asset data. Many problems trace back to a specific configuration item — a failing disk model, an outdated driver, a misconfigured network device. Odysseus, the endpoint asset discovery solution from ITDEVTECH, pushes live hardware and software inventory into TIKTING, so when a problem record is raised the team can immediately see which assets share the same profile and assess blast radius before the next incident hits.
A Step-by-Step Process for Running Both Practices Together

Use this sequence to keep incident and problem management connected without letting one consume the other.
- Step 1 — Log every incident with consistent categorisation. Inconsistent categories make trend analysis impossible and problem triggers unreliable.
- Step 2 — Apply a workaround or fix and restore service. Close the incident ticket with the resolution method documented.
- Step 3 — At the end of each week, run a trend report. Flag any category that has exceeded your recurrence threshold.
- Step 4 — For each flagged category, open a problem record. Link all related incident tickets to it.
- Step 5 — Assign a problem owner and a target date for root cause analysis. Use a structured technique: five whys, fishbone diagram or fault tree analysis — whichever fits the complexity.
- Step 6 — Document the root cause and the workaround in a known error record. Publish a knowledge base article so agents can apply the workaround immediately on the next occurrence.
- Step 7 — If a permanent fix requires a configuration or infrastructure change, raise a formal change request through your change management workflow. Do not implement fixes ad hoc.
- Step 8 — After the fix is deployed, verify through monitoring that the incident category drops. Close the problem record only when evidence confirms the root cause has been removed.
- Step 9 — Conduct a quarterly review of all open and recently closed problem records. Identify any known errors that have been open too long without a permanent fix being scheduled.
Common Mistakes That Keep the Cycle Running

Even teams that understand the theory often fall into the same traps.
- Closing problem records when a workaround is found rather than when the root cause is permanently resolved. A workaround is not a fix.
- Letting the incident queue owner also own the problem queue. The urgency of incidents will always win.
- Running root cause analysis only after major incidents. Chronic low-severity incidents that never trigger a problem record quietly consume the most agent time over a year.
- Failing to update known error records when the environment changes. A workaround that worked on an old OS version may not work after a platform upgrade.
- Not linking problem records to the CMDB. Without knowing which configuration items are involved, the scope of a fix is guesswork.
The ITSMF and AXELOS both emphasise that problem management maturity is one of the clearest indicators of overall ITSM programme health. Organisations that invest in it consistently see lower incident volumes, shorter resolution times and higher user satisfaction over time.
Frequently Asked Questions
What is the main difference between incident management and problem management?
Incident management focuses on restoring service as quickly as possible after something goes wrong. Problem management focuses on finding and removing the underlying cause so the same thing stops happening. Incident management is reactive and time-critical; problem management is investigative and longer-term. Both practices are defined separately in ITIL v4.
When should an incident trigger a problem record?
Most teams use two triggers: a recurring incident category that exceeds a defined threshold within a rolling time window, or any major incident that caused significant business or user impact. The exact thresholds depend on your environment, but the trigger rules should be documented and applied consistently rather than left to individual judgment.
Who should own problem management in a service desk team?
Problem management works best when it has a dedicated owner — a senior analyst or problem manager — who is not responsible for the day-to-day incident queue. Incident pressure consistently crowds out investigation time when the same person owns both. In smaller teams, a scheduled protected block of time for problem work is the minimum viable alternative.
What is a known error record and how is it different from a problem record?
A problem record documents an investigation in progress. A known error record is created once the root cause has been identified but a permanent fix has not yet been deployed. It contains a tested workaround that agents can apply immediately when related incidents occur, reducing resolution time while the fix is being developed or scheduled.
How does problem management connect to change management?
When a root cause requires a change to infrastructure, configuration or software, the problem manager raises a formal change request. This ensures the fix goes through proper impact assessment and approval rather than being applied ad hoc. The problem record stays open until the change is deployed and verified to have resolved the underlying cause.
How often should open problem records be reviewed?
A monthly review of all active problem records is a common baseline. Records with no activity beyond a defined period — typically 30 days — should escalate to the problem manager or service desk manager. A quarterly review of the full problem backlog, including known errors awaiting a permanent fix, helps prevent chronic issues from being forgotten.
Key Takeaways
- Incident management restores service fast; problem management removes root causes permanently. Running only one practice means either chronic repeat incidents or slow recovery times.
- A clear trigger rule — based on recurrence thresholds or major incident impact — is what turns incident data into problem records reliably.
- Known error records are the bridge between investigation and resolution: they cut incident resolution time while a permanent fix is in progress.
- Problem management needs a dedicated owner and protected time. Incident queue pressure will always crowd it out otherwise.
- Asset data matters. Knowing which configuration items share a profile with the failing one is essential for scoping a fix correctly. Odysseus provides that live inventory data directly inside TIKTING, so problem investigation starts with facts rather than guesswork.


































































































