Most SLA breach reviews produce a cause that cannot be acted on. “The ticket was complex” and “we were short-staffed” are true, unfalsifiable, and lead nowhere. The fix is to stop asking what went wrong and start asking where the clock was running.
SLA breach root cause analysis works at two scales and they need different methods. For a single serious breach, walk the ticket's timeline and ask why at each handoff. For a period of breaches, classify every one by where the time was spent rather than what broke, then look at the distribution. One breach is an anecdote; thirty breaches sorted into six buckets is a diagnosis.
Classify by Where the Time Went
The usual taxonomy — people, process, technology — is easy to apply and almost useless afterwards, because every bucket contains several unrelated fixes. A time-based classification is harder to argue with and each bucket points at one owner:
| Where the clock ran | What it looks like in the ticket | The fix |
|---|---|---|
| Wrong target from the start | Priority or category set incorrectly at creation, so a target applied that never matched the work | Triage quality, intake form, or automatic classification |
| Sat unassigned | In a queue, nobody picked it up. Long gap between creation and first touch | Routing rules and queue ownership |
| Assigned but untouched | An owner existed and did nothing for hours. Usually workload, sometimes absence | Capacity, or reassignment on inactivity |
| Waiting on someone external | Requester, approver or vendor. Clock ran while nobody on your team could act | Pause conditions, and an underpinning contract that fits the SLA |
| Actively worked, genuinely slow | Continuous engineer activity, still overran | Skills, tooling, or accept that the target is wrong |
| Target was never achievable | Nobody could have met it — overnight, weekend, or a 30-minute target with no cover | Change the target or extend the cover |
The last two look similar and are opposites. One is a capability problem you can invest in; the other is a promise that should never have been made. Confusing them produces improvement plans aimed at teams that were never the constraint.
Split Response From Resolution First
Before classifying anything, separate which clock breached. A response breach and a resolution breach on the same ticket have almost no causes in common — the first is about attention, the second about capability. Aggregating them produces a cause list that is an average of two unrelated problems.
In practice this single split resolves a surprising share of reviews on its own: a queue that looks slow usually turns out to be fine at one clock and poor at the other. The detail is in first response time vs resolution time.
A Record That Survives the Meeting
Keep it short enough that someone will actually fill it in during a busy week.
Per-breach record
Ticket: [ID] - [PRIORITY] - [CATEGORY] Clock breached: [RESPONSE / RESOLUTION] Target: [TARGET] | Actual: [ACTUAL] | Over by: [DURATION] Where the time went: [ONE OF THE SIX BUCKETS] Evidence: [TIMESTAMP TO TIMESTAMP, FROM THE TICKET HISTORY] Was this preventable by us? [YES / NO / PARTLY] Action: [SPECIFIC CHANGE, OR "none - accepted"] Owner: [NAME] | Review by: [DATE]
Two fields do the work. Evidence forces a timestamp range from the ticket history rather than a recollection, which is what stops the analysis becoming a discussion of who was busy. “None — accepted” as a legitimate action prevents the ritual of inventing an improvement for every breach; some breaches are the cost of a target you have chosen to keep.
One Breach or Thirty?
The two scales genuinely need different tools.
A single severe breach — something customer-visible or contractually significant — deserves a timeline walk. List every state change with its timestamp, find the largest gap, and ask why at that gap rather than at the outcome. Five Whys works here, provided you point it at the gap and not at the breach itself; asked about the breach it reliably terminates at “we were busy”.
A period of breaches needs counting, not depth. Classify all of them into the six buckets and look at the shape. Any bucket above roughly a third of the total is your systemic problem, and the other five are noise you should deliberately ignore until it is fixed.
The mistake is doing deep analysis on every breach. It is expensive, it produces thirty unrelated actions, and none of them get done.
When the Cause Is the Target
Expect a meaningful share of breaches to land in the last bucket, and prepare for the discomfort of that.
If a target was agreed in a workshop, never revisited, and quietly cannot be met by the rota you actually staff, the honest root cause is the target. That finding is unpopular because it reads as excuse-making, so teams instead record “capacity” and commission an improvement plan that cannot work.
The test is arithmetic rather than opinion: take the target, take the hours you have cover, and check whether the promise is achievable in the window. Where suppliers are involved the same test extends down the chain — see SLA vs OLA vs underpinning contract for why a supplier's contracted response time can make your SLA impossible before anything goes wrong.
Closing the Loop
An RCA process that produces findings and no measurable change is worse than none, because it consumes the goodwill you need for the next one. Three habits keep it honest:
- Re-run the classification next period. The distribution moving is the only evidence your fix worked. A single review tells you what was wrong last month; the second one tells you whether you fixed it.
- Cap open actions. Three at a time, with owners and dates. Thirty findings and no capacity is how RCA becomes theatre.
- Report the accepted ones. Breaches you have consciously decided not to prevent should appear in the report as a number, so nobody later mistakes them for oversights.
The output that matters is not a cause list. It is a smaller number in the same bucket next quarter — which is also what makes the compliance report worth producing at all.
Where InfraCue Fits
InfraCue records status history on every ticket, so the evidence line in that template comes from timestamps rather than memory: when it was created, assigned, moved to a waiting state and resolved. SLA targets are set per priority and optionally per category and sub-category, counted in working time against a per-customer schedule, so the sixth bucket — a target nobody could meet overnight — largely stops occurring. Ticket reports can be filtered by SLA state and exported to CSV for period analysis, and the dashboard and NOC wallboard show breached and at-risk counts live. Pricing starts at $39 per month (₹2,999), with a 30-day trial and no credit card.
This guide is maintained by InfraCue, an IT operations platform with ticketing, SLA tracking and device monitoring in one console.