Governing the Corrective Action Backlog
A backlog nobody ages gets tidied by an audit date rather than by a decision. Governing corrective actions is unglamorous work: named people, two clocks of which only one can be reset, four numbers, and the nerve to refuse a closure with no evidence behind it.
Series: The Improvement Engine, Post #5
Governing the Corrective Action Backlog
I was handed an improvement register that had been running for four years. Several hundred open corrective actions, the oldest raised before most of the current management team joined, many already quietly done and never recorded, and an audit close enough to cause anxiety.
That register was not governance. It was a compost heap with a change history, and its problem was never its length. Almost nothing in it was true. I use “register” for the whole list and “corrective action backlog” for the part with a cause, an owner and a date. Most organisations have only the first.
The Asymmetry Nobody Prices
Raising an action costs a sentence, closing one costs a fortnight
Raising an action takes thirty seconds and earns a reputation for rigour. Completing one costs a fortnight of an engineer’s capacity, and almost no register accounts for the difference. In the continual improvement model “take action” is a single step among seven, which rather flatters it: that one step is where the whole capacity argument lives, and a backlog is what accumulates while the argument goes unhad.
Plenty of corrective actions are a configuration change and an afternoon, which is why you never see them: cheap ones get done, and being done removes them from the list. What survives is the tail, so the mean cost of an action is a figure of no interest — a backlog is by construction made of the expensive ones. The three-way sort belongs here as costing rather than filing: one pile carries the fortnight.
- A backlog with no service rate is not a backlog — it is an arrivals log, and status reporting will not convert one into the other.
- Closure needs help, not chasing — the constraint is engineering time, so what works is a reserved allocation, not another status call.
- Only one pile costs a fortnight an item — which makes classification a cost control rather than a filing preference.
In my opinion this asymmetry stalls more registers than any shortage of process. The cure is not tidier reporting, it is funding the closing end, since the raising end has never needed help.
Every Reassignment Resets the Clock
Handover is treated as tidying when it is a governance event
Capacity is the constraint, but it is not what turns a shortfall into a four-year-old record. That is ownership decay. An action starts with someone who agreed to it; they change team or hand over a queue, and the record acquires a new name and nothing else.
Nobody logs a reassignment, because it looks like housekeeping rather than a change of commitment. The incoming owner never agreed the date, was not in the room, and can decline at no cost. Three or four of those is how an action reaches its second birthday without anybody behaving badly.
- An owner who did not agree the date will not defend it — acceptance is a conversation, and only one held out loud survives a bad quarter.
- Reassignment requires re-acceptance — a new name means a new conversation, a restated date and a recorded reason.
- Count owners, not months — how many people an action has belonged to predicts failure better than its age does.
The only version I have seen hold is the handover conducted in front of somebody else: the outgoing owner states the condition of the work, the incoming owner states what they will do and by when. Anything less is a name change with a timestamp.
Ageing Rules That Do Not Ask Permission
Escalation should be a property of the system, not an act of courage
Two clocks are printed on every open action and only one can be reset. Age runs from the date it was raised; overdue runs from the current due date, which is why moving that date repairs it. A third clock runs that no register prints: tenure, how long the work has sat with its current owner, which every handover silently returns to zero. Escalate on age; report all three.
Left discretionary, escalation falls to whoever has the least influence to lose, so make it mechanical. Thirty days past the agreed date notifies the line manager, sixty puts the action on the service review agenda by name, and ninety days of age, whatever the due date says, requires written justification from the accountable director.
- Trigger on the third owner as well as the clock — an action that has outlived two handovers needs the room whatever its age band says.
- Escalate to whoever can reprioritise the owner’s queue — not whoever sits above them, since notifying somebody with no lever is correspondence.
- Make date changes visible — one recorded extension is fine, three silent ones are how a backlog reports itself green.
The thresholds belong to assurance, not delivery, because the people a rule names must not be the people able to amend it. Delivery will call that unrealistic, which is fair and beside the point: a threshold its named party can move is a preference, not a control. Automate it, because whoever escalates pays the whole social cost while the organisation collects the whole benefit: urging people to escalate never works, and a rule that fires by itself does.
Three Counts and a Ratio
Enough to govern a backlog, and far too boring to narrate
The instruments are unremarkable and the data already sits in your register. The three counts are open by age band, completions per period, and reopened. The ratio is completions against intake. Those four govern a backlog better than any suite sold for the purpose.
None of it is original, just queue arithmetic from flow management. Divide the open count by monthly throughput and you have months to clear: two hundred open against eight closed a month is over two years, and raising fourteen a month means it never clears. Throughput is the service rate the first section said most registers lack.
- Open count by age band — the distribution, not the total: forty actions none older than a quarter beats twelve that are two years old.
- Completions against intake — the ratio, because anything below one means the arithmetic is deciding your future rather than you are.
- Reopened rate — the cheapest test of whether “closed” means anything, though not of effectiveness.
What none of them tells you is when you fail an audit: one agreed action never implemented is enough. They govern flow, not exposure. Actions agreed with an auditor, a regulator or a customer are a named subset, run item by item alongside the arithmetic rather than inside it. The failure mode of good reporting is a green dashboard beside the one open item that costs you the certificate.
Out of the Sidecar and Into the Room
Improvement dies in meetings that only improvement people attend
Improvement fails in two rooms, for opposite reasons. The room it owns is the monthly continual improvement forum, chaired by whoever owns quality and attended by four loyal souls: an immaculate record, and authority over nothing, since the engineers it needs are allocated elsewhere by people not in it.
The room it visits has the opposite fault. Everyone with authority is in the post-incident review, the resolve is genuine, and nothing produced there outlives the week. Unless somebody turns the findings into owned, dated entries before it empties, they live only in the minutes. One venue keeps the record and has no power; the other has the power and keeps no record.
- Borrow an agenda people already attend — five minutes at a meeting senior people turn up to beats an hour at one they avoid.
- Same queue as incidents and changes — improvement competes for that capacity, so make it visible where capacity is allocated.
- Post-incident actions reach the queue first — one recorded in the minutes and nowhere else is already lost.
Improvement needs standing on an agenda that already holds the ranking and the capacity. I disagree, fairly strongly, that it warrants a governance structure of its own: give it one and you solve the record problem while making the power problem permanent, because a separate forum is not protection, it is quarantine.
The Pre-Audit Purge
Mass closure before an audit is misrepresentation with better formatting
Three weeks before an audit, somebody proposes a clean-up, and a block of actions is closed at speed as stale, superseded, or covered by another initiative. Be fair about it: this is nearly always proposed in good faith. If the open count is the number under scrutiny, the purge is not a lapse but the behaviour the reporting design paid for, and nobody asks whether the cause was addressed.
It is still not housekeeping. Closing an action because its age is embarrassing rather than because the cause was dealt with puts a false statement about the control environment into the record. It is not fraud in any sense a lawyer would recognise, but the effect on everyone who later relies on that record is the same.
- Closure needs evidence, not a comment — “no longer required” is a claim; the change record at least proves the work was done.
- Withdraw with a stated reason — an obsolete action can leave honestly, provided the reason and whoever decided it stay attached.
- Watch for clustering — an unusual closure spike deserves a sample check, and everybody should know it happens.
So purge-resistance is a reporting choice made months earlier, not nerve in the week before the audit. A raw open count rewards any closure, so it funds the purge; completions against intake and the reopened rate need closures to be real, since a withdrawal is not a completion. Eighty honest, ageing, uncomfortable actions beat a tidy twelve curated for presentation. The first is a problem you can work on. The second is one you have agreed to stop looking at.
Where I Would Start
Five moves before the next service review, and the order matters
If your register has stopped moving, none of the remedies below are difficult, merely unpopular for a month or two. Reclassify before you reserve: you cannot size the capacity until you know what is in the backlog.
- Reclassify into three piles — corrections already done, which close on evidence; opportunities, which leave the backlog; and corrective action, your real number.
- Reserve the hours for closing — even half a day a fortnight per team, because ageing rules on unfunded work produce a well-documented graveyard.
- Re-own your twenty oldest, out loud — a named person, a date they chose, a count of previous owners; leave the rest unowned, since a blank beats a team name.
- Publish the four numbers — open by age band, completions per period, reopened rate and the intake ratio, on one slide.
- Hand the thresholds to assurance — thirty, sixty and ninety days plus the third owner, into an existing agenda rather than an inbox.
The most useful hour I have spent on a stalled backlog went on counting the owners of the ten oldest actions, and reading the dates they were raised rather than the dates now due. Neither figure had ever been reported.
The purpose of governing a backlog is to make the register true, not to make it small, and the two pull against each other. Every mechanism above adds to the visible count or holds items there longer: an unresettable age, a re-accepted handover, a closure obliged to show its work. A manager who judges the governance by whether the number came down has already bought the purge. Closing an action and proving it worked are not the same thing, and that deserves a conversation of its own.
Hopefully this has been useful to you and I wish you well on your ITSM journey…