Proving It Worked
Closing a corrective action proves that somebody did something. Proving that the cause has stopped producing the effect is a separate act, performed months later by somebody else, and it is the step almost every improvement register leaves out.
Series: The Improvement Engine, Post #6
Proving It Worked
The most useful thing I have watched an improvement forum do was take delivery of a negative result. An action closed four months earlier, well written and genuinely done, had at last been revisited, and the fault it was meant to remove had recurred twice inside the window. The room’s first instinct was to argue with the window.
Nobody there was being dishonest, only unaccustomed to being told, because almost every register I am handed closes on the day the change is made and then proves nothing else.
Implemented Is Not Effective
The tidy register is the more dangerous of the two
Start with the hazard nobody expects. A visibly stuck register will eventually irritate somebody senior into intervening, which is uncomfortable but productive. One in which everything closes on time attracts no help, because nobody asks a second question of good news.
The frameworks always asked the question. In the ITIL 4 continual improvement model it is step six, and admirably blunt — did we get there? The Plan-Do-Check-Act lineage sits a generation back, with the v3-heritage 7-Step Improvement Process. Neither is missing a step. What the question lacks is an artefact: nothing specifies what “there” would have looked like, which is why it can be answered sincerely, at length, and without evidence.
- A clean closure record defunds the check — a register reported upwards as complete reads as proof that verification is redundant, and buys the withdrawal of the very attention that would have funded looking.
- The failure mode is silent — nothing alerts you when a closed action fails, because by then nobody is watching the symptom.
- Recurrence arrives disguised as new work — a fresh reference number, a fresh investigation, and the same analysis bought twice.
The commonest defect I meet in improvement governance is also the least discussed, because from any distance it looks like competence.
Effectiveness Criteria Are Not Acceptance Criteria
Two tests, two dates, two people, and one field where most registers keep both
Acceptance criteria answer whether the change was built as specified, checked at delivery by whoever built it. Effectiveness criteria answer whether the failure stopped, checked months later by somebody else. Most registers give the two one field between them, so the first closes the record and the second never happens.
Criteria invented at closure are worse, because you settle for whatever evidence is to hand. Written in advance, the check runs to three lines. The observable: the nightly billing extract does not abort on a duplicate record. The window: two month-end cycles. The verifier: the service owner, running the baseline query. Close on implementation, but raise the check as a separate, linked, dated record, so the obligation outlives the closure.
- Name the observable — the specific failure mode this action is meant to stop, not an aspiration towards better service.
- Fix the window before anybody knows the answer — long enough to contain the interval between the failures you have seen.
- Nominate a verifier who did not choose the remedy — marking your own homework is optimism with a signature on it.
The strongest objection is not bureaucracy. It is that ninety days on, the service has been re-platformed, the observable describes nothing that exists, and the check is unanswerable, so the whole apparatus was theatre. The expensive part, though, was never the check: it was deciding what would count, free when the action is raised and impossible afterwards — cheap enough that I verify in full only what follows a major incident, an audit finding or a repeat failure, and sample the rest. An observable overtaken by a re-platform is itself a finding: the cause was removed by demolition, not by your action, so you do not get to bank it.
What Actually Counts as Evidence
A screenshot of a configuration change proves that the configuration changed
The change record, the screenshot and the closure comment attest to one fact: somebody did what they said, which is not the same as it working. For a frequent fault you can count before and after; the counting is the evidence. For a rare one you cannot, because nobody proves the absence of an event rarer than their own patience. There the honest move is to verify the control: fire the alert deliberately, or submit the record that should now be rejected and watch it rejected.
The awkward part is attribution. The hardware refresh nobody mentioned will happily take credit for your process fix. Two defences, one of them free: hold the result against the effect you predicted, so an outsized win reads as a flag rather than a triumph; and where a comparable service went untreated, check whether the fault moved there too.
- Recurrence against a baseline — the named fault counted in comparable windows, on rules nobody changed in between.
- Demonstrated behaviour of the control — for faults too rare to count, prove the mechanism does what it was built to do.
- Rule out the coincident change — if the platform team touched the same estate inside your window, the result is not yours.
A removed cause produces silence; so does a fault that has not come round yet. Only the window you chose in advance tells you which of the two you are holding.
What an Auditor Actually Asks For
Nobody is testing whether the action worked, only whether you can show it
Ask for evidence that a corrective action was effective and what arrives is a change record and a closure comment: paperwork about the work, not the outcome. The obligation is to act on the cause so that the nonconformity does not, in the standards’ own words, recur or occur elsewhere. Nobody reads that last phrase slowly. It makes verification a question about two places. Both ISO/IEC 20000-1:2018 and ISO 9001:2015 then require a review of the effectiveness of the corrective action, and that review is what almost nobody can produce.
The duty is narrower than people fear. It bites on corrective action raised against a nonconformity, not on every good idea typed into a register, and satisfying it takes a dated note of who looked, at what, and what they found. Few hold that note, not because writing it is expensive but because nobody was told to look.
- The requirement is the review, not the verdict — an action that failed and was honestly reviewed stands better than one that quietly worked.
- Two places, not one — the check has to reach the adjacent service built the same way, because that is where the cause surfaces next.
- Evidence must outlive its author — attached to the action itself and findable by a stranger who was never in the room.
One caveat, before somebody quotes me at an assessment. The standards require the review and the evidence; the fields, the window and the named verifier are mine, defended on practical grounds rather than clause numbers. My experience has been that audit is a poor motivation for good practice and an excellent detector of its absence: the finding is rarely that an action failed, but that nobody can tell.
When the Honest Answer Is No
A failed check is worth more than a passed one
A negative result is the only one that sends anybody back to the causal analysis. There are three explanations, and they differ wildly in what they cost to rule out. The change may never have been adopted — the new alert routes to a mailbox nobody reads. It may have addressed one condition out of several. Or the diagnosis was wrong. Most people jump to the third and reach for a larger remedy when the first can be settled in an afternoon.
I have watched a fix applied faithfully to an interface when the condition producing the failure lived in the credential lifecycle behind it. The interface failed again on the old schedule. What matters then is that the thread survives. A fresh reference number discards the attempt, the evidence and the reasoning behind it, so reopen the original where you can, and where you cannot, let the principle outrank the mechanism.
- Test adoption before the diagnosis — the cheapest explanation for a flat result is a change that never reached the people it was written for.
- Keep the thread, whatever your tool allows — reopen if you can, or link the successor and record the failed check against the original.
- Record partial success honestly — “reduced but not eliminated” is a legitimate outcome that a binary closed state destroys.
A register in which everything worked is not an improvement mechanism but a marketing document, and everybody reading it knows as much. I would sooner present one with three failed checks.
Verification Is How You Learn to Predict
The gap between claimed and delivered benefit is the only honest input to the next decision
Every improvement decision is a forecast: this change, on this cause, will produce roughly this effect. Organisations make that forecast constantly, defend it energetically to win capacity, and never score it. Scoring it privately is a hobby. Calibration is worth nothing until it costs somebody something, so my version has a price: after a dozen verified outcomes, the ratio of delivered to claimed benefit becomes a discount on whoever proposes the next action, said out loud in front of whoever decides next time.
Two objections, both fair. It punishes the ambitious estimator, who guessed high because the problem deserved it, and it teaches everybody to sandbag. My answer to both is that the discount is already applied. Whoever decides next time is running one from memory, on a recollection of who disappointed them two years ago, under no obligation to admit it. An open ratio is cheaper, because it can be contested and because it moves: deliver accurately twice and it follows you up.
- Attach the discount to the proposer, not the proposal — a delivered-to-claimed ratio is only calibration if it changes what the next claim is worth.
- Store the prediction beside the result — the same record, both figures, and whoever made the call, so the comparison exists without anybody assembling it.
- Publish outcomes, not activity — a short list of faults that demonstrably stopped happening beats any chart of closures.
No organisation I have worked in could tell me how close its improvement estimates ran to its results, which is a fair part of why they never got closer.
Where I Would Start
What to add to a register that currently closes on implementation
The expense here is not tooling. It is the ninety days between closing an action and being allowed to believe it.
- Add three fields to the action record — observable, window and verifier, all required before an action enters the backlog.
- Raise the verification as its own dated record — linked to the closed action, with its own owner and date.
- Sample five closed actions, starting with your own team’s — you are testing whether the register answers the question, not whether anybody did their job.
- Diary the check so somebody is told — if the platform cannot schedule against the record, raise a dated task at the moment of closure.
- Where nobody independent exists, fix the query in advance — a verifier who chose the remedy can be honest if the answer does not turn on judgement.
Verification is the difference between a mechanism and a filing system. Ideas were never the scarce thing in improvement work. Proof was, and a closed record has stood in for it throughout my career.
The register stops being a list of things done and becomes a catalogue of faults that no longer occur, evidence attached. The forum I opened with did not settle because somebody argued better, but because the window had been fixed before anybody knew the answer, so there was nothing left to argue about. It produces no dashboard anybody wants on a wall. It does mean you stop paying twice for the same fix, which is an agreeable place to finish.
Hopefully this has been useful to you and I wish you well on your ITSM journey…