Guided Workflow Organizational diagnosis · Evidence discipline

Transfer Claim Test

How would you know if the replacement actually did the job?

Why this matters

You removed a review step, automated part of a workflow, or replaced work an expert used to do. You identified something important the old way was providing, and decided the new system needs to keep providing it.

Now comes the harder question: how would you actually know if it did?

That question looks simple. It's surprisingly easy to answer too confidently, because the easiest evidence to gather is silence, and silence is ambiguous. No reopened decisions, no escalations, no incidents: that could mean the replacement is working. It could also mean nothing has happened yet that would have tested it.

The deeper problem is that evidence arrives after you've already made the claim. That leaves room for the claim itself to shift. You might start with one idea of what counts as success, one idea of what a fair test looks like, and one idea of how much evidence is enough — then quietly adjust each of those as results come in. Each adjustment can sound reasonable on its own. Together, they can make it almost impossible for the evidence to count against the original claim.

Governing principle

Evidence can change what you claim next. It cannot change what you claimed before the evidence arrived.

This tool exists to prevent that drift. Before you watch anything happen, you write down exactly what you expect, what would prove you wrong, and how much evidence is enough either way. Then you let the results speak, keep an honest record of what they actually showed, and never quietly rewrite what you said before you knew the answer.

The Hidden Function Audit asks what must survive a change. This one asks how you'd know whether it did — and keeps that judgment accountable as more evidence comes in.

The diagnostic

Step 1

Say exactly what you expect the replacement to keep doing

Start with the thing the Hidden Function Audit told you the new system must not lose. That's the preservation requirement you're carrying forward. Now write one sentence: what's now supposed to perform that function, where is it supposed to work, and to what standard?

If there are different kinds of situations where this might work differently — routine cases versus rare ones, for instance — name them now. This matters more than it sounds: once you start watching results, you can't add a new situation to the list after the fact and pretend you always meant to test it separately.

We'll call this sentence your transfer claim. The worksheet also asks you to name any situations you want to test separately.

Output: your transfer claim, plus the situations you'll test separately.

Step 2

Decide what would support the claim or show it isn't working — before you look

Before you start watching results, decide what would actually test the claim. What kind of event would show it isn't working? What counts as a real test of it — and just as importantly, what looks relevant but doesn't actually test anything? How many real tests do you need before a quiet result starts to mean something? And how many failures would be enough to conclude the replacement isn't working, even before you've hit that number of tests?

Different situations may need different answers. A rare, high-stakes case probably shouldn't need as many failures to raise concern as a routine, low-stakes one.

Once you start watching results, changing these answers doesn't revise your test — it starts a new one. The old answer stays on the record.

The worksheet records these as: what counts as a real test, what doesn't count, how many real tests are enough to draw a conclusion, and how many failures are enough to say it's not working.

Output: a plan for what would support the claim or show that it fails, decided before you look.

Step 3

Keep a record of what actually happens

Keep what happens separate from what you decided in advance. For each relevant event, note whether it was actually a real test of the claim, which situation it belongs to, what happened, and whether it counts as evidence against the claim.

If something unexpected shows up — a pattern you didn't plan to test for — write it down, but don't treat it as proven yet. An unplanned pattern can become a new thing to test. It can't quietly become a confirmed result under the old plan.

Output: a running record of what's actually happened, kept separate from the plan.

Step 4

Decide what the evidence actually lets you say

Check each situation you're testing against the plan you made in Step 2. If different situations are turning out differently, don't force one overall verdict — each one gets its own answer, and the two thresholds from Step 2 don't run on the same clock:

How the verdict gets decided

Supported: you've run enough real tests, and you haven't hit the failure line. The evidence currently supports the claim for this situation — that's different from proving it works in general.
Not working: you've hit the failure line — even if you haven't run many tests yet.
Not enough evidence yet: neither bar has been reached.

Situation testedHow we knew to test itEnough real tests?Hit the failure line?What happenedWhat we can say
Routine casesPlanned in advanceYesNoNo failures observedSupported
Cross-functional casesPlanned in advanceYesYes1 failureNot working
Exception casesPlanned in advanceNoNo2 real tests; 2 failures, below the lineNot enough evidence yet

A failure stays on the record even when you don't yet have enough tests to draw a conclusion. What decides the verdict is whether you've crossed the failure line or run enough real tests — not just what you've happened to observe so far.

Output: a verdict for each situation you're testing — and only for the ones where a threshold has actually been crossed.

Step 5

Decide what to do next

If a situation is Supported, keep it as-is. If you still don't have enough evidence either way, keep watching.

If a situation isn't working, work through these questions in order:

  1. Is it only happening in one particular situation, while others you've tested enough are fine? If yes, narrow the claim to what's actually working. If you can't tell yet because you haven't tested that situation enough, leave it open rather than deciding.
  2. Can you trace the failure to something specific that was built or run wrong — something fixable? If yes, fix it and test again.
  3. If it was built and run exactly as intended and still failed, ask whether that changes what you believed about the original problem, back in the Hidden Function Audit. If it does, that's where you go next. If it doesn't, redesign your approach and test again.

Output: what to do next, and a record of how you got there.

Step 6

Keep the old claim on the record

Don't erase what you originally claimed just because the evidence changed what you can now say. Keep three things visible: what you claimed before you knew anything, what the evidence currently justifies saying, and any new claims — narrower scopes, new mechanisms, whatever came out of Step 5 — that still need their own test.

Evidence rule

Narrowing changes the claim going forward. It does not rewrite the claim that was tested.

Output: an honest record that can change without rewriting its own past.

Worksheet

The worksheet is built around four zones, because the risk in this tool isn't misunderstanding the sequence — it's quietly "clarifying" something that was supposed to be locked in. Each zone marks a different rule about what can change and when.

Locked before you start watching results

Sections 1–2. No edits once you start — a change here begins a new round of testing, it doesn't revise this one.

1. Your transfer claim

What's being preservedfrom the Hidden Function Audit
What has to stay truethe Hidden Function Audit's observable answer
What's doing the job now
Where this is supposed to hold
Situations to test separatelylist each
What standard it has to meet

Transfer claim: [what's doing the job now] will keep [what has to stay true] holding across [where], to [what standard].

2. Decide what would support the claim or show it's wrong

One card per situation you're testing separately — different situations may need different thresholds.

Situation: _______________

What would show this isn't working?

What counts as a real test?

What doesn't count as a real test?

Enough real tests to conclude anything
Enough failures to say it's not working

Added as you go

Section 3. Add rows as events happen — never revise Sections 1–2 to fit what's showing up here.

3. Keep a record of what happens

EventSituationReal test?What happenedEvidence against the claim?

Something unexpected? _______________
Treat it as: ☐ Just an observation   ☐ Worth testing on its own

Decided once the evidence allows

Sections 4–5. A situation can be settled as soon as its own plan's rule is met — running enough real tests and hitting the failure line aren't on the same clock.

4. What the evidence lets you say

Situation testedHow we knew to test itEnough real tests?Hit the failure line?What happenedWhat we can say
Planned in advance / Noticed along the waySupported / Not working / Not enough evidence yet

A situation you noticed but didn't plan to test can only be logged as an observation here — it can't earn a verdict under this test.

5. What to do next

Only situations that aren't working need a row here. Fill in only as far as the questions take you.

Situation that's not workingOnly this situation?Something specific and fixable?Does this challenge the original diagnosis?What to do
yes/no, or not enough evidence to tellyes/noyes/no

Kept on the record for good

Section 6. Only ever added to — nothing here gets overwritten, including a failed result.

6. Keep the old claim on the record

A version changes only when what you're claiming changes — a narrower scope, a new situation split off to test on its own, or a redesigned mechanism. Fixing something without changing the claim starts a new round of testing for that situation, not a new version. Watching longer under an unchanged plan stays in the same round.

VersionSituationTesting roundResultWhat changed / what's next
V1originalRound 1

Evidence rule

Narrowing changes the claim going forward. It does not rewrite the claim that was tested.

Worked example — Decision rights & coordination

Following on from the Hidden Function Audit's Decision rights & coordination example.

Step 1

What has to stay true (from the Hidden Function Audit): ambiguous or contested decisions reach a clear, final decision that teams know they can act on without reopening the issue or looking for someone else to overrule it — call this authoritative closure for short.

Transfer claim: a documented decision-rights matrix plus a designated escalation owner will provide equivalent authoritative closure for product decisions, across routine decisions, cross-functional disagreements, and exception cases.

Situations to test separately: routine, cross-functional, exception.

Step 2

What would show this isn't working: an in-scope decision is reopened after resolution, stalls on unclear authority, or triggers further escalation because a team doesn't accept the designated resolution as final.

What counts as a real test: only decisions that actually require interpretation or a competing claim of authority — routine decisions everyone already agrees on don't count.

Enough real tests to conclude anything (illustrative): at least six qualifying cases per situation, over one full planning cycle. Enough failures to say it's not working (illustrative): one qualifying reopening for routine and cross-functional cases; three qualifying reopenings for exception cases, since the organization expects more variation in genuinely novel decisions.

Step 3 — Round 1

12 qualifying routine cases, no reopenings. 7 cross-functional cases, 1 reopening traced to an outdated owner listed in the matrix. 2 exception cases, both reopened because neither team accepted the escalation owner's authority to settle them.

Step 4

Situation testedHow we knew to test itEnough real tests?Hit the failure line?What happenedWhat we can say
RoutinePlanned in advanceYes (12 of 6)NoNo failuresSupported
Cross-functionalPlanned in advanceYes (7 of 6)Yes (1 of 1)1 failureNot working
ExceptionPlanned in advanceNo (2 of 6)No (2 of 3)2 real tests, 2 failuresNot enough evidence yet

Step 5

Cross-functional — enough tests, and it's crossed its failure line. Traceable to a specific problem: an outdated name in the matrix. → Fix it, test again.

Exception — neither line crossed yet → keep watching; still not enough evidence either way.

Step 6 — Round 1

Original claim unchanged. Currently Supported: routine. Still to resolve: cross-functional (being fixed, will be retested), exception (still watching).

Round 2

Cross-functional retest passes with the corrected ownership — now Supported. Exception reaches 6 qualifying cases, 6 reopenings — both lines are now crossed for exception too, and every case failed. Step 5 runs again for exception: not localized to a sub-case, since exception is already the smallest unit being tested. Built and run exactly as intended — the escalation owner ruled on all six exactly as designed, so nothing was broken. Does this change what we believed in the Hidden Function Audit? Not yet — the failures are consistent with "authoritative closure" being the right thing to preserve, just not delivered by this particular mechanism. → Redesign, test again.

Step 6 — Updated record

A version changes only when what's being claimed changes. Fixing the matrix didn't change the claim, so cross-functional stays under V1 with a second testing round. Different situations were on different timelines after Round 1, so each is tracked separately rather than treating the whole claim as retested together.

VersionSituationTesting roundResultWhat changed / what's next
V1RoutineRound 1SupportedKeep as-is
V1Cross-functionalRound 1Not workingFix matrix entry, test again
V1ExceptionRound 1Not enough evidence yetKeep watching
V1Cross-functionalRound 2SupportedFix confirmed; keep as-is
V1ExceptionRound 1, continuedNot workingClose this portion; redesign creates V2
V2ExceptionRound 1TestingNew plan being tested

Final twist

V2 transfer claim (per Step 1–2): the redesigned exception mechanism will maintain authoritative closure for exception cases. Enough failures to say it's not working, same as before: 3 qualifying reopenings.

The redesigned mechanism successfully produces authoritative closure: teams agree who has final authority, the decision is explicitly recorded as final, and no additional authority is sought. But over the next planning cycle, 4 qualifying exception cases are reopened anyway — crossing V2's own failure line. V2 is not working.

Is it localized to a smaller situation? No — it's uniform, not tied to a sub-case. Is it something specific and fixable? No — the mechanism operated exactly as designed, producing and recording resolution every time.

A closer look at the four reopened cases shows they cluster around decisions where affected teams weren't involved before the decision, while equally ambiguous decisions with early stakeholder participation remain settled regardless of who makes the final call.

So: does this change what we believed in the Hidden Function Audit? This time, yes. V2 achieved the literal thing we were trying to preserve — authoritative closure — and the failure persisted anyway. That's real evidence against the original diagnosis, not just against this replacement: authoritative closure may not have been the condition preventing decisions from reopening. Shared context or participation before resolution may have been doing part of that work. → Return to Hidden Function Audit.

VersionSituationTesting roundResultWhat changed / what's next
V2ExceptionRound 1Not workingBuilt and run as intended; evidence challenges the original diagnosis → return to Hidden Function Audit

Closing the loop

Back in the Hidden Function Audit, the new evidence becomes input to a fresh diagnosis. It suggests that participation or shared context before resolution may have been part of what the old mechanism was providing — but this tool doesn't get to decide that. The Hidden Function Audit has to determine whether that function is real, what it actually protects, and whether the evidence is strong enough to make it something that must be preserved.

If that renewed diagnosis produces a revised answer, that answer becomes the input to a new transfer claim. The process then starts again: say exactly what you expect, decide what would prove you wrong before you look, watch what happens, and let the thresholds — not either tool's say-so — decide whether the new claim is Supported, not working, or still unresolved.

That's the full cycle the tool pairing is built to support. This tool can surface evidence that the original diagnosis was incomplete, but it can't rewrite that diagnosis itself. The Hidden Function Audit can revise what needs to be preserved, but it can't declare that a new replacement actually preserves it. Each tool hands the other exactly what it needs, and neither one gets to finish the other's job — including finishing this example with an ending the evidence hasn't earned yet.