Back to Writing
AI & Transformation

Proof of Life

A skill can pass every quality gate in your system and still be dead — never firing, or firing on the wrong work. Nothing inside the system can tell you. The evidence has to come from outside: evals, run at merge time, on every change.

September 5, 2026
10 min read

This is the fourth part of a series I published as three.

The skill was finished. It had a signed contract and a named owner. It had a judge that scored every draft before a human saw one, and it sat at the second rung of the autonomy ladder — drafts only, nothing acted on without a person in the loop. By the standard of the three pieces I had just published, there was nothing left to do to it.

Then I went looking for the evidence that it worked, and there wasn’t any.

Not evidence that the output was good — I had that, a judge produces it on every draft. Evidence that the skill ran. That when someone asked the question it was built for, this skill and not some adjacent one loaded into context and did the work. That when someone asked a question it had no business answering, it stayed quiet. That an agent equipped with it did better than an agent without it. I had built two libraries and a contract binding them, and I had never once checked whether the thing was alive.

Part 1 put autonomy and quality on two separate axes. Part 2 argued the judges belong in their own library, independent of the skills they officiate. Part 3 described the contract every component signs before it joins either library. All three are about what a skill is. None of them can tell you whether it does anything — and I want to be precise that this is a gap in what I built and shipped, not a gap I noticed in someone else’s work.

That check has to come from outside the skill. It has a name in software already, and once I borrowed it the whole problem reorganized itself.

Running and dead look identical

Anyone who has operated a distributed system has met this failure.

A process is up. The port answers. The health check returns 200. And the process is doing no work at all — wedged on a lock, spinning on a queue it will never drain, alive in every sense the machine can observe and useless in every sense that matters. The reason this is hard is structural: a process cannot certify its own liveness. Whatever is stuck is also what would have to notice it was stuck. The check has to be external, it has to be periodic, and it has to test for work actually happening rather than for the process still existing. That is why liveness probes are a separate mechanism in every orchestrator worth using, and why a probe that only checks “is the port open” is the classic way to run a dead service for a month without knowing.

An AI skill fails the same way, and the two things you would naturally reach for cannot catch it.

The harness cannot. It runs the skill; it does not vouch for it. Asking the harness whether a skill is working is asking the process whether it is stuck.

The judge cannot either, and this is the part that took me longest to accept, because it undercuts the piece I was proudest of. A judge scores a draft. It runs after the skill has produced something. If the skill never loads — because its description didn’t match how the person actually phrased the request — there is no draft. No draft means no judge. No judge means no score, no failure, no signal of any kind. The entire quality layer, the referee pool I spent a whole post arguing for, reports nothing at all. Silence from the judges is indistinguishable from everything being fine.

So the failure is invisible in exactly the place you’d look for it. In practice a large share of skill failures are trigger failures rather than output failures, and not one of them shows up in output judging. A skill can clear every quality bar in the system and still be dead.

What you need is a probe: something outside the skill, run on a schedule, that tests whether work is actually happening. For skills, that probe is a suite of evals, run at merge time, on every change. They are the skill’s proof of life.

Six moves toward a proof of life

1. Score the outcome; record the trigger separately.

The instinct is to make “did the skill load?” the pass condition. Resist it. If an agent completes the task correctly without ever loading your skill, that is a pass — and a useful signal that the skill may be doing less than you assume. Requiring the author’s expected path punishes a capable agent for finding a different successful one. Record whether it fired as a diagnostic, kept separate in the results, because a trigger failure and an output failure need completely different fixes and a blended accuracy number hides which one you have.

2. Treat the description as routing logic, not marketing copy.

The description is the highest-leverage field in the whole component and the easiest to leave ungoverned. It is paid for on every model call, whether or not the skill loads, and it is the only thing deciding whether it does. “Use for web development” pulls React guidance into Angular work. What it needs, and what mine did not have, is an explicit negative boundary — a plain statement of what the skill is not for. The dangerous cases are always the adjacent ones: tasks that share vocabulary with the real job and are genuinely out of scope.

3. Write the “should not fire” cases, and make them outnumber the others.

For every request the skill should handle, write the neighbor it should ignore. A sourcing skill should fire on add an open role for a solutions architect in Austin and stay quiet on prepare interview questions for a solutions architect candidate — nearly every content word shared, completely different jobs. A negative case with no vocabulary overlap tests nothing.

I am not first to this. Philipp Schmid of Google DeepMind makes the case directly in a talk called “Don’t Ship Skills Without Evals”: the description field is the trigger mechanism, negative tests are what catch the over-triggering a vague one causes, and ablation is how you decide when a skill should retire. I read it as a description of the gap I had just found in my own work.

I would have argued my policy checks were sound by reading them. Then the first test run caught a guardrail meant to keep protected attributes out of a candidate search happily matching the substring “age” inside the word “storage.” Every architecture doc in the corpus mentions storage. Review would not have found that; running it did, in about a second. Over-triggering is the failure we rarely test for, and unlike a bad draft it costs context on every single call.

4. Run the suite without the skill.

This is the honest test, and the uncomfortable one. Run the whole thing twice — once with the skill, once with it removed — and compare. If the delta is negligible, the skill is not earning its place in the context window, however well it reads. Until you have run this, the efficacy of your skill is unknown, which is a different state from acceptable and should be reported as such.

5. When the skill retires, keep the eval.

Capability skills — the ones teaching a model something it cannot yet do reliably — are temporary by design. Models improve, and last quarter’s scaffolding becomes this quarter’s dead weight. Preference skills, the ones encoding your workflow and house style, are durable, because no foundation model is going to learn your organization’s private conventions on its own. Knowing which kind you have is what lets you read a flat delta correctly: for a capability skill it means the model caught up, and for a preference skill it more likely means your cases stopped exercising the preference.

Either way, the eval stays. It still expresses behavior the organization needs, and if performance degrades later it is the thing that tells you.

6. Make the merge rule explicit.

A change must leave the deterministic checks green and must either improve the evaluated cases or add new ones. “This wording reads better” is a hypothesis, not evidence. That single rule is what converts skill maintenance from taste into engineering, and it is the reason the evals have to exist before the rewrite rather than after — evals written afterward get shaped to fit the skill you already wrote, which is the opposite of the point.

The judges need one too

The uncomfortable corollary is that a judge is a component like any other, which means it can be confidently dead in exactly the same way.

I built one to score search-query quality: a weighted rubric, four dimensions, a 7.0 pass bar. Before letting it gate anything I wrote five calibration examples and scored them by hand, spanning clear pass, clear fail, and the genuinely ambiguous middle. The point was to check that the judge tracked human judgment rather than merely scoring fluently.

Two things fell out, and neither was findable by reading the rubric.

Four of the five stated totals did not reconcile with the rubric’s own weights. I had written the scores, then written the total, and the arithmetic simply did not connect them — which meant my anchors were teaching the judge a target no human actually held.

The second was worse. One case scored 7.4, comfortably above the pass bar, while specifying no location and no seniority level at all. For a people search that supports no structured filters, a constraint absent from the prose is a constraint not applied, so that query would retrieve the wrong people entirely. The weighted average had quietly traded a fatal omission against strong scores elsewhere. The fix was structural: some dimensions have to gate rather than average, because a deficiency there makes the output wrong instead of merely worse.

Both problems came out of doing the sums, not from reviewing the rubric. A rubric that scores fluently without tracking human judgment is worse than no rubric, because it launders taste as measurement — and I would have shipped mine.

What I’m still working out

Two things here are bets rather than results, and I would rather name them than let them pass as findings.

The first is the retirement threshold. Our working rule is that an ablation delta at or below two percentage points means the skill is not measurably helping and should be considered for retirement. I chose that number because it is small enough to catch dead weight and large enough to survive noise across a few trials. I have no principled defense of it, and I would revise it against anyone’s real data.

The second is whether the merge rule holds under deadline. It is easy to honor when the change is substantive and genuinely annoying when you can see the improvement and the suite cannot. The failure mode I am watching for is the rule quietly becoming advisory — enforced on other people’s changes and waived on your own. I do not yet know whether a small team can hold that line for years rather than quarters.

The eval library outlives the skill library

The three verbs of this series were produce, judge, and climb. What I missed is that all three are claims, and a claim without evidence is just a well-organized assertion.

Skills have a lifecycle. They get created, they get better, and eventually the models underneath them make them redundant whether anyone attends to it or not. Judges last longer, because an encoded standard is genuinely yours. But the thing with the longest life is the one I built last and should have built first. An eval outlives the skill it was written for, survives its retirement, and keeps expressing what the organization needs long after the component that once satisfied it is gone. The skills produce, the judges score, and the evals are the only artifact that can tell you either of them is still alive.

The hard part, I am finding, is not writing the eval. It is admitting the skill needs one — which means admitting you shipped something you could not prove worked.

So, one exercise. Open a skill your team relies on, the one you would put on a slide. Find evidence it fired on the right work this month — not that its output was good, that it ran, on the tasks it was built for and not on their neighbors. Then ask the harder question: if it had stopped firing three weeks ago, what in your system would have told you? If the answer is a person who happened to notice, you don’t have a proof of life. You have a skill that is probably fine, and probably fine is the same shape as dead until something outside it checks.


The artifact: an eval spec

Each piece in this series ships one structural template. Part 2 shipped the evaluator; Part 3 the contract manifest. This one is the eval spec — the external probe, one file per skill, versioned beside the component it proves.

---
name: candidate-sourcing-eval
type: eval                                # first-class — a sibling to skills and judges
proves: candidate-sourcing                # the component this vouches for
owner: talent-ops                         # a person owns this evidence
harnesses: [claude-code, cursor]          # reliability is per-harness, not universal
trials_per_case: 3                        # agents are nondeterministic; one run proves little
last_run: 2026-09-02
---

## Should fire
Requests this skill exists to handle. Score on OUTCOME — whether the task
succeeded. Record trigger state separately, as a diagnostic.

- "Add an open role for a solutions architect in Austin, senior, hands-on."
    expect: a role config; names the level and the segment
- "Why did this candidate rank where they did?"
    expect: reads the stored score breakdown
    forbid:  hedging language — guessing is the failure this case exists to catch
- "Can we stop sourcing from companies where we've had bad hires?"
    expect: declines, and explains why a permanent employer screen is prohibited

## Should not fire — the adjacent neighbors
Must be at LEAST as numerous as the should-fire cases. The useful ones share
vocabulary with a real case and are genuinely out of scope. A negative case
with no lexical overlap tests nothing.

- "Prepare interview questions for a solutions architect candidate."
    # shares nearly every content word with case 1. The hardest negative here.
- "Draft an offer letter for a candidate we're hiring."
    # downstream of sourcing; shares "candidate"
- "What's the difference between a senior and a principal engineer?"
    # this skill holds a level mapping, which makes it a real over-trigger risk

## Ablation
The honest test: is the skill earning its place in the context window?

run: whole suite, twice — with the skill present, and with it removed
compare: outcome accuracy, not trigger accuracy
cadence: on every model change, and quarterly
retire_below: +2 percentage points        # a working threshold, not a proven one
on_retire: DELETE THE SKILL. KEEP THIS FILE.
           It still expresses behavior the organization needs, and it is what
           will tell you if performance later degrades.

## Merge rule (policy, not preference)
A change to the skill, its config, or its scripts must:
  1. leave the deterministic checks green, AND
  2. improve the evaluated cases OR add new ones.

"This wording reads better" is a hypothesis. It is not evidence.

Notice what the spec does not contain: any judgment about whether the output is good. That belongs to the evaluator from Part 2. This file answers one question only — is the thing alive — and it answers it from outside, which is the only place the answer can come from.

Enjoyed This Article?

Subscribe to receive long-form essays on strategy, leadership, and systems thinking. Published twice a month with insights you won't find anywhere else.

Olawale Oladehin

About Olawale Oladehin

Olawale is a strategist, speaker, and thought leader who works with organizations to navigate complexity and build systems that create lasting value. He writes about strategy, leadership, and decision-making.

Learn more

Continue Reading