Back to Writing
AI & Transformation

A Library of Prompts Is Not a System

Two skills each called themselves "level three" and meant opposite things. The fix wasn't better prompts — it was a contract: a short, boring, non-negotiable declaration every component signs before it enters the library at all.

August 19, 2026
10 min read

This is the third and final part of a three-part series.

The moment our AI skill library almost died was a disagreement about the number three.

Two teams were presenting their skills at the same review. The first had built a signal-triage skill — it watches an account base for inflections and recommends a play. “It’s at level three,” the builder said. “It acts, then logs what it did.” The second team had built a proposal-drafting skill. “Also level three,” that builder said. “It drafts, and a human approves everything.”

Same number. Opposite meanings. One team’s three meant acts without asking. The other team’s three meant never acts at all. Each skill had invented its own autonomy scale — its own x-axis — privately, inside its own prompt. Reasonable on its own, incompatible together. Multiply that across a set of skills, and a whole organization of people setting their own trust levels per skill, and the math was obvious: we weren’t building a system. We had disparate prompts sharing loose context and varying levels of quality, all happening to live in the same folder.

The fix wasn’t better prompts. The fix was a contract — a short, boring, non-negotiable declaration that every component signs before it’s allowed into the library at all. This final piece is about that contract: what belongs inside a skill, what has to be pulled out of every skill and owned by the framework, and why the least glamorous artifact in the whole architecture may be the one doing the most work. Part 1 put autonomy and quality on two separate axes. Part 2 argued the judges deserve their own library. This part is about the standard that lets both libraries compound instead of sprawl.

The box that changed everything was boring

The most instructive precedent I know comes from shipping.

Before the 1960s, cargo moved as break-bulk — barrels, bales, crates, each shaped differently, each loaded by hand, each port a bespoke negotiation. Then came the container: a dumb steel box with standardized dimensions and corner fittings. Marc Levinson’s history of it, The Box, makes the point that transformed my thinking about skill libraries: the revolution was not the box. Anyone can build a box. The revolution was the standard — the agreement on exactly where the corner castings sit, so that any crane in any port on earth can lift any container off any ship onto any truck, without anyone asking what’s inside.

What’s inside the box is the shipper’s business. How the box interfaces with the world is the standard’s business. That split — contents free, interface fixed — is what let the system compound. Every new port, ship, and crane that adopted the standard made every existing container more valuable, and freight costs fell so far that the world economy reorganized around them.

That collision — two skills, two different autonomy scales — was a break-bulk problem. Each skill was a beautifully hand-shaped crate, and every user was a dockworker learning each crate’s quirks by hand. Skills, like containers, can vary all they want on the inside; what has to be standard is the corner casting. The contract is that casting: fix the interface, free the contents.

What lives where

Here is the split as we now enforce it.

The skill owns its craft. How to triage a signal, how to structure a proposal, what a good first draft looks like for its specific capability — the contents of the box. This is where the domain expertise lives, and the contract has no opinion about any of it.

The framework owns the semantics of trust. The maturity ladder — suggest, draft, act-with-review, act-and-log — is agreed once across the organization and owned by the framework, with one meaning. A skill does not get to invent its own. It declares what its behavior is at each rung of the shared ladder, the way a container declares its weight without redesigning the corner castings. When a person moves a skill from draft to act-with-review, that step means the identical thing on every skill they run — which is the only reason a non-expert can safely operate fifteen of them. The burden of knowing what to type never transfers to the user; the burden of conforming sits on the component, where it belongs.

The alternative designs both fail in instructive ways. One skill per level — fifteen skills times five levels — is seventy-five things to maintain and a migration every time the ladder changes. One skill with “just prompt it more precisely” puts the expertise burden on exactly the people who don’t have it yet on day one. One skill, one shared interface, framework-owned semantics is the only shape that scales down to a new hire and up to an organization.

Amazon ran this movie internally two decades ago with microservices and two-pizza teams. Small teams, each owning a service end to end, each free to build its internals however it chose — and required to connect to everything else through strongly defined APIs. No back doors, no shared databases, no side channels. It’s the same contents/interface split: full freedom inside the service, none over how it connects. That constraint is a large part of why the internals could later be recomposed into entirely new businesses — services built for internal teams could be offered to customers because the interfaces were already contracts. Constraint at the interface is what buys freedom everywhere else.

The contract, in five clauses

Every component in our library — skill or evaluator, because the judges from Part 2 sign the same contract — must declare five things before the review gate opens. Here is the shape of a conforming manifest:

# conformance contract v1 — no component enters the library without this
name: signal-triage
type: skill                          # skill | evaluation — same contract for both
facing: internal                     # internal | external — this one word sets the ceiling

# 1. Maturity: conform to the shared ladder — never invent your own
maturity:
  interface: framework/maturity-v1   # semantics owned by the framework
  behaviors:                         # what THIS skill does at each shared rung
    suggest:          "flag the inflection, name the likely play"
    draft:            "draft the recommended motion for approval"
    act_with_review:  "prepare and queue the actions; a human releases them"
    act_and_log:      "execute internal actions; log and surface the decision"

# 2. Telemetry: trust must be observable, or promotion is a guess
telemetry:
  emits: [approvals, edits, overrides, time_to_approval]

# 3. Ceiling: enforced by the framework, not remembered by people
#    facing: external hard-caps at act_with_review — no accrued evidence overrides it

# 4. Context: declare what you read and write — no undeclared reach
context:
  reads:  [account-memory, usage-signals]
  writes: [account-memory, activity-log]

# 5. Review: a named gate and a named owner before entry
review:
  gate: library-review
  owner: field-ops

Each clause exists because its absence produces a specific failure. Without clause one, you get the level-three collision. Without clause two, autonomy promotions run on vibes — the “twenty clean approvals” mechanism from Part 1 only works if approvals and edits are emitted and countable, so telemetry is a condition of entry, not a nice-to-have. Without clause three, the most important safety property in the system — customer-facing components can never fully auto-act, regardless of how much trust they’ve banked — depends on every builder remembering it, and a rule that lives in memory is a rule that’s already been broken somewhere. Without clause four, you can’t answer “what does this thing touch?” — which becomes the first question that matters the day something goes wrong. And without clause five, the library’s front door is open, and everything upstream of it erodes.

Notice what the contract never mentions: quality of craft, choice of wording, how the skill thinks. Contents free. Interface fixed.

The bureaucracy objection

The pushback I take most seriously: this is how governance kills momentum. Contracts grow clauses. Review gates grow queues. Somewhere out there is a version of this framework with a forty-field manifest and a committee, and nobody builds anything on it anymore.

I feel the pull personally, because the sixth clause is always reasonable. Someone proposes adding a cost declaration, a latency budget, a data-classification field — each defensible alone, each another pound on the front door. Our working rule is that a clause earns its place only if its absence produces a failure we have actually seen, not one we can imagine. Five clauses have cleared that bar so far. I’ve vetoed two more, and I’m not certain I was right both times. Keeping a standard minimal turns out to be harder than writing one — the container succeeded partly because the committee resisted almost everything except the corners.

The boring artifact is the moat

Levinson’s deeper point about the container was economic: once the interface standardized, the system — not any box, ship, or port — became the asset. Boxes were commodities the day they were standardized. The network they clicked into was priceless.

The same inversion is coming for AI operating models, and I think it’s the quiet answer to the question every leader building one should be asking: what, exactly, compounds here? Individual skills don’t — they have a lifecycle, created and evolved and eventually commoditized as the models underneath them improve whether you do anything or not. What compounds is the governed system: a shared trust ladder the whole organization conforms to, telemetry that makes promotion evidence-based, ceilings enforced structurally, dependencies declared, and one gate at the door — with a producing library and a judging library both signed onto the same contract. That’s the difference between an organization that owns an operating model and an organization that owns a folder of prompts.

The three verbs of this series come together in the contract. Skills produce. Evaluators judge. And the ladder they climb means one thing everywhere — because both libraries live under one standard that fits on a page.

One exercise to close the series. Open any two AI workflows your organization runs — the two you’d put on a slide. Ask one question of both: does “trusted to act on its own” mean the same thing in each? If yes, you have the beginnings of a system, and your job is to guard the standard. If no, you have break-bulk — beautifully crafted crates, loaded by hand, and a port that will never get cheaper to run.

Enjoyed This Article?

Subscribe to receive long-form essays on strategy, leadership, and systems thinking. Published twice a month with insights you won't find anywhere else.

Olawale Oladehin

About Olawale Oladehin

Olawale is a strategist, speaker, and thought leader who works with organizations to navigate complexity and build systems that create lasting value. He writes about strategy, leadership, and decision-making.

Learn more

Continue Reading