Seven plays that take one agentic workflow from a working demo to a system that survives customers, auditors, and on call. Prompts are persuasion. Harnesses are enforcement.
Bring the seven plays, the 12-row scorecard, and the fifteen-minute audit into the room. Complete the form and the download appears immediately.
Your request is sent to Spillwave at [email protected]. The download link is shown only after the form is submitted.
Decide, act, observe, repeat until a goal or a stop.Can we name the stop?
An open loop. The model picks the next step.Do we need an open loop, or a route?
A fixed route. The engineer draws the path.Can we draw the path on one page?
The governed runtime around either shape.What is true even if the model argues?
You own an agent, an agentic workflow, or a coding-agent rollout, and the room already recognizes at least one of these.
Do not run this playbook if you only want a bigger model. The model is not the missing piece. Audience: a VP, a director, an engineering lead, and the operator who will live with the loop.
The harness owns these six. Put them in the runtime, not in the prompt. Treat it as a cockpit, not a cage: the measure is larger, longer, more autonomous work completed, not simply fewer crashes. Your framework hands you a loop. It does not hand you a harness.
| Job | Buyer risk if missing | What you hold if present |
|---|---|---|
| Context assembly | The window is not full and the agent still forgets. | The model sees the right slice. Large output goes to a file. A short summary comes back. |
| Tool contracts | Wrong tool, wrong argument, silent side effect. | Each tool says when not to use it. Policy lives in the contract. |
| Memory | Storage is not memory. The fact was in the store. The agent still bought the shrimp. | Tiers of memory. Durable state across sessions. Retrieval you can audit. |
| Observability | Green dashboards. Customers never booked. | Outcomes and drift, not only uptime. Claims are not evidence. Files are. |
| Recovery | Naive retry books the same flight twice. | Idempotency, checkpoints, last-known-good, replay. |
| Orchestration | Split the work and pay 15 times the tokens. | Scoped subagents. Human gates on irreversible actions. |
Figure 1One pass of a governed loop. The runtime counts the retries. The model does not.
Princeton held the model fixed in 2024 and changed only the interface. SWE-bench resolved issues rose from 3.8 percent to 12.47 percent. Same model. The cockpit moved the number.
Public incidents map to missing jobs, not missing models.
Do not start with a platform rewrite. Start with one loop. Every play names when to run it, who must be in the room, the work in order, the artifact you hold after Spillwave leaves, a check a skeptic can run, and what changes for the business.
The owner of the workflow. One engineer who has been babysitting it. One person who feels the cost or the risk.
A scored card and three named gaps.
Anyone in the room can point at the three gaps without a slide.
You stop arguing about models. You have a shared list of missing controls.
Product owner. Engineering lead. The person who will approve irreversible actions.
A one-page route or a written open-loop charter, plus named stops.
A new hire can tell agent from workflow in one sentence for this job.
You buy the simpler shape when it is enough. You do not pay open-loop cost for a fixed route.
Engineering lead. Security or risk if the loop can spend money, change data, or talk to customers.
Put the six jobs above in the runtime, not in the prompt: context assembly, tool contracts, memory, observability, recovery, orchestration.
A harness around the chosen workflow. Stops and budgets the model cannot talk past.
An irreversible action is blocked when the gate is off, even if the prompt says yes.
You can operate the loop. You are not hoping the prompt behaves.
The person who currently holds the procedure. The lead who must scale it.
An agent is the worker. A skill is the playbook that worker loads when the task starts. A skill is not a longer system prompt. It loads only when the task needs it, which keeps the window small.
Versioned skills and plugins for the workflow. Inventory the org can count, not heroics.
A second person runs the same skill and gets the same kind of result.
Know-how survives vacation, attrition, and a new model.
Engineering lead. Whoever owns quality or compliance.
Maker and checker are doctrine, not job titles. Self-check without an external signal can make results worse. A wish is an acceptance criterion that cannot fail a test. If the new tests are not red before the code exists, the loop stops.
A maker path, a checker path, and a stop the maker cannot lift.
The checker has no write path to the work it grades, and a false "done" does not ship.
Quality is a control. It is not a vibe.
The workflow owner. A knowledge owner if you have one. Rick and Chris on the graph design.
Stand up the stack for this workflow only. Do not boil the company ocean. Self-improving loops write back into it: a judged finding becomes a claim, a closed ticket becomes a decision, and next week's agent loads that instead of a prompt you typed again.
A small graph for the workflow, with owners, and an agent that can retrieve it.
A new session answers "why did we decide this" from the graph, not from Slack.
The org compounds. The agent does not reset to zero.
Engineering lead. On-call owner. Risk if the loop can act in production.
A scheduled or event-driven loop with the same gates as the attended run.
You can point at a night or weekend run that stopped correctly, with a receipt.
Autonomy is a decision with a perimeter. It is not a default.
One workflow, twelve rows. Score 0 if the control is absent, 1 if partial, 2 if it is enforced in the runtime rather than requested in a prompt. Your total picks the play that runs first.
Score the twelve rows and the playbook tells you where to start.
Both windows are scoped to the one workflow you picked in Play 1. The plays run in order, and Play 7 waits until the attended loop is honest.
Figure 2The typical sequence. One workflow the whole way through, never the platform.
Bring the scorecard. Bring the workflow. The clock is the agenda.
| Minutes | What happens |
|---|---|
| 0 to 2 | Agree to use this playbook. Pick the workflow. |
| 2 to 6 | Name the three lowest rows. What breaks. Who complains. |
| 6 to 10 | Impact in ranges. Incidents per month. Hours of babysitting. Revenue or brand at risk. |
| 10 to 13 | Which play runs first. What 30 days can change. |
| 13 to 15 | If there is fit, book a 45-minute working session on that workflow. |
Questions we will ask: how many people touch the agent in production, what happens if this is still true in six months, who else must see a lift before you act, and whether a prompt-only fix has already been tried.
Spillwave helps leaders move past experiments and install real AI capability: harnesses, skills, knowledge, and training. Rick designs the harness, the skills, and the knowledge graph. Chris co-authors the operating model with your team in the field. They have worked together since 2003. They do not throw a deck over the wall.
Claude Certified Architect (CCA-F), member of the Anthropic Partner Network, and co-author of Manning's Harness Engineering for Production AI Agents. He built Skilz and helps run a large agent-skill marketplace, and teaches loop engineering and second-brain design. Earlier he shipped production systems as an approved vendor at Apple and the NFL, and as Senior Distinguished Engineer at Capital One on transaction intelligence and model governance. He has sourced and delivered enterprise work himself since 2003.
Rick's long-running Spillwave business partner and co-author of the Manning book. Rick hired Chris as a trainer in 2003, and they have designed and delivered together since. Chris is the field pair on the loops.
We do not sell a model. We sell the cockpit.
That row is the conversation. Bring your total and the row number, and we will run the fifteen minutes on one workflow. On LinkedIn, comment or DM PLAYBOOK or HARNESS for this file, and reply AUDIT with your score.