Books / Agents Decide, Tools Do, Gates Prove

Cover of Agents Decide, Tools Do, Gates Prove: A Field Guide to Building Agentic Applications by Walter Meyer

Agents Decide, Tools Do, Gates Prove

A Field Guide to Building Agentic Applications

By Walter Meyer · ThatWaltGuy · First edition, 2026

Every few weeks the focus changes: a new model, a new framework, a new acronym. Teams stop mid-build to rebuild their demos. The demos get better every cycle. The production numbers stay where they were. Then somebody asks the question that matters: how do you know?

Agents Decide, Tools Do, Gates Prove is a field guide to building agentic applications that survive that question. The method fits in six words. Agents handle judgment. Tools handle repetition, as plain code with no model inside. Gates handle trust, and every gate writes an artifact someone else can examine. Build that way and three properties show up without being retrofitted: the system scales, average unit cost falls, and the audit trail starts with the first run.

No models are named. No vendors. No frameworks. That is the method, not caution: a system designed around whatever shipped this afternoon inherits its shelf life. Five parts walk the work in order (decide what to build, design the solution, build the machine, prove it works, run it) and four appendices carry the instruments: a fit test card, a scoring sheet, a gate checklist, and a complete worked example of an invoice chain built on the method.

Written for practitioners who own a build: architects, senior engineers, engineering leaders, technical product owners, and the consultants who have to stand behind a number in front of a client or an auditor. If you are the executive funding one of these builds, this is what your team should be able to prove to you.

"A team can adopt every release for two straight years and end exactly where it started, with a pile of impressive demos that didn't move the needle in any measurable way. Other than cost. That one definitely moved and someone wants answers."

Key takeaways

  • 01 Agents decide, tools do, gates prove. Judgment goes to models, repetition goes to plain code, and every claim of correctness gets a gate that writes an artifact.
  • 02 Three questions decide whether work qualifies: can a machine check it, could someone who did not do it judge it, and can "done" be documented before it starts? A no is a finding worth as much as a yes.
  • 03 The highest use of an agent is rarely doing the work. It is building and maintaining the machine that does the work, then judging the output. Put the model in the workshop, not in the loop.
  • 04 Fix the generator, never the output. A hand-patched result is one the pipeline cannot reproduce, and the record now claims otherwise.
  • 05 Nobody grades their own work. Separate the builder from the judge and enforce it in the dispatcher, not by politeness.
  • 06 Verify on three independent surfaces: does the data exist, is the output right, and can the person who inherits it work with it. An unchanged save changes no business meaning.
  • 07 Total spend is a headline. Cost per accepted outcome, against a declared comparator, is the diagnosis.

Reader questions

Which models, vendors, or frameworks does it cover?

None, on purpose. A system designed around a particular model or framework inherits its shelf life. The book designs around the shape of the work, keeps every model swappable behind a capability tier, and treats each swap as a revalidation, not a redesign. Everything in it applies whatever you are running this afternoon.

Is this a book about AI agents in general?

No. The focus is bounded, repeatable workflows whose outcomes have to be evaluated and defended: invoice intake, catalog and content migrations, recurring compliance reporting, translation review, test evidence. Novel strategy, one-off decisions, and relationship work stay with people, and the book says so.

Do I need to be an engineer to read it?

It is written for people who own a build, and it assumes you have seen one go wrong. The examples are invented and described in plain language, and the instruments are cards, sheets, and checklists rather than code. An executive funding an agentic build will find the list of what the team should be able to prove.

What is the invoice chain?

An invented running example: accounts payable intake at a mid-size distributor, four thousand invoices a month, a team of four. It enters at the fit test in chapter 2 and runs through the complete worked example in Appendix D, so every rule in the book gets applied to the same case. The distributor, its invoices, and its people are imagined. The failure modes are not.

Why "gates prove" rather than "tests pass"?

A gate is a check with authority: work does not proceed until it passes, and every run writes a record of what was checked, against what criteria, with what result. Skip is not pass. The record says what was not measured. That is what turns "trust me" into "here is the record," and what makes a system auditable by construction instead of by reconstruction.

How does it relate to Inflicting Change?

Same author, same standard: "done" means demonstrated. Inflicting Change is about delivering change through people and politics. This book is about building agentic systems you can prove work. Either can be read first.

About the author

Walter Meyer leads enterprise transformation programs and builds agentic systems. His career began in the U.S. Navy Nuclear Power Program, where "done" meant demonstrated, and has run through a startup he took to profitability and acquisition and senior roles at two creative digital consultancies. He is the author of Inflicting Change and Agents Decide, Tools Do, Gates Prove. More →