← all posts

Three files keep my agents on goal

2026-05-22

Give an agent a goal and a list of candidate tools, and you'll get one of two failure modes. Accumulation — it tries to integrate all of them. Or drift — it solves an adjacent problem and reports back with something tangential to what you asked for.

For exploration-shaped work I've landed on a three-file shape that prevents both. Nothing fancy, just markdown. The example below is from a small repo where I surveyed about twenty AI-dev tools and composited a working harness from the survivors — but the pattern transfers to any "evaluate N candidates, keep what's good" task.

The three surfaces

vision.md is the north-star brief. About thirty lines: the goal in one paragraph, two domains of focus, the seed list of candidates. It's stable — I don't rewrite it to record decisions. Its job is to be re-readable by the agent at any point, so that "what are we doing here?" has a deterministic answer.

explorations/ holds one note per candidate, plus an index. Every note follows the same shape — what it is, install, strengths, weaknesses, overlap with other candidates, verdict — and ends with one of five labels: adopt, trial, hold, reject, reference. The shape is enforced by a CLAUDE.md at the repo root that prescribes the per-candidate flow: fetch, summarise, map to goals, capture, update index.

harness/ is the composited result. A single result.md describes the stack a colleague would adopt: chosen tools, why each was picked, install steps, integration. A decisions.md next to it logs cross-cutting tradeoffs that don't belong in any one tool's note. Every time a candidate earns an adopt verdict, the harness updates immediately. It's continuously written, not big-bang at the end.

The thing that makes this work isn't the file layout. It's that each file has one job and one update trigger. vision.md changes when the goal changes. An exploration note changes when its upstream tool changes. result.md changes when something is adopted or replaced. Without that separation, source notes silently absorb adoption opinions, and the agent can't re-orient cleanly on a re-read.

The loop in practice

The typical exchange is short: "check this out, github.com/some-framework". The agent reads vision.md and the relevant existing exploration notes, fetches the candidate, writes a per-project note in the prescribed shape, lands on a verdict, updates the index. If the verdict is adopt, it touches result.md. If it's reject, it doesn't — but the note stays as a record so the same candidate isn't re-litigated three months later.

One example. I evaluated Superpowers, a Claude Code skills plugin. The exploration note ended with: adopt the test-driven-development skill standalone, reject the rest of the plugin. The reason — its auto-trigger skills would collide with the workflow tool I'd already picked — didn't belong in the Superpowers note (that note is about Superpowers, not about what I do with it), so it went into decisions.md. The TDD skill went into result.md with its own block, pinned by commit. A reader of result.md now knows what I'm using and why; a reader of the exploration note knows what Superpowers is. Different lifecycles, different files.

What this doesn't solve

The seed list isn't a scope fence. Adjacent tools surface during exploration and the agent will add them — that's intentional, but it means "done" is a judgment call, not a checkbox.

Verdicts also aren't always binary. Partial adoption (one skill from a plugin, the memory layer from a tool but not its planner) is common, and the exploration note has to be explicit about what's in and what's out, or the next session will get it wrong. The five-label ladder helps; carve-out language in the verdict line — "adopt — TDD skill only" — does the rest.

And the pattern presupposes that exploration is the work. If exploration is a side-quest to a bigger project, a lighter shape probably wins; this one earns its overhead when the survey is what you're delivering.

What to take from this

You don't need this exact layout. What matters is that three things have separate, durable homes: the goal, the per-candidate evaluation, the composited result. Each file has one job. Each file has one update trigger. Agents work well when the surfaces they're re-reading are predictable — and predictable files are what kept the survey from collapsing into either a feature checklist or a tangent.