Blog
How to break down a spec for AI coding agents
Give an AI coding agent a spec for a whole feature and the first hour usually goes well. Then the details from the start of the spec fade from its context, the diff grows to thirty files, and the review turns into an archaeology project. The fix is not a better prompt. It is a smaller unit of work.
This article is a practical method for breaking a large specification into pieces an agent can implement, verify and hand back in one run — and for keeping those pieces connected, so the whole still adds up to what you meant. We'll use one example throughout: adding a checkout flow to an online shop.
Why big specs fail with AI agents
Three things go wrong when one spec covers too much:
- Context fades. An agent works within a limited context. The longer the session, the more early requirements compete with code it has just read and written. Thoughtworks calls out "instruction bloat" and "context rot" as practical risks of spec-driven workflows.
- Instructions dilute. Addy Osmani describes a "curse of instructions": the more rules you put into one prompt, the worse the model follows each of them. A large spec is a large prompt.
- Review breaks down. A change touching thirty files cannot be reviewed against the spec line by line. People skim it, and mistakes ship.
The opposite failure exists too. Birgitta Böckeler describes a small bug fix that a spec tool turned into "4 'user stories' with a total of 16 acceptance criteria." Breaking work down is not about making more documents. It is about making each piece the right size.
Signs a spec is too big for one agent run
Before handing a spec to an agent, check it against these signals. Two or more usually means it needs to be broken down:
- It describes several things a user can do, not one.
- Its acceptance criteria run past a screen, or fall into clear groups.
- It touches several parts of the system that could be built and tested separately.
- You can't name the one check — a test, a command, a visible behaviour — that proves it works.
- You already know some parts must exist before others can start.
- You'd want different people, or different agents, working on parts of it at the same time.
The checkout flow fails most of these: a cart, a payment step and a shipping step are each things a user does, each with its own edge cases, and payment can't be finished until there is a cart to pay for.
Step 1: Clarify the whole before you split it
It is tempting to split first and ask questions later. Do it the other way round. Questions about the whole feature change how it should be split.
For the checkout flow, an agent asked questions like these before writing anything:
- Should guest checkout skip account creation entirely, or offer it as an optional step?
- Should shipping and billing addresses share one form, or stay as two separate steps?
- Do we need saved payment methods for returning users, or only one-time card entry for now?
Each answer moves a boundary. "One-time card entry for now" removes a whole stage for saved payment methods. "Two separate steps" for addresses makes shipping its own piece of work. Write the answers into the parent spec: they are the context every child will inherit.
Step 2: Split by what the user can do, not by layer
There are two common ways to cut a feature:
- By layer — database, then API, then UI. Each piece is technically coherent, but none of them does anything on its own, and nothing can be verified end to end until the last piece lands.
- By capability — cart, then payment, then shipping. Each piece is a thin slice through every layer it needs, and each can be tried, tested and reviewed on its own.
Prefer capabilities. For the checkout flow, that gives three stages:
- Checkout flow
- Cart
- Payment
- Shipping
Each stage still gets its own questions, at its own level of detail. For the cart: should quantity changes update totals live, or only on an explicit "Update cart" click? For payment: should a declined card retry in place, or send the user back to the cart? These questions would have been noise at the feature level; at the stage level they are exactly what the agent needs to know.
Step 3: Keep going only where a piece is still too big
Decomposition is a tree, not a fixed number of levels. Apply the size check from above to each child and split again only where it fails:
- Checkout flow
- Cart
- Payment
- Card entry and validation
- Declined-card handling
- Order confirmation
- Shipping
Here the cart and shipping are small enough as they are; payment, with its error paths and confirmation, gets one more level. Uneven depth is a good sign: it means the tree follows the actual complexity instead of a template.
Step 4: Write a leaf an agent can finish
A leaf is a piece you will hand to an agent as it is. A good leaf passes all of these:
- One outcome. It describes one change in behaviour, in one or two sentences.
- Testable acceptance criteria. For example in Given / When / Then form, or as "WHEN [event] THE SYSTEM SHALL [behaviour]" in Kiro's EARS style.
- Explicit boundaries. What is out of scope, and what the agent must not touch.
- A way to verify. The test, command or manual check that proves it works.
- A reviewable result. You expect a diff a person can read against the spec in one sitting.
For example:
# Declined-card handling
Goal: when a card payment is declined, the customer can fix it without losing their cart.
Acceptance criteria:
- Given a declined payment,
when the payment provider returns a decline,
then the payment step stays open with the card form and an error message.
- The cart, addresses and shipping choice are unchanged after a decline.
- After three declines in a row, the customer is offered a different payment method.
Out of scope: saved payment methods, retries in the background.
Depends on: Card entry and validation.
Verify: the payment integration tests pass, including the three new decline cases.Step 5: Record dependencies as blockers
Pieces of a tree are rarely independent. Declined-card handling needs Card entry and validation; Order confirmation needs a successful payment; Payment as a whole needs a Cart to pay for.
Write these down as blockers — "this spec can't start until that one is done" — instead of relying on the order of a list. Blockers give you three things:
- A build order. The next piece to implement is always one whose blockers are all done.
- Safe parallelism. Pieces without blockers between them — say, Shipping and Card entry — can be built at the same time, by different people or agents.
- Unattended runs. An agent can work through the tree on its own, picking the next ready leaf each time, without you sequencing the work by hand.
Step 6: Give each leaf only the context it needs
A leaf should not carry the whole project. It needs:
- its own spec;
- the decisions from its ancestors that affect it — for Declined-card handling, that guests can check out and that there are no saved cards;
- the specs it depends on, or at least what they produced.
Everything else is noise that competes for the agent's attention. This is the main reason to keep the tree instead of flattening it into a task list: the tree tells you which context belongs to which leaf.
Common mistakes
- Splitting before clarifying. You end up with pieces drawn around assumptions that the first answered question overturns.
- Splitting by layer. Nothing works end to end until everything is done, and nothing can be verified on its own.
- Uniform depth. Forcing every branch to three levels creates trivial leaves in simple areas and oversized ones in complex areas.
- Hidden dependencies. If the order of work only exists in your head, an agent working alone will pick the wrong piece first.
- Too much ceremony for small changes. A one-line fix needs one short spec, or none. Decompose when the size check says so, not by default.
- Letting the tree go stale. When a requirement changes, change the spec first, then the code — otherwise the next agent run implements yesterday's plan.
How to do this with Athenode
Athenode is built around this method, so each step maps to a built-in skill:
- Clarify the whole:
/mindmap-specification I want to add a checkout flowasks rounds of questions grounded in your codebase and writes the parent spec from your answers. - Split and clarify each piece:
/decompose-specificationproposes the stages, asks each stage its own questions, and finalizes them — as many levels deep as each branch needs. Blockers between specs are recorded as you go. - Implement:
/apply-specificationtakes a whole subtree and works through every prepared leaf in blocker order, planning, implementing and reviewing each one without stopping to ask you — for hours on a large tree. Each spec keeps a record of the files that changed and what was done. - Small change? Skip decomposition entirely: mind-map one spec, answer a couple of questions, and apply it straight away.
The tree lives in Athenode, shared by your team, and the skills work in 9 AI coding tools, including Claude Code, Cursor and Codex. See the Getting Started guide to try it, or read what spec-driven development is for the bigger picture.
Checklist
Before you hand a piece of work to an agent:
- The parent spec's open questions are answered and written down.
- The work is split by user-facing capability, not by technical layer.
- Each leaf has one outcome, testable acceptance criteria, boundaries and a way to verify it.
- Each leaf is small enough to review against its spec in one sitting.
- Dependencies are recorded as blockers, not implied by list order.
- Each leaf carries only the context it needs from its ancestors and blockers.
Sources
- Thoughtworks, Technology Radar: GitHub Spec Kit (on instruction bloat and context rot)
- Addy Osmani, How to write a good spec for AI agents, O'Reilly Radar
- Birgitta Böckeler, Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl, martinfowler.com
- Kiro, Feature Specs (EARS notation) and Specs best practices