AthenodeAthenode

Guide

Spec-driven development: a practical guide

Spec-driven development (SDD) is a way of building software with AI coding agents in which you write a structured specification — what to build, why, and how you will know it works — before any code is generated. The spec, not the chat history, becomes the source of truth that both you and the agent work from.

This guide explains where spec-driven development came from, how it differs from vibe coding and test-driven development, what a good spec for an AI agent contains, how to break a large spec into pieces an agent can finish in one run, and where the approach goes wrong.

What is spec-driven development?

Definitions vary, because the practice is young. Three widely cited ones:

  • Birgitta Böckeler on martinfowler.com: "Spec-driven development means writing a 'spec' before writing code with AI ('documentation first'). The spec becomes the source of truth for the human and the AI."
  • The Thoughtworks Technology Radar (November 2025, ring "Assess"): workflows that "begin with a structured functional specification, then proceed through multiple steps to break it down into smaller pieces, solutions and tasks."
  • GitHub, announcing Spec Kit: "We're moving from 'code is the source of truth' to 'intent is the source of truth.'"

What they share: intent is written down and agreed on first, the work is broken into smaller steps, and the written spec — not a prompt you typed an hour ago — is what the agent implements and what you check the result against.

Why spec-driven development took off in 2025

Andrej Karpathy coined "vibe coding" in February 2025: describe what you want in a prompt, accept what the model produces, and keep prompting until it seems to work. It is fast for prototypes, but as GitHub put it, the output often "looks right, but doesn't quite work" — the agent fills every gap in a vague prompt with its own guesses.

Spec-driven development emerged as the answer for work that has to reach production. Within a few months:

  • July 2025 — AWS released Kiro, an agentic IDE built around requirements, design and task documents.
  • September 2025 — GitHub open-sourced Spec Kit, and Tessl launched its spec framework and registry.
  • November 2025 — Thoughtworks put spec-driven development on its Technology Radar, and its engineers called it one of the most important practices to emerge in 2025.

The underlying idea is not new. Behaviour-driven development, design by contract and model-driven development all put a specification ahead of the code. What changed is who writes the code: an AI agent takes a spec far more literally than a human colleague does, so the quality of the spec now directly decides the quality of the result.

Spec-driven development vs vibe coding vs TDD

Vibe coding Test-driven development Spec-driven development
What you write first A prompt A failing test A specification
What drives the code The conversation Tests that must pass The spec and its acceptance criteria
Who writes the code An AI agent Usually a human An AI agent
Where intent lives In the chat history In the test suite In a versioned spec
Best for Prototypes, throwaway scripts Well-understood units of logic Features that must work and be maintained

These are not mutually exclusive. A good spec includes acceptance criteria that become tests, so spec-driven development and TDD fit together naturally — the spec says what and why, the tests prove it.

The three levels of spec-driven development

Böckeler distinguishes three levels of commitment to the spec:

  1. Spec-first — you write a spec for the task at hand and use it to guide the agent. After the task, the spec may be thrown away.
  2. Spec-anchored — the spec is kept and maintained after the task, and used again as the feature evolves.
  3. Spec-as-source — the spec is the main artefact; humans edit only the spec and the code is regenerated from it.

Most teams today work at the first or second level. Spec-as-source is the most ambitious, and Böckeler warns it risks combining "the downsides of both MDD and LLMs: inflexibility and non-determinism." A practical middle ground is spec-anchored work with a tree of specs: each spec stays attached to the code that implemented it, so you can trace any change back to the intent behind it.

The spec-driven development workflow

Tools name the steps differently, but nearly every SDD workflow follows the same shape:

  1. Clarify the intent. Pin down what you are building, for whom, and what is out of scope — before discussing implementation. Questions from the agent are a feature here: every question answered now is a guess the agent won't make later.
  2. Write the specification. Capture goals, scope, acceptance criteria and constraints in a structured document.
  3. Decompose it. Break the spec into smaller specs until each piece is small enough for an agent to implement and verify in a single run.
  4. Order the work. Record which pieces depend on which, so nothing is built on top of something that doesn't exist yet.
  5. Implement piece by piece. Hand each ready piece to the agent, with its spec and only the context it needs.
  6. Review against the spec. Check each result against the acceptance criteria, not against your memory of what you meant.
  7. Keep the spec alive. When requirements change, change the spec first, then the code.

What a good spec for an AI agent contains

A spec is not a long requirements document. It is the smallest amount of writing that removes the agent's need to guess. Drawing on Addy Osmani's guidance, GitHub's analysis of 2,500 agent instruction files and Kiro's EARS format, a useful spec covers:

  • Goal and why — one or two sentences on the problem and who it is for.
  • Scope and non-goals — what is in, and explicitly what is not.
  • Acceptance criteria — testable statements, for example Given / When / Then or EARS ("WHEN [event] THE SYSTEM SHALL [behaviour]").
  • Constraints and boundaries — what the agent must always do, must ask about first, and must never do.
  • Dependencies — which other specs or existing parts of the system it relies on.
  • How to verify — the commands, tests or checks that prove it works.

A leaf-level spec can be short:

# Reward notification email

Goal: tell a referrer by email when one of their invites earns them a reward.

Scope: one transactional email, sent once per reward.
Non-goals: marketing emails, push notifications, digests.

Acceptance criteria:
- Given a reward is credited to a referrer's balance,
  when the ledger entry is created,
  then exactly one email is sent to the referrer within 5 minutes.
- The email shows the reward amount and the invitee's first name only.
- No email is sent if the referrer has unsubscribed from transactional mail.

Depends on: Ledger entry per reward.
Verify: the integration test for the reward flow passes; the email renders in the preview tool.

Two warnings from practitioners. First, more instructions are not always better: Osmani calls it the "curse of instructions" — the more rules you pile into one prompt, the worse the model follows each of them. Second, keep the what separate from the how: GitHub recommends leaving technical decisions out of the first, functional version of a spec.

How to break a large spec into agent-sized pieces

A spec for a whole feature is too large for one agent run: the agent loses track of earlier details as its context fills up, and the review becomes a single enormous diff. The fix is decomposition — but into a tree, not a flat task list.

Take a Referral rewards program. Decomposed level by level, it might look like this:

  • Referral rewards program
    • Referral link generation
      • Unique invite codes
      • Share buttons on the profile page
    • Reward payout rules
      • Qualifying purchase checks
      • Minimum order amount
      • Refund clawback window
    • Account balance ledger
      • Ledger entry per reward
      • Crediting the account balance
      • Reward notification email
    • Referral stats dashboard
      • Invites and conversions chart
      • Export to CSV

Rules of thumb that make the tree work:

  • Stop when a leaf fits one agent run. A leaf is small enough when an agent can implement it, run its checks and hand back a reviewable change without needing the rest of the tree in its context.
  • Decompose only where needed. Some branches stop at the second level; others need a third. Depth should follow complexity, not a template.
  • Record dependencies as blockers. Crediting the account balance depends on Ledger entry per reward; the dashboard depends on both payout rules and the ledger. Blockers give you a build order, and they let several agents work in parallel on leaves that don't depend on each other.
  • Let children inherit context. A leaf should carry what it needs from its parent and its blockers — not the whole project.
  • Ask questions at every level. The questions worth asking about the program ("can rewards be clawed back?") are different from those worth asking about one email ("which name do we show?"). Clarify each spec at its own level of detail.

Common pitfalls and criticisms

Spec-driven development has sharp critics, and their points are worth taking seriously:

  • Too much process for small changes. Böckeler describes a small bug turning into "4 'user stories' with a total of 16 acceptance criteria." Not every change needs a spec — a one-line fix needs a good commit message.
  • Markdown bloat and review burden. Marmelab reported a tool producing 8 files and 1,300 lines of text for a feature that displays the date. If reviewing the spec takes longer than reviewing the code, the spec is too big.
  • Spec drift. Specs and code diverge as soon as someone changes one without the other. Thoughtworks engineers note that "spec drift and hallucination are inherently difficult to avoid."
  • Instruction bloat and context rot. Thoughtworks flags growing instruction files as a risk: past a point, the agent follows each instruction less reliably.
  • Waterfall in disguise. Gojko Adzic and Marmelab warn that writing everything up front brings back the worst parts of waterfall. Others argue the opposite: with an agent implementing each piece in minutes, feedback loops stay short.

Most of these come from the same root cause: one big document written once. Smaller specs, clarified one level at a time, kept attached to the code they produced and updated when the code changes, avoid most of them.

Spec-driven development tools

The main options as of October 2026:

Tool Made by What it is Main artefacts
GitHub Spec Kit GitHub Open-source (MIT) toolkit of agent commands Constitution, spec, plan and task Markdown files
Kiro AWS Agentic IDE and CLI requirements.md, design.md, tasks.md per feature
OpenSpec Fission AI Open-source (MIT) workflow aimed at existing codebases Current specs plus change folders with proposal, design and tasks
BMAD Method BMad Code Open-source (MIT) method with role-based agents Planning documents produced by product, architecture and development roles
Tessl Tessl Spec framework and registry, now focused on agent skills Specs tied to individual code files
Cursor Plan Mode Anysphere Planning mode inside the Cursor editor An editable Markdown plan
Athenode Athenode Shared specification tree plus a CLI and MCP server for 9 AI coding tools A tree of specs from idea to leaf, linked by blockers

Most of these keep specs as Markdown files in one repository and focus on one feature or change at a time. They differ in how much structure they impose, whether they lock you into one editor, and whether the spec outlives the task. For a side-by-side comparison, see spec-driven development tools compared.

How Athenode approaches spec-driven development

Athenode is built around the tree described above:

  • What before how. /mindmap-specification draws a spec out of you through rounds of questions, grounded in your existing codebase, before any implementation plan exists.
  • A tree, not a flat list. /decompose-specification breaks a spec into stages, each clarified with its own questions, as deep as each branch needs — up to a whole project in one tree.
  • Blockers set the build order. Leaves wait until what they depend on is done, so agents only pick up work that is ready.
  • Unattended implementation with review. /apply-specification plans, implements and reviews every prepared leaf in a subtree, and records the changed files and a summary against each spec.
  • Shared by the team. The tree lives in Athenode, not on one laptop, so everyone sees what is specified, decomposed and implemented.
  • Works with the agent you use. The CLI installs the skills, agents, MCP servers and AGENTS.md into 9 AI coding tools, including Claude Code, Cursor and Codex. Your code stays in your repository.

To try it, follow the Getting Started guide or run npx @athenode/cli init in your project.

Frequently asked questions

Is spec-driven development just waterfall?

It can become waterfall if you write one large spec up front and never revisit it. Done well, it is the opposite: small specs, clarified one level at a time, each implemented and reviewed within minutes, with the spec updated whenever the requirements change.

Do I need a spec for every change?

No. Small, well-understood fixes don't need one. Specs pay off when a change has several parts, touches existing behaviour, or will be maintained by more than one person.

What is the difference between a spec and a prompt?

A prompt is a one-off instruction that lives in a chat. A spec is a structured, versioned document with acceptance criteria that outlives the conversation: you can review it, share it, implement it again and check the result against it.

Which AI coding agents work with spec-driven development?

Any agent that can read files and follow instructions. Spec Kit and OpenSpec work with many agents through slash commands; Kiro and Cursor build it into their own editors; Athenode supports Claude Code, Cursor, Codex, GitHub Copilot, Kiro and four more tools through its CLI.

Sources

Ready to ship better, together?

Spec it. Decompose it. Ship it. All with your AI agent.

Start for free

Join engineers building with Athenode today.