AI-Driven Developmentintermediate8 min

Spec-Driven Development

Turn a rough idea into a written spec the AI builds against — the source of truth that makes everything downstream go smoothly.

Hand an AI a one-line request — "add discount codes to checkout" — and it will cheerfully build something. The trouble is that it had to make dozens of small decisions you never stated: what the discount comes off, whether two codes can be combined, what happens when a code has expired, who is allowed to create one. Some of those guesses will be wrong, and because the code runs and looks reasonable, you'll only find out after it ships.

Spec-driven development fixes this by inserting a deliberate step between the idea and the build: a written spec. The spec is the source of truth — a clear description of the desired behavior that the AI builds against, so the decisions get made on purpose instead of by accident.

The spec is the source of truth

A spec isn't a vague paragraph of hopes; it's a concrete description of what the thing should do. What's in scope and what isn't, the key behaviors, the edge cases, the constraints. Its real value shows up in the questions it forces into the open. A good spec doesn't paper over ambiguity — it surfaces open questions so you can answer them deliberately. "Can a customer use two codes on one order?" is a cheap question to answer in a spec and an expensive one to discover in shipped code.

The most useful lines in a spec are its acceptance criteria: short statements the finished code can either pass or fail. "Shipping is never discounted" is one. "Discounts work correctly" is not, because no build could ever fail it. Each real criterion does two jobs. Before the build, it tells the AI what to make. After the build, it becomes a check you can run against what the AI made, so "is this right?" turns into a question with a yes-or-no answer.

Step through the same small feature built twice. The first time it starts from a one-line prompt. You can see the rules the product owner has in mind, but the AI can't, and before the bug surfaces you'll be asked to find it yourself by checking the code against those rules. Then switch to Spec first to replay it with the rules written down, and compare when each version finds out something is wrong.

Check yourself

A demo of the new discount code works perfectly, but nobody has run the spec's acceptance criteria yet. What does the passing demo tell you?

How it works: idea to build

The flow is short and worth following in order. You start with a rough idea — fuzzy and one-line, the way real features begin. You turn it into a detailed spec, ideally with the AI asking clarifying questions rather than guessing: a question loop where it raises what's unclear, you answer, and the spec tightens with each pass. Once the spec is solid, you break it into user stories — concrete, buildable slices of behavior, each carrying its own acceptance criteria. Only then does the build begin, with the spec and stories as the reference the AI works against.

The one constant throughout is the human at the review gate. You read and refine the spec before any code is written, because that's the cheapest possible place to catch a misunderstanding. A wrong rule caught while it's still a sentence costs you an edit; the same rule caught in production costs a fix, a release, and whatever the bug did in the meantime.

Note

In our stack — Claude Code running Anthropic's Claude models works best when the spec lives in a Markdown file in the repo. You ask Claude to draft the spec and surface open questions, answer them in the file, and then point it at that file to generate stories and build. The file becomes shared context both you and the model return to — far more durable than instructions buried in a long chat.

Watch out

A criterion that can't fail isn't a criterion. The most common way to get a spec wrong is to write it in the same vague words as the prompt: "handles discounts properly," "export is fast," "works for large accounts." They read like requirements but rule nothing out, so the AI still guesses and nothing can catch a bad guess. If you can't picture an input that would make a line fail, rewrite it with a number, an example or a named edge case: "exports up to 50,000 rows," "a $50 order with $10 shipping totals $55."

Keep the spec in a file, not a chat

It's tempting to keep all of this in the conversation — describe the feature, let the AI build, correct it as you go. But chats are a poor home for a source of truth. They grow long, and important decisions get buried under dozens of later messages; they're lossy, as earlier context falls out of the window or gets summarized away; and they're hard to edit, since you can't cleanly revise a decision you made twenty turns ago.

A file has none of those problems. It's stable, it's the whole spec in one place, and you can edit any line of it directly. When you revise the file, you're updating the single artifact everything else builds on — no archaeology through chat history required. This file-based spec is the backbone of the larger AI workflows: the greenfield workflow starts from one, and review gates check work against one. Get the spec right, in a file, and the rest of the process has something solid to stand on.

Check yourself

Twenty messages into a chat with an AI, you change your mind about how expired codes should behave. What's the most reliable way to make sure the build follows the new rule?

Key takeaways

  • In spec-driven development the spec is the source of truth: the AI builds against a written description of the desired behavior, not against a vague request.
  • The flow runs idea → spec → user stories → build, with a human reviewing the spec before any code is written.
  • A good spec surfaces open questions instead of guessing — ambiguity caught here is far cheaper than ambiguity discovered in the finished code.
  • Acceptance criteria must be able to fail: each one tells the AI what to build, then becomes a check you run against what it built.
  • Keep the spec in a file, not a chat: chats grow long and lossy, while a file is a stable, editable artifact you and the AI can both return to.

Keep going