Skip to content

Spec-driven development

Updated 26 September 2026 Reviewed by Teemu Malinen

What is Spec-driven development?

A method where the specification, not the code, is the primary artefact. Developers write and refine a structured spec, and an AI coding agent generates code to match it. An open-source toolkit released in 2025 runs a Spec, Plan, Tasks, Implement flow. It is the fast-rising discipline for keeping AI-built software true to intent, instead of drifting prompt by prompt.

Why it matters

Building by prompt has a failure mode that only shows up over time. Each request nudges the code a little, the reasoning behind each nudge lives in a chat log nobody keeps, and after a few weeks no single artefact explains why the system does what it does. Putting the specification first fixes that by moving the source of truth out of the conversation and into a document you can version, review and hand to someone new. When the requirement changes you edit the spec and regenerate, rather than patching code and hoping the intent survives. The discipline it demands is real, since a vague spec produces vague software just as surely as a vague prompt did. The gain is that the thing you argue over and sign off is the thing that actually drives the build.

Spec-driven development went from a niche idea to mainstream vocabulary in less than a year. Amazon’s Kiro launched in July 2025 with specs at its centre, GitHub announced its open-source Spec Kit in September 2025, and Thoughtworks placed spec-driven development in the Assess ring of its Technology Radar in November 2025. By the April 2026 edition (Volume 34), Thoughtworks was describing spec-driven development as one of the controls teams use to keep increasingly capable coding agents in check.

What is spec-driven development?

Spec-driven development is a way of building software in which a written specification, not the code, is the source of truth for both the people and the AI coding agent. Birgitta Böckeler, writing on Martin Fowler’s site in October 2025, describes a spec as a structured, behaviour-oriented artefact written in natural language that expresses what the software should do and guides the agent. IBM defines the method as one where a detailed specification is authored and agreed before development begins. Thoughtworks notes that the term’s definition is still evolving, but that it generally means a workflow that starts from a structured functional specification and breaks it down into smaller pieces, solutions and tasks.

The idea is older than AI coding tools. Ostroff, Makalsky and Paige presented an agile approach to specification-driven development in 2004, combining test-driven development with design by contract. What changed in 2025 is that a coding agent can take a spec and produce working code from it, which makes the quality of the spec the main lever on the quality of the result.

What are spec-first, spec-anchored and spec-as-source?

Spec-driven development comes in three levels of commitment, a split introduced by Böckeler and used since by IBM and in academic write-ups. Spec-first means a well thought-out spec is written before the task and used to guide the AI, but may be dropped once the work is done. Spec-anchored means the spec is kept after the task and maintained as the feature evolves; IBM adds automated tests that link it to the code. Spec-as-source is the most radical: only the spec is edited by people, and the code is regenerated from it. Of the three tools Böckeler reviewed, only Tessl explicitly aimed beyond spec-first, and Thoughtworks noted it was still in private beta as of September 2025. Spec-as-source remains the experimental end of the range.

How does spec-driven development work in practice?

A typical spec-driven workflow runs in four steps: specify, plan, break down into tasks, implement. GitHub’s Spec Kit makes this explicit. A project gets a “constitution” once, a set of principles on code quality, testing and maintainability that every change must respect. Each feature then goes through specify, plan, tasks and implement, and the current version adds a converge step that checks whether the implementation matches the artefacts. Kiro follows a similar path, producing a requirements file, a design file and a task list; at launch it wrote acceptance criteria in EARS (Easy Approach to Requirements Syntax) notation.

Tools aside, the working pattern is the same. The spec lives in the same repository as the code, under review like any other change. A new requirement lands as an edit to the spec, discussed in a pull request, then the implementation is regenerated or updated to match. Six months on, an engineer who never saw the original conversation can read the spec and know what the system is meant to do, and why.

What makes a good spec for an AI coding agent?

A good spec describes what the software does and leaves how it does it to the plan. Liu Shangqi of Thoughtworks recommends that a spec define the software’s external behaviour: inputs and outputs, preconditions and postconditions, invariants and constraints. It should use the business’s own domain language rather than technology-specific terms, and scenarios are often written in the Given/When/Then style familiar from behaviour-driven development. Markus Eisele, writing for O’Reilly in August 2026, puts the target as enough structure to constrain the work, enough examples to make intent concrete and enough executable checks that review does not turn into guessing. A spec the agent can check its own work against, through tests or contracts, is worth more than a long document it can only read.

How does spec-driven development relate to vibe coding and agentic coding?

Spec-driven development sits at the opposite end from vibe coding. Vibe coding means accepting what the AI produces without reviewing the code; spec-driven development adds requirements analysis, design and human sign-off before code is generated. GitHub’s own framing is that a vague prompt forces the model to guess at potentially thousands of unstated requirements, and a spec removes that guesswork.

Agentic coding describes how autonomously the AI works, and spec-driven development describes what it works from. The more steps an agent takes on its own, the more it matters that the goal is written down in a form the agent and the reviewer share. Thoughtworks’ April 2026 Radar groups spec-driven development with agent skills as feedforward controls that steer an agent before it acts, alongside feedback controls such as mutation testing, which prompt the agent to correct itself before a person reviews the work.

When is spec-driven development worth it?

Spec-driven development pays off where intent is easy to lose: new products, features added to an existing system, and modernisation of legacy code. GitHub names these three as the main use cases, noting that a spec can capture the business logic of an old system without carrying over its technical debt. It pays off less on small, well-understood changes. Böckeler found that using Kiro on a small bug fix produced 16 acceptance criteria, far more ceremony than the problem needed. Eisele recommends matching the amount of spec to the work: light specs for exploratory work, clear acceptance criteria for a bounded task, strong contracts and tests for deterministic work such as integrations, and typed, validated contracts when several agents hand work to each other.

What are the risks and limitations?

The main risks are review overhead, agents ignoring the spec and specs drifting from the code. Böckeler found Spec Kit generated many verbose, repetitive markdown files that were tedious to review, and saw agents fail to follow all the instructions despite the templates and checklists. Liu Shangqi notes that spec drift and hallucination are hard to avoid, so teams still need highly deterministic CI/CD to guarantee quality. Thoughtworks warns the industry may be relearning a bitter lesson: handcrafting detailed rules for AI does not scale. Böckeler also draws a parallel with model-driven development, whose spec-as-source ambitions ran into inflexibility, and adds that AI brings non-determinism on top. Eisele points to a quieter problem: when prose piles up beside the code, the team ends up with competing sources of truth that confuse the model instead of guiding it.

What does the research say?

Evidence for spec-driven development is still mostly practitioner experience, and controlled studies are few. One of the first quantitative results comes from a preprint by Pardis Taghavi and Santosh Bhavani (arXiv, April 2026). Across 128 runs covering 32 features in five repositories, adding repository-grounding steps to a pipeline modelled on Spec Kit improved judged quality by 0.15 points on a 1–5 scale, about 3% of the full score, while keeping 99.7–100% compatibility with existing tests. Quality was scored by a language model, not by people, and the paper has not been peer reviewed. The result supports a narrower claim than much of the enthusiasm around the approach: grounding a spec in the real codebase helps, and the size of the effect is modest.

Frequently asked questions

Is spec-driven development just waterfall again?

No, although the risk is real. Waterfall fixed requirements long before delivery. Spec-driven development writes a spec per feature and turns it into code straight away, and Thoughtworks’ Liu Shangqi describes its feedback loops as shorter than waterfall’s. The warning sign is a spec that grows larger than the change it describes.

What is the difference between spec-driven development and vibe coding?

Vibe coding accepts AI-generated code without reviewing it and keeps the intent in the prompt history. Spec-driven development writes the intent down first, reviews it and checks the code against it. Vibe coding suits throwaway prototypes; spec-driven development is meant for software that has to be maintained.

Do you need a special tool for spec-driven development?

No. Kiro, Spec Kit and similar tools add templates and workflow steps, but the practice works with any coding agent: a versioned spec in the repository, a plan and task list derived from it, and tests the agent must pass. What the tools add is repeatability.

How is spec-driven development different from TDD and BDD?

Spec-driven development builds on both. Test-driven and behaviour-driven development already put the expected behaviour before the code; spec-driven development widens the spec to cover design and task breakdown, and hands the implementation to an AI agent. Given/When/Then scenarios and automated tests remain the most checkable parts of a good spec.

Does spec-driven development replace code review?

No. It moves part of the review earlier, to the spec, where mistakes are cheaper to fix. The generated code still has to be reviewed and tested, because agents do not always follow the spec.

Sources

Otto Sunnari, myynti ja kumppanuudet, Sofokus / Otto Sunnari, Sales and partnerships at Sofokus

Ready to start leveraging AI?

Call, email, or book a time straight from my calendar.

Otto Sunnari

Sales and partnerships