Agentic coding
Also known as: agentic engineering
What is Agentic coding?
A way of working where an AI agent takes a goal, plans the steps, edits files, runs commands and iterates until tests pass, with the developer directing and approving rather than typing each line. It moves past autocomplete to autonomous task execution. The shift is from writing code to reviewing proposed changes.
Why it matters
Handing over the keystrokes changes the shape of the work more than its speed. When an agent plans, edits and runs on its own, the developer’s day fills with a different task: reading proposed changes and deciding whether to trust them. That reads as easier and often is not, because judging a diff you did not write, across files you were not watching, takes real attention. The method rewards tasks with a clear finish line, a failing test to turn green or a well-bounded feature, where the agent can check its own work. Point it at something open-ended with no acceptance criteria and it will produce something plausible and confident that misses the point.
The same holds at team level. Google and DORA’s 2025 report on AI-assisted software development, based on nearly 5,000 respondents, concludes that AI does not fix a team but amplifies what is already there. A team with solid tests, small changes and working review gets faster with agents. A team without them gets its existing problems faster too.
How does agentic coding work?
At its core, agentic coding is a loop. The developer gives a goal, the agent reads the relevant parts of the codebase, plans the steps, edits files, runs commands and tests, reads the results and corrects itself until the task is done or it needs help. According to IBM, agentic coding centres on coding agents: AI systems that combine the reasoning of a large language model with access to coding tools and an execution environment. That access is what separates an agent from an autocomplete tool: rather than only suggesting text, it acts on the repository and sees what happens.
Anthropic separates agents from workflows. In a workflow, the steps are fixed in advance by code. An agent decides its own next step and which tool to use. Software is a natural fit for agents because the result can be checked: tests pass or fail, and the agent can use that signal to iterate. Anthropic adds that human review still matters, because passing tests do not prove that a change fits the wider system.
How much autonomy does a coding agent have?
How much a coding agent does on its own is a setting, not a fixed property of the tool. Kevin Feng, David McDonald and Amy Zhang (2025) describe five levels of agent autonomy by the role the person plays: operator, collaborator, consultant, approver and observer. Their key point is that an agent’s level of autonomy can be chosen as a deliberate design decision, separate from what the agent is technically capable of. The same agent can be run on a short leash for a payment module and a long one for a throwaway script.
In day-to-day use the split is often between supervised and autonomous agents. Birgitta Böckeler of Thoughtworks describes supervised coding agents as interactive agents that a developer drives and steers in the editor, and autonomous background agents as headless agents sent off to work through a whole task in their own environment, usually ending in a pull request. Agents can also work for longer stretches than before. METR’s March 2025 study found that the length of tasks frontier agents could complete with 50% reliability had doubled roughly every seven months over six years, with the caveat that the estimate depends on which tasks are measured.
What is agentic engineering?
Agentic engineering is a synonym for agentic coding that stresses the professional discipline around it. Andrej Karpathy, who coined vibe coding in February 2025, put forward agentic engineering in early 2026 as the name for its professional counterpart: engineers orchestrating agents that write the code while they provide oversight. In his April 2026 notes from Sequoia’s AI Ascent, Karpathy calls it the discipline of coordinating fallible agents while preserving correctness, security, taste and maintainability, and writes that the engineer is still responsible for the software, just as before.
Practitioners picked the term up within weeks. In February 2026 Simon Willison began publishing a guide called Agentic Engineering Patterns, defining the practice as building software with coding agents that can both generate and execute code, used by professional engineers to amplify existing expertise. IBM published its own explainer the same month. In practice, agentic coding names the way of working and agentic engineering emphasises the skills: writing specs, supervising plans, reading diffs, building tests and managing permissions.
How does agentic coding differ from AI-assisted development and vibe coding?
Within AI-assisted software development, agentic coding marks a level of autonomy, while vibe coding marks the absence of review. AI-assisted software development is the umbrella: any building of software with AI in the loop, from line completion to fully delegated tasks. Agentic coding sits at the more autonomous end of that range, where the AI carries out multi-step work instead of suggesting the next line. The dividing question is how far the AI may act before a person steers or approves.
Vibe coding answers a different question: is the code reviewed at all? Simon Willison defines vibe coding as building software without reviewing the code the model writes. IBM draws the contrast in similar terms: vibe coding is the informal end of the spectrum, driven by prompts and speed, while agentic coding adds discipline, with agents given defined roles and constraints and producing code that can be tested and iterated on. An agent can be used in a vibe-coding style, by merging whatever it produces. Autonomy says how much the agent does. Review decides whether the result is engineering.
Review becomes the developer’s main job
Once agents write most of the code, the constraint in agentic coding moves from writing code to accepting it, and reviewing becomes the developer’s main contribution. Faros AI’s AI Engineering Report 2026 draws on two years of telemetry from 22,000 developers in more than 4,000 teams. As teams moved from low to high AI adoption, the acceptance rate of AI-generated code rose from 20% to 60%, merged pull requests per developer grew by 16.2%, and median time in review rose by 441.5%. Pull requests merged without any review, human or agentic, rose by 31.3%, and incidents per pull request by 242.7%. Faros sells engineering-metrics software and the report is not peer reviewed, but telemetry is firmer evidence than a survey.
Full delegation remains rare even among heavy users. In Anthropic’s December 2025 study of its own staff (a survey of 132 engineers and researchers, 53 interviews and 200,000 Claude Code transcripts), more than half said they could fully delegate only 0–20% of their work. Between February and August 2025, the maximum number of consecutive tool calls per transcript, meaning actions such as editing files or running commands without human input, rose from 9.8 to 21.2, while the average number of human turns fell from 6.2 to 4.1. Agents act longer between check-ins, so each review covers more work.
Which tasks suit agentic coding?
Tasks with an end state the agent can verify for itself suit agentic coding best. Give an agent a bug with a failing test and a clear definition of done, and it can loop on its own until the test passes and the fix holds. Give it “clean up the payments module” with no test and no boundary, and it churns out a large, reasonable-looking diff that changes behaviour nobody signed off on. Same agent. The scoping made the difference.
Good candidates include bug fixes with reproducible tests, test coverage for existing code, well-specified features, dependency upgrades and refactoring behind a stable test suite. Weak candidates are tasks with unclear business rules, architecture decisions and anything whose correctness only a domain expert can judge. Böckeler’s experiments at Thoughtworks show why the environment matters: when an agent could not run the tests, it could not catch its own regressions before opening a pull request.
What are the risks of agentic coding?
The main risks of agentic coding are plausible but wrong changes, duplicated code, security exposure and eroding skills. In Böckeler’s June 2025 comparison of autonomous agents on the same task, only two of six runs found existing code to reuse; the other four created duplicate functionality. In a follow-up experiment building applications end to end, her team saw agents add unrequested features, make unsupported assumptions about business logic and claim success while tests were failing. Their conclusion was that a human in the loop to supervise generation remains essential.
Security risks grow with the permissions an agent holds. OWASP’s 2025 Top 10 for LLM applications lists excessive agency, with three root causes: excessive functionality, excessive permissions and excessive autonomy. Simon Willison’s “lethal trifecta” names the dangerous combination: access to private data, exposure to untrusted content and the ability to communicate externally. An agent that reads issues, web pages or dependencies is exposed to prompt injection, and Willison notes that no reliable defence exists yet. Anthropic also flags higher cost and compounding errors, and IBM warns that over-reliance can erode developers’ own skills.
How to adopt agentic coding safely
Safe adoption of agentic coding starts with limits, not with speed. Give agents the least privilege the task needs and require human approval for high-impact actions, as OWASP recommends; in software work these include deployments, database changes and anything touching secrets. Anthropic recommends extensive testing in sandboxed environments. Make sure the agent can run the test suite before it starts. Simon Willison’s Agentic Engineering Patterns guide has chapters on running the existing tests first and on red/green test-driven development, and lists inflicting unreviewed code on collaborators as an anti-pattern.
Keep changes small enough to review properly, and set the autonomy level per task type rather than per team. Plan review capacity before scaling agent use, since faster generation lands on reviewers first. Measure delivery outcomes such as change failure rate and time in review, not just the amount of code produced.
Frequently asked questions
What is the difference between agentic coding and agentic engineering?
They describe the same practice. Agentic coding names the way of working, where an AI agent plans and carries out coding tasks. Agentic engineering, popularised by Andrej Karpathy and Simon Willison in 2026, emphasises the professional discipline: specs, oversight, tests and responsibility for the result.
Is agentic coding the same as vibe coding?
No. Vibe coding means accepting AI-generated code without reviewing it. Agentic coding describes how much the AI does on its own. An agent’s output can be reviewed carefully or merged unread, and only the first is professional software development.
Does agentic coding replace developers?
Not on current evidence. In Anthropic’s study of its own engineers, more than half could fully delegate only 0–20% of their work. The developer’s role shifts from typing code to scoping tasks, reviewing changes and owning the result.
Is agentic coding safe for production code?
It can be, with the same controls as any other change plus a few more. Code from an agent needs review, tests and security checks before production, and the agent itself needs limited permissions, an isolated environment and human approval for high-impact actions.
Why does agentic coding slow down code review?
Because agents produce more and larger changes than reviewers can absorb. In Faros AI’s 2026 telemetry, median time in review rose by 441.5% as teams moved to high AI adoption, while merged pull requests per developer rose by only 16.2%.
Sources
- IBM Think (2026): What is agentic coding? – coding agents combine LLM reasoning with coding tools and execution environments; agentic coding versus vibe coding; generated output treated as a first draft requiring review.
- IBM Think (2026): What is agentic engineering? – term coined by Andrej Karpathy in 2026; engineering expertise used to orchestrate and oversee AI agents.
- Andrej Karpathy (2026): Sequoia Ascent 2026 summary – agentic engineering as the discipline of coordinating fallible agents while preserving correctness, security, taste and maintainability.
- Simon Willison (2026): Writing about Agentic Engineering Patterns – agentic engineering defined as building software with coding agents that can generate and execute code.
- Simon Willison: Agentic Engineering Patterns (guide) – chapters on running the tests first and red/green TDD; unreviewed code inflicted on collaborators as an anti-pattern.
- Anthropic (2024): Building effective agents – workflows versus agents; tests as feedback for coding agents; human review, cost and compounding errors; extensive testing in sandboxed environments.
- Feng, McDonald & Zhang (2025): Levels of Autonomy for AI Agents – five levels (operator, collaborator, consultant, approver, observer); autonomy as a design decision separate from capability. Preprint.
- METR (2025): Measuring AI Ability to Complete Long Tasks – length of tasks agents complete with 50% reliability doubling roughly every seven months over six years; methodological caveats.
- Birgitta Böckeler, martinfowler.com (2025): Autonomous coding agents — a Codex example – supervised versus autonomous background agents; 2 of 6 runs reused existing code.
- Birgitta Böckeler et al., martinfowler.com (2025): How far can we push AI autonomy in code generation? – unrequested features, false success claims; human supervision remains essential.
- Faros AI (2026): The AI Engineering Report 2026 — The Acceleration Whiplash – telemetry, 22,000 developers / 4,000+ teams; acceptance 20% → 60%, merged PRs per developer +16.2%, median time in review +441.5%, PRs merged without review +31.3%, incidents per PR +242.7%. Vendor report, not peer reviewed.
- Anthropic (2025): How AI is transforming work at Anthropic – survey of 132 staff, 53 interviews, 200,000 Claude Code transcripts; over half can fully delegate only 0–20% of work; consecutive tool calls 9.8 → 21.2. Vendor study of its own staff.
- Google / DORA (2025): State of AI-assisted Software Development – nearly 5,000 respondents; AI amplifies existing team strengths and weaknesses; negative relationship with delivery stability.
- OWASP (2025): LLM06:2025 Excessive Agency – excessive functionality, permissions and autonomy; least privilege and human approval for high-impact actions.
- Simon Willison (2025): The lethal trifecta for AI agents – private data, untrusted content and external communication; prompt injection not reliably solved.
- Simon Willison (2025): Not all AI-assisted programming is vibe coding – vibe coding defined as building software without reviewing the code.