In this article
Vibe coding: a threat or an opportunity?
Both. What decides it is what the output is used for. Vibe coding means producing software with AI without reviewing the source code. In an experiment it can be a sensible choice: you get an idea tested fast.
If that same output moved into even limited ongoing use, an AI review gap would appear — meaning that the proper verification of functionality, security and other critical areas would be left undone.
Vibe coding is, then, a new model for software development worth taking seriously — but only when applied correctly. This article walks through where vibe coding fits and where it earns a place, even in a professional’s toolkit.

Key takeaways
- Vibe coding means producing software with AI so that the source code isn’t reviewed.
- For experiments, vibe coding is a good fit. You get an idea tested fast, and the code doesn’t end up in continuous use.
- When an unreviewed output moves into limited use or production, an AI review gap appears. Functionality, security and other critical areas then go unverified.
- Experience in software development decides whether the review level can be matched to the purpose.
- When you buy software development, the key thing isn’t whether AI is used. What matters more is the process — for example, how the quality of the produced code is ensured.
Why does vibe coding mean different things to different people?
In its original sense, vibe coding refers to a tightly defined way of working: the source code isn’t read. Andrej Karpathy, who coined the term, described a way of working where changes are accepted without inspecting the diffs and errors are fixed by trial and error. He himself limited the phenomenon to low-stakes work. Simon Willison has put it more precisely: “building software with an LLM without reviewing the code it produces.” Note that if a software professional reads the code, tests it and takes responsibility for the result, it isn’t vibe coding, even if the model wrote every line.
In everyday use, though, the term has stretched to cover almost all AI-assisted software development. Collins named it the word of the year for 2025, and in the market it’s often used to describe development done with AI regardless of who does it or how they work.
From a decision-maker’s point of view, this creates two different ways to misunderstand it:
- professional AI-assisted development can be seen as a mere throwaway experiment, or
- unreviewed AI-generated source code can be wrongly accepted as production-ready.
Review decides — not the person, not the tool
The confusion comes from lumping three independent things behind the colloquial term “vibe coding”: the person doing the work, the level of autonomy of the AI agents, and the level of review of the end result. But vibe coding has nothing to do with who does it or how autonomous the agent is. It is the watershed on the review axis: the point where the source code isn’t read.
Because of this, vibe coding as a way of programming can carry either very little risk or very high risk. Nor is greater autonomy in agentic coding automatically better. The full breakdown, with its four types of author, lives in the glossary term AI-assisted software development. For this article, one conclusion is enough: risk is decided by review relative to how the output is used.
Where does unreviewed code become a risk?
Unreviewed code becomes a risk the moment you step outside experimentation. The “AI review gap” below sums this up in a simplified picture. Its horizontal axis is the output’s purpose, split into three areas: experiment, limited use and production use. Its vertical axis is review coverage, split into three levels: no review, spot checks, and full review with testing and acceptance criteria.

At a high level, the chart answers the question of which review level makes each purpose possible. The top row matches traditional work by an experienced software professional, where the source code is read, tested and accepted against criteria agreed in advance. In this model, risk stays under control from experiments all the way to production.
In the middle row, not all source code is read; the checks focus on certain areas, spot checks or the like. This model can be perfectly sufficient for, say, an experiment or limited use, but in production the risk level would rise significantly.
The bottom row is vibe coding. Here, the risk is low only in prototype-style experimentation, regardless of whether the author is a layperson or a professional. Moving a vibe-coding output into even limited use raises the risk level significantly.
The review-gap boundary is marked with a dashed line: to its right, review always lags behind what the purpose would actually require.
So vibe coding isn’t unequivocally a bad way to produce software, but it’s usable in only one column of the chart if you want to keep risk levels under control.
The AI review gap applies an established principle: the depth of inspection is scaled to how critical the subject is. In safety-critical design, aviation’s DO-178C grades the number of required objectives by the system’s assurance level, and IEC 61508 and ISO 26262 classify methods as recommended or highly recommended based on the risk level.
In ordinary software development, the same principle shows up in review practices and automated inspection processes, where the most thorough assessment is aimed at the riskiest changes.
Vibe coding may well be the right choice when the output stays in the experiment column: a prototype for testing whether an idea works, a one-off helper script of your own, or a demo whose code can be thrown away afterwards. Here speed and making things concrete are worth more than code quality. The condition, of course, is that the output doesn’t leak into wider or longer-term use.
How will AI change your business? The AI guide for business leaders explains how to lead AI adoption starting from business goals and how to recognise your organisation's current level.
Read the AI guideWhy does an experiment quietly turn into production?
A review gap rarely comes from a decision. More often, an experiment built with vibe coding is first taken up by a few people, found useful, and finally creeps into the team’s everyday work without anyone ever making a decision to put it into production.
The shift can be hard to spot if you’re not a software professional. Applications usually have two sides: the visible user interface, and the security, software architecture and the like that stay under the hood.
The first is visible right away in the first demo: how the application works. The second only surfaces under load, in maintenance, in a security review, or as use expands.
Looking finished says nothing about production readiness. In a peer-reviewed experiment by Stanford researchers, “participants who used an AI assistant wrote less secure code than those who worked without one.” Trust in the output grew exactly where there was reason to be less trusting.
Can professionals vibe code too?
Of course, and sometimes it can even be the most sensible choice (for instance, when you want a quick, throwaway, visible prototype). The difference is that a professional knows how to choose an adequate review level, whereas someone who can’t program does not.
A small caveat is in order, though. In May 2026, Simon Willison wrote that as agents grow more reliable he has started leaving some production code, too, unreviewed line by line, and that “the line between vibe coding and professional agentic work is blurring in a way I don’t like.”
The observation is a human one, and it underlines that responsibility for the outcome stays with the person, whether or not the code is read. At the same time it shows that professional software development is changing fast at the upper end of the review axis as the technology develops.
What does this mean for someone buying software development?
When you buy software development, the question “do you use AI?” tells you almost nothing about the risk.
For the buyer, the decisive question isn’t primarily whether AI is used, but how professional a process has been built around the code being produced.
So it’s worth finding out, among other things, how AI-generated code is reviewed, who is responsible for the result, and how the use of AI is woven into the software development process.
Frequently asked questions
Unreviewed code in production is the serious-risk square of the AI review gap. Code meant for production has to be read, tested and accepted against agreed criteria — at which point it’s no longer vibe coding, even if AI wrote every line. Vibe coding suits experiments, not production.
Yes, when the purpose is limited to experimentation — testing an idea, say, or a throwaway demo. In limited use and in production, unreviewed code is a significant or serious risk, so there the code has to be reviewed properly, tested and accepted against agreed criteria.
A person. Responsibility doesn’t transfer to the AI model, whether or not the code was read before it was accepted. When you buy software development, it’s worth confirming who is accountable for the result and how AI-generated code is reviewed before it ships.
Not on its own. A machine checks what rules can catch — whether the tests pass and whether the code contains known vulnerabilities. A person checks whether the change does what it’s meant to and whether the remaining risk is acceptable. Both are needed for production code.
About the author
Founder & Chairman
Teemu Malinen
Teemu writes about digital business trends, modern company culture and startup investments.
Related terms
From the glossary
Short definitions for the concepts this article leans on.
- Vibe coding Building software by describing it in plain language, without reviewing the code.
- Large language model (LLM) The model type behind chatbots and copilots.
- AI-assisted software development Building software with AI in the loop.
- Agentic coding AI agents carry out coding tasks under developer direction.


