Skip to content

AI-generated code security

Updated 26 September 2026 Reviewed by Teemu Malinen

What is AI-generated code security?

The distinct security risks that come with AI-written code. Models reproduce insecure patterns from their training data, and developers tend to trust the output more than they should. A large share of AI-generated code samples carry a common vulnerability. The answer is not to stop, but to scan, review and test AI output as rigorously as any other code.

Why it matters

The individual risk is well known; the organisational one is arithmetic. When a team ships two or three times as much code per sprint, the rate of insecure patterns need not rise for the absolute number of vulnerabilities to climb sharply, because there is simply far more code carrying them. That means more attack surface and more to review before anything reaches production. On top of that sits a provenance question the speed obscures: generated code can pull in dependencies nobody chose deliberately, including packages that did not exist until an attacker registered the name. Suggestions can also match existing public code. GitHub’s Copilot documentation says such matches typically occur in less than one percent of suggestions, and its code reference log names the license of the matching code, so the team can decide what attribution the code needs. None of this is an argument to stop. It is an argument to scale the guardrails with the output, so scanning, dependency checks and human review grow at the same pace as the code, instead of staying sized for the volume a team used to write by hand.

What makes AI-generated code a distinct security risk?

AI-generated code security is the practice of finding and preventing the vulnerabilities that code assistants and coding agents introduce. AI-written code carries a distinct risk for three reasons. First, models learn from public code, insecure patterns included, and reproduce them fluently. Second, the code looks finished, which lowers the reader’s guard. Third, a model can suggest dependencies that are outdated, vulnerable or entirely invented. OWASP’s Top 10 for LLM Applications 2025 names the third case directly: under “unsafe code generation” it lists models that suggest insecure or non-existent code libraries. Taken one at a time, these failure modes are old. Injection flaws, weak input handling and risky dependencies existed long before AI. What changes is the rate at which they arrive and how confident the code looks when it does.

What does the research show about vulnerabilities in AI-generated code?

Academic, policy and industry studies keep finding a high share of vulnerable output. In a study published at the IEEE Symposium on Security and Privacy 2022, Pearce and colleagues generated 1,689 programs with GitHub Copilot across 89 security-relevant scenarios and found about 40% of them vulnerable. Georgetown University’s Center for Security and Emerging Technology (CSET) tested five models in November 2024 and reported that almost half of the code snippets contained bugs that could potentially lead to malicious exploitation. Veracode’s 2025 GenAI Code Security Report, a vendor study covering more than 100 models in Java, Python, C# and JavaScript, found that 45% of code samples failed its security tests, and that models had become better at writing working code but no better at writing secure code. These are benchmark and lab results, and the rate varies a lot with the task and the language, but the direction is consistent across researchers.

Do developers write less secure code with an AI assistant?

In the best-known controlled study, yes. Researchers at Stanford University ran a user study with 47 participants on five security-related programming tasks in Python, JavaScript and C, published at ACM CCS 2023. Participants who had an AI assistant wrote significantly less secure code than those who did not, and they were also more likely to believe their code was secure. That combination is the hazard: weaker code paired with higher confidence means fewer second looks. The same study offers the counterweight. Participants who were sceptical of the assistant and kept refining their prompts produced code with fewer vulnerabilities. For teams, the lesson is that the review habit matters more than the tool, and that security review has to be designed for code that looks more polished than it is.

What are hallucinated packages and slopsquatting?

A hallucinated package is a dependency that a code model recommends but that does not exist in the package registry. Slopsquatting is the attack that exploits it: someone registers the invented name and publishes malicious code under it, so a developer who installs the suggestion installs the attacker’s package. The term was introduced in April 2025 by Seth Larson of the Python Software Foundation. The scale is measurable. A study presented at USENIX Security 2025 analysed 576,000 code samples from 16 code-generating models and found an average hallucinated-package rate of at least 5.2% for commercial models and 21.7% for open-source models, with 205,474 unique invented package names. OWASP describes the same scenario in its LLM09 entry: attackers collect commonly hallucinated names, publish malicious packages under them and wait for developers to follow the assistant’s suggestion.

How do you check AI-generated code for security issues?

Treat AI-generated code as untrusted input and run it through the same controls as any third-party contribution, sized for the higher volume. OWASP’s guidance on improper output handling puts it plainly: treat the model as any other user and take a zero-trust approach to what it produces. In practice that means four layers. Static application security testing (SAST) runs in the pipeline on every change, so injection flaws and unsafe functions are flagged before merge. Dependency scanning checks that every package the code imports exists, is the one intended and has no known vulnerabilities; for AI-suggested dependencies, confirming that the package is real and established is the specific defence against slopsquatting. Dynamic testing exercises the running code. And a human reviewer who understands the change approves it, with security-sensitive code (authentication, input handling, cryptography, access control) getting a second, deliberate look rather than a skim.

How does AI-generated code security map to OWASP and NIST?

OWASP and NIST both cover AI-generated code, though neither treats it as a separate discipline. The OWASP Top 10 for LLM Applications 2025 addresses it in two entries: LLM05 Improper Output Handling, which calls for validating model output before it reaches other components, and LLM09 Misinformation, which includes unsafe code generation and recommends secure coding practices, human oversight and training reviewers against over-reliance on AI suggestions. NIST’s Secure Software Development Framework (SP 800-218, version 1.1) has no AI-specific carve-out, and its practices apply whoever or whatever wrote the code: reuse well-secured components and verify third-party software (PW.4), review and analyse human-readable code (PW.7), and test executable code (PW.8). NIST SP 800-218A (July 2024) is a different thing: it adds practices for producers of AI models and AI systems and for acquirers of those systems; it does not cover secure use of AI coding tools as such.

In practice

A team celebrates shipping far more per sprint, and its security scanning, unchanged, keeps flagging the same small percentage of issues. The percentage is steady; the absolute count is not, because the denominator doubled. The fix is not complicated: make security review scale with throughput, not with headcount, or the extra speed just ships extra risk.

Frequently asked questions

Is AI-generated code less secure than human-written code?

Controlled research points that way. In a Stanford user study, developers with an AI assistant wrote significantly less secure code than those without one and were more confident in it. Benchmark studies also find a high share of vulnerable AI output, though rates depend heavily on the task and language.

What is slopsquatting?

Slopsquatting is registering a package name that AI models tend to invent and filling it with malicious code, so developers who install the suggested dependency install the attacker’s package. The defence is to verify that every AI-suggested dependency exists and is the established package you meant, before installing it.

Do newer AI models write more secure code?

Not reliably. Veracode’s 2025 testing of more than 100 models found that newer models wrote more functional code but no more secure code. Security still has to come from the process around the model, not from the model choice.

How do you scan AI-generated code for vulnerabilities?

Use the same tools as for any code: static analysis (SAST) in the pipeline, dependency scanning for vulnerable or non-existent packages, and dynamic testing of the running application. Automated scanning does not replace a human reviewer who understands what the change does.

Does OWASP have guidance on AI-generated code?

Yes. The OWASP Top 10 for LLM Applications 2025 covers it in LLM05 Improper Output Handling and LLM09 Misinformation, which names unsafe code generation and hallucinated packages as risks and recommends secure coding practices and human oversight.

Sources

Otto Sunnari, myynti ja kumppanuudet, Sofokus / Otto Sunnari, Sales and partnerships at Sofokus

Ready to start leveraging AI?

Call, email, or book a time straight from my calendar.

Otto Sunnari

Sales and partnerships