Skip to content

Blog

Is AI-generated code secure? What to ask the software house

Veracode’s 2026 report and Stack Overflow’s 2025 survey show where AI-generated code gets security wrong. What to ask a software house that uses AI assistants to write a web app: checks, review, user input, logs, ISO/IEC 27001.

Christian Ascone

Not always. In the tests of Veracode’s 2026 GenAI Code Security Report, about 44% of AI code generation tasks introduced a security vulnerability, while the syntax, according to the same report, is correct in almost 100% of cases. It pays to ask the software house which checks the code goes through before production and who reviews it.

The figures come from two sources: Veracode’s report, which has AI models write code and measures its security, and the Stack Overflow Developer Survey 2025, a survey of developers.

What did Veracode measure in the 2026 report?

How often an AI model, given a programming task, produces code with a known vulnerability. In the 2026 report Veracode tested 11 new models on 80 tasks. Counting the earlier editions too, the models tested over four years number more than 100.

The main result is in the blog post presenting the report, of 28 July 2026: “roughly 44% of AI code generation tasks introduced a risky security vulnerability in tests”. The average security pass rate across the models is 56%, almost the same as the 55% of the first report, the 2025 one.

The 44% refers to the tasks of Veracode’s benchmark, that is, programming tasks prepared for the test, and it should be read this way: it measures how often the models get those tasks wrong, and says little about how much of the code written with AI in companies today is vulnerable. The comparison between the two editions says that in a year the average moved by one point.

Is syntactically correct code also secure?

Not necessarily. According to Veracode, today’s models generate syntactically correct code in almost 100% of cases, and security follows a different path: “Syntax is effectively solved. But secure coding is not following the same curve”.

The report’s blog post also has the sentence that concerns whoever receives the code: code that compiles, runs and looks clean can still introduce exploitable weaknesses.

Where does AI-generated code go wrong most?

On cross-site scripting and log injection. The report breaks the results down by type of vulnerability, and the gap is large:

Type of vulnerability Average security pass rate
Cryptographic algorithms 87%
SQL injection 83%
Cross-site scripting (XSS) 15%
Log injection 12%

The data are in Veracode’s blog post of 28 July 2026.

Cross-site scripting arises when data entered by a user comes back into a page without being handled; log injection when data entered by a user ends up in the logs as it is. In a web app this data comes from forms, comments, search fields and usernames.

The first report, the 2025 one, pointed the same way. In the 2025 report, across more than 100 models in Java, Python, C# and JavaScript, 45% of code samples failed the security tests and introduced OWASP Top 10 vulnerabilities. On cross-site scripting (CWE-80) the tools failed in 86% of the relevant samples.

The language matters too. In the 2026 report Java comes last, with an average pass rate of 30%.

Is it enough to choose the best AI model?

The choice of model matters, and on its own it leaves many tasks uncovered. In the summer 2026 dataset the best model reaches 68%, while six models out of eleven sit between 50 and 53%, according to the report page. Even the best gets almost one security task in three wrong.

Specialisation changes little. Models trained for code have an average pass rate of 51%, general-purpose ones 52%. Veracode sums it up like this: “Being trained to write code faster does not mean writing it safer”.

Which AI assistant a supplier uses is therefore a useful question, to ask together with the one about the checks the code goes through after it has been written.

Do developers trust AI-generated code?

Less than they use it. In the Stack Overflow Developer Survey 2025 84% of developers use or plan to use AI tools. 46% do not trust the accuracy of their output, against 33% who do; only 3% trust it highly.

The most cited frustrations point the same way. 66% name AI solutions that are “almost right, but not quite”. 45% say that debugging AI-generated code takes more time.

The survey is from 2025, and the question about trust concerns the accuracy of the output in general. The figure to keep is the ratio between the two groups: among developers, those who distrust the output of these tools outnumber those who trust it.

What to ask a software house that uses AI to write code?

How the code is checked before it goes into production. These are the questions that, having read the sources, we would ask any supplier of a web app:

  • Which security checks does the code go through before release? Static analysis, security testing, and at what point in the work: on every change or only at the end. If the answer is that the code is tested before delivery, ask what is tested and with which tools.
  • Who reviews the code, and what do they look for? Ask whether there is a review dedicated to vulnerabilities, on top of the one on functionality.
  • How is user input handled? This is the question about cross-site scripting. Ask for a concrete example on a form in your project.
  • What ends up in the logs, and how? This is the question about log injection, the type of vulnerability with the lowest pass rate in the 2026 report.
  • Which language will it be written in, and with which specific checks? In the 2026 report Java is the language with the lowest average pass rate: if your project is in Java, the question weighs more.
  • In which parts of the project do you use AI assistants? It tells you where to focus the other questions, and whether whoever reviews the code knows which parts were generated.

The answers, given in writing, can be compared across different suppliers. For each one it pays to ask for an example taken from a project already delivered.

What is a software house’s ISO/IEC 27001 certification for?

It says that the company has an information security management system verified by a third-party body, with an expiry date. You can ask for the certificate, the body that issued it and the expiry date. On the code of a single project the questions in the previous section still need asking.

dotenv also uses AI assistants to write code, and the same questions can be put to us. Our certification is UNI CEI EN ISO/IEC 27001:2024, issued by Dasa-Räegister S.p.A., a body accredited by ACCREDIA, expiring on 15 December 2028.

Where do you start?

With the six questions, put in writing to your current supplier or to the ones you are evaluating, before signing.

If you are evaluating a new web app, the page on web app development shows what kind of web apps we build, with three projects, and from there you can ask for a first meeting. If for now you need a document to attach to the evaluation, the certifications page has the standard, number and expiry date of the ISO/IEC 27001 certificate.

A similar problem, in your company?