Research

We re-scanned 393 AI-built repos without using AI. 1 in 8 shipped a critical flaw.

VibeSafe · October 3, 2026 · 8 min read

In September we published what a language model found across 400 AI-built repositories. The obvious criticism of that study was fair: a model's judgement is not reproducible, and you cannot audit it. So we did the whole thing again with deterministic rules, where every single finding traces back to a rule and a line number. Here is what changed, what held, and the number that went down.

Why run it twice

The first pass asked a model to read each file and describe what was wrong with it. That produces rich, readable findings — and it has two problems if you want to publish a statistic. Run it again and you may get a slightly different answer. And when someone asks "how exactly did you count that?", the honest reply is "the model decided", which is not a method anyone can check.

The second pass uses 26 deterministic rules. Same frozen corpus, same files, no model involved. Every finding carries a rule id and a line number, so any number below can be traced to the exact pattern that produced it. Run it tomorrow and you get the same answer.

The corpus

400 public repositories built with Lovable, Bolt and v0, selected by build artefact rather than by README text — a repo qualifies because it contains .lovable, lovable-tagger or .bolt, not because someone wrote "built with Lovable" in the readme. One repo per owner, forks excluded, up to three files each, frozen on 5 September 2026 so the sample cannot drift.

Of 1,185 files, 1,163 were readable across 393 repos. The missing 22 are files deleted from GitHub since the corpus was frozen. The whole run takes about 90 seconds.

What the rules found

309 of 1,163 files — 27% — had at least one finding. 58 files had something critical, spread across 49 repos. That is 1 in 8 repositories carrying at least one critical flaw, by rules alone.

Cookie set without Secure or SameSite       111
Access-Control-Allow-Origin: *               74
dangerouslySetInnerHTML from a variable      60
RLS policy written as using (true)           39
SQL built by template interpolation          14
innerHTML assigned from a variable            6
Google API key committed in source            3
eval(), weak token randomness, no RLS         7

The finding that matters most

39 files contained a row-level security policy written as using (true).

This is worse than it looks, because it is the signature of someone who tried. They read that Supabase tables need row-level security. They enabled it. Then they wrote a policy that returns true for every row — which passes every check, for every user, every time. The dashboard shows RLS as enabled. The table is wide open.

-- Enabled, and completely ineffective
alter table orders enable row level security;
create policy "read orders" on orders for select using (true);

-- What it needs to say
create policy "read own orders" on orders
  for select using (auth.uid() = user_id);

An AI builder will happily generate the first version when asked to "add RLS", because it satisfies the literal request. Nothing in your app breaks. Nothing warns you. Your customers' rows are readable by anyone holding the public key, which ships in your frontend bundle by design.

The number that went down, and why that is the point

Our September post reported that 59% of scanned repos had a critical issue. This pass says 12%. Both are true, and the gap is the most useful thing in this post.

On the 115 files both passes read, the model flagged something in 113. The rules flagged something in 32. They agreed on the verdict only 30% of the time. Rules catch syntactic patterns — a cookie without a flag, a wildcard CORS header, SQL glued together with a template literal. They are blind to everything that requires reading intent: an endpoint that never checks who is asking, user input reaching a sink three functions later, an error handler that leaks a stack trace to the browser.

So treat 12% as a floor, not an estimate. It is the share of repositories where a pattern-matching rule, with no understanding of the code, could prove something was wrong. The real figure is higher. We publish the floor because it is the number we can defend line by line.

One more thing worth noticing

Only three secrets turned up in source files across 1,163 files. That sounds like good news and is not. Neither pass opens .env files — and our earlier study found 18.5% of these repositories had committed one. Secrets in these projects are not pasted into components. They sit in the environment file that got committed on the first push.

How to check your own app in about a minute

Check your own app free →

10 scans a month · No card needed

Questions

Can I see the raw data?

Yes — ask and we will send the per-file and per-finding CSVs. Every row carries the repository, the file, the rule id and the line number, so any finding can be re-checked against the public repository it came from.

Are you naming the repositories?

Not publicly. These are real projects by real people, and most of the owners do not know. The aggregate is the useful part; publishing a list of vulnerable apps would just be handing someone a target list.

Does this mean Lovable and Bolt are insecure?

It means code generated quickly and shipped without review tends to carry the same handful of defects, which has been true of every generation of development tooling. The sample is heavily Lovable-weighted — 849 of the files — and only 24 came from v0, which is far too few to say anything about v0 at all.

Related: