A Dialogs explainer · June 2026
If you've built an app with an AI builder — or you're about to bet a business on one — there's a question you need answered before real customers, real data, or real money touch it: is it actually safe to ship? Independent security research has been looking hard at exactly that, and the honest answer is worth ten minutes of your time.
First, credit where it's due: AI app builders — Lovable, Base44, Bolt, Replit, Cursor, and the rest — are genuinely remarkable. A non-technical founder can describe an idea in plain language and have a working application an hour later. That's not hype; it's a real shift in who gets to build software, and it deserves the excitement it's getting.
But "works in a demo" and "safe in production" are two different bars, and the gap between them is wider than most people building this way realize. This isn't an argument against AI builders. It's a look at what the research has actually found when those builds meet the real world — and what it means for you if you're planning to put a business on top of one.
The case that opened the conversation
The first widely documented example became a formal vulnerability: CVE-2025-48757, rated 8.26 on the standard severity scale. In May 2025, a security researcher built an automated scanner and pointed it at 1,645 applications featured on Lovable's own public showcase. About 10.3% of them — 170 apps — had a critical database-exposure flaw, across more than 300 vulnerable endpoints.
The flaw was a misconfiguration of Row-Level Security (RLS) — the database rules that decide which user is allowed to see which rows of data. Without correct RLS policies, the public key embedded in the app's front-end let anyone query the database directly and pull entire tables. The exposed data reportedly included names, emails, phone numbers, home addresses, financial records, and developer API keys. The issue was first disclosed by researcher Matt Palmer; an engineer at a major tech company independently reproduced the attack, and the formal CVE was published at the end of May 2025.
Two things make this the right anchor for the whole topic. First, the 10.3% figure is precise and verifiable — but it's also narrow: it measured one platform's curated showcase against one class of vulnerability. It is not a claim that one in ten of every app ever built with AI is exposed. Second, and more importantly, it turned out not to be a one-off.
The pattern repeated — across apps and across platforms
The same shape of problem kept surfacing on different platforms and at higher stakes:
- February 2026 — a featured education app. Researcher Taimur Khan examined a Lovable-built application used to create exam questions and manage grades across universities including UC Berkeley and UC Davis, plus K-12 schools. He reported 16 vulnerabilities (6 critical) and an exposure of roughly 18,700 user records, with the app permitting unauthenticated bulk email, account deletion, and grade manipulation. The Register covered the disclosure.
- July 2025 — a platform-wide auth bypass. Wiz Research found that Base44 (since acquired by Wix) had an authentication-bypass flaw: a publicly visible app ID was enough to create a verified account on private apps — the equivalent of getting into a locked building by knowing the room number. Wix fixed it within 24 hours of disclosure.
- April 2026 — a mass exposure. A researcher disclosed a breach affecting Lovable projects created before a certain date, in which a free account could reportedly read other tenants' source code, database credentials, and customer data. It was widely described as the platform's third major security event in just over a year.
No single incident proves much on its own. The point is the consistency — the same root causes (missing access controls, shallow authentication, exposed secrets) recurring across tools, builders, and use cases.
The bigger picture: research at scale
Individual incidents make headlines; systematic scans tell you whether they're typical. Several independent efforts now point the same direction.
Escape.tech (October 2025) ran the most comprehensive scan to date: 5,600 publicly deployed vibe-coded applications across multiple platforms. They found more than 2,000 high-impact vulnerabilities, over 400 exposed secrets (API keys, credentials, tokens), and 175 instances of exposed personal data — including medical records and bank account numbers — all in live production systems. The researchers stressed their method was passive and conservative, calling the results a lower bound rather than the full extent of the risk. Investors took the thesis seriously enough that Escape raised an $18M funding round on it in early 2026.
RedAccess ("Shadow Builders," 2026) widened the lens further, identifying roughly 380,000 publicly accessible assets built with tools like Lovable, Base44, Replit, and Netlify. About 1.3% of them — on the order of 5,000 — held sensitive corporate data. Notably, this one was independently verified by both Axios and Wired, which is rare for vendor research and worth weighting accordingly.
Controlled studies point the same way. A Carnegie Mellon analysis found that while about 61% of AI-generated code functions correctly, only around 10.5% passes a security review — fewer than 11 of every 100 snippets meeting basic security standards. Veracode's GenAI Code Security Report found roughly 45% of AI-generated code samples introduced a known (OWASP Top 10) weakness. A December 2025 study by Tenzai had five widely used AI coding agents each build the same applications from identical prompts and found 69 vulnerabilities across 15 apps, with every tool introducing server-side request forgery and none setting security headers by default.
The risk extends to dependencies, too. Endor Labs analyzed over 10,000 repositories in late 2025 and found that only about one in five of the package versions recommended by AI assistants were safe — the rest were either hallucinated or carried known vulnerabilities. And IBM's Cost of a Data Breach research reported that a meaningful share of organizations — roughly one in five — have already experienced a breach linked to AI-generated or "shadow" AI code.
Different teams, different methods, consistent conclusion: insecure-by-default is currently the norm for AI-built software, not the exception.
Why this happens — and why it isn't a moral failing
It would be easy to read all this as "AI writes bad code." That's not quite right, and the real explanation matters because it tells you exactly where the risk lives.
AI builders are optimized to make something work. Security requirements, by contrast, are mostly unstated — nobody types "and make sure users can't read each other's data" into the prompt, because an experienced engineer would never need to be told. Those non-functional requirements (correct access controls, deep authentication, input validation, security headers, secrets kept off the front-end) are precisely the things that get skipped when the only goal expressed is the feature itself.
A useful mental model, popular among engineers who use these tools heavily: treat AI-generated code like a very fast junior developer who has no threat model and was never told yours. You wouldn't merge a junior's work touching authentication or payments without reading it. The same caution applies here — except the person shipping an AI build often isn't able to read it, which is the heart of the problem. The builder is, by design, non-technical. They can't easily tell a secure app from an insecure one that looks identical in the demo.
This is also why the failure tends to be structural rather than incidental. A normal bug is a logic error. These are whole security layers that were simply never built, because nothing in the conversation asked for them.
What the platforms are doing — and where responsibility sits
To their credit, the platforms are responding. Wix patched the Base44 flaw within a day. Lovable shipped built-in security scanners (a basic and a deep scan) that check schema and RLS configuration, flag exposed secrets, and audit dependencies.
But two things are worth reading carefully. First, Lovable's own documentation is explicit that these tools don't replace a thorough security review and that the user remains responsible for ensuring an app is safe for its use case — especially when it handles sensitive data. Second, security researchers broadly agree that "secure by default" for AI code generation is a multi-year problem, not a next-release one, because it depends on changing what the underlying models are trained on.
The practical conclusion is unavoidable: for now, security responsibility sits with whoever ships the app — not with the tool that generated it. If that's a non-technical founder, the responsibility lands somewhere they can't realistically discharge alone.
What this means if you're building a business on an AI app
None of this is a reason to stop using AI builders. It's a reason to be clear about the moment things change. The instant real customers, real data, or real revenue start flowing through what you built, it stops being an experiment and becomes a production system — and it deserves a production security posture.
In practice, a genuine readiness review for an AI-built app looks at, at minimum:
- Access controls (RLS/authorization): does the database actually prevent one user from reading or modifying another's data — verified by testing, not just by RLS being "enabled"?
- Authentication depth: does login hold up under real conditions — multiple users, token expiry, password resets, rate limiting — not just a single happy-path test?
- Secrets: are API keys and credentials kept server-side, not embedded in the front-end bundle where anyone can extract them?
- Dependencies: are the third-party packages free of known vulnerabilities, and do they actually exist?
- Input handling and headers: is user input sanitized, and are basic security headers set?
- Ongoing monitoring: because a clean scan today says nothing about a new dependency vulnerability next month.
An answer you can't trust is worse than no answer, because you'll make decisions — and stake your reputation — on it. The goal isn't fear. It's knowing what you have before you bet on it.
Where Dialogs fits
We think this gap is entirely predictable, and entirely fixable. AI did the part it's brilliant at: it got you to a working idea faster than anyone could have a year ago. The remaining work — making it safe, making it real inside your actual business, and keeping it that way — is a different job, and it's the one we've done for 30 years.
That's what our Readiness review is for: a senior engineer tells you, in writing, what's safe, what's real, what's missing, and what it takes to go live. From there, Last Mile makes it production-ready inside the systems you already run, and Steward keeps it healthy for the long haul. If you've built something with AI and you're not sure whether it's safe to ship, that's exactly the conversation we're built for.