Software has a new co-author, and it doesn’t read its own work.
In Y Combinator’s Winter 2025 batch, 25% of startups had codebases that were 95% AI-generated. GitHub now reports that 46% of all newly committed code is AI-generated. The tools driving this shift have collectively turned a once-specialized skill into something a marketing manager can do over the weekend. On Lovable’s platform alone, 200,000 vibe-coded projects are built per day.
That’s the part that’s going well. The other part: Veracode’s 2025 GenAI Code Security Report tested over 100 large language models across 80 real-world coding tasks and found that AI-generated code introduced security flaws in 45% of cases. Cross-site scripting failed 86% of the time. Java sat at a 72% security failure rate. The pattern holds across model sizes, vendors, and release dates, so this isn’t something the next model will quietly fix.
That’s the gap secure code review is meant to close. And it’s the reason the practice itself needs an update.
This guide covers what secure code review actually is, how it changes when AI writes most of the code, the most common vulnerabilities showing up in vibe-coded apps, what a modern review process looks like, and how teams can build a workflow that keeps pace with AI-accelerated development without grinding it to a halt.
What Is Secure Code Review?
TL;DR: Secure code review is the practice of analyzing source code for security vulnerabilities before that code reaches production. It’s distinct from a standard code review because the goal isn’t readability or correctness, it’s identifying weaknesses an attacker could exploit.
Secure code review (also called security code review or code security review) is a structured examination of source code with one job: find the security flaws before they ship. It blends automated tools, like Static Application Security Testing (SAST) and Software Composition Analysis (SCA), with human expertise to surface vulnerabilities that automated scanners alone tend to miss.
A regular code review checks whether the code works, if it’s readable, and if it follows team conventions. A secure code review asks a different question: can someone who shouldn’t have access get access? That includes things like authentication flaws, injection paths, insecure data handling, hardcoded secrets, and logic errors that create exploitable conditions.
Why Secure Code Review Exists
The shift-left movement in DevSecOps is built on a simple economic argument. Catching a vulnerability while a developer is still writing the code is dramatically cheaper than catching it in production. Black Duck reported that the cost of fixing a bug found during implementation is roughly 6x the cost of fixing one identified during design, and the cost in the testing phase climbs to about 15x. Those multipliers compound when the issue is a security flaw, because the cost isn’t just engineering time, it’s potentially data breach exposure, regulatory fines, and customer trust.
Secure code review exists because waiting for a penetration test (or worse: a real attacker) to find your application’s weaknesses is the most expensive way to find them.
How Teams Identify Vulnerabilities Before Production
A mature secure code review process layers several techniques:
- Automated SAST scanning runs against every commit or pull request, flagging known-bad patterns in the code itself.
- Software Composition Analysis scans third-party dependencies for known vulnerabilities.
- Secrets scanning catches API keys, tokens, and credentials before they hit version control.
- Manual review by engineers or security professionals evaluates business logic, authentication flows, and architectural decisions that scanners can’t.
- Runtime validation confirms whether flagged vulnerabilities are actually exploitable in the running application.
The combination matters. Automated tools generate volume. Human reviewers (or sufficiently sophisticated AI agents) generate context. Neither alone is enough, and that gap is exactly where most vulnerabilities slip through.
Why AI-Generated Applications Create New Security Risks
TL;DR: AI coding assistants optimize for code that runs, not code that’s secure. When developers ship that code without reviewing it, the result is a measurable spike in vulnerabilities.
AI coding tools are remarkably good at producing functional code. They’re noticeably worse at producing secure code, and the way most teams use them makes the problem worse.
There are four structural reasons for this:
1. AI coding tools optimize for the wrong thing.
When a developer asks an LLM for “a login form” or “a database query,” the model generates the code that statistically matches that request based on its training data. That training data includes billions of lines from open-source repositories, including a lot of insecure code. As Veracode’s CTO Jens Wessling put it, “developers do not need to specify security constraints to get the code they want, effectively leaving secure coding decisions to LLMs.” Default outputs are functional. They are not, by default, hardened.
2. Developers accept code they don’t fully understand.
The 2025 Stack Overflow Developer Survey found that only 33% of developers trust AI-generated code for accuracy, while 46% actively distrust it, and 81% worry about security and privacy. And yet adoption keeps climbing because the velocity gains are real. The result is a workforce shipping code it doesn’t fully trust, often without time to verify it line by line.
3. Non-technical builders are now in production.
Vibe coding platforms like Lovable, Replit, and Bolt let designers, marketers, and founders ship apps without writing a single line of traditional code. Recent research from security firm Red Access found 380,000 publicly accessible assets built with these tools, including roughly 5,000 containing sensitive corporate data. When the person shipping the app doesn’t know what role-based access control is, asking them to review their own code for security issues is a non-starter.
4. Prototype code keeps becoming production code.
It always has, but the speed has changed. A weekend project goes live on Monday because customers showed up. The “we’ll harden it later” promise rarely gets kept, and now there’s a CVE in production that nobody scoped, designed, or reviewed.
The Tools Behind the Shift
The AI coding ecosystem has consolidated around a handful of major players, each contributing to the new reality:
- GitHub Copilot has crossed 20 million cumulative users and sits inside 90% of Fortune 100 companies. GitGuardian found public repositories using Copilot had a 6.4% secret leakage rate.
- Cursor hit $1B annualized revenue in under 24 months and has been adopted as a standard AI coding tool at NVIDIA, per Opsera’s analysis.
- Claude Code powers a growing share of agentic development workflows. GitGuardian’s research found that commits co-authored by Claude Code leaked secrets at roughly double the GitHub baseline rate.
- Replit, Lovable, and Windsurf are the consumer-facing edge of the trend, where non-developers ship directly to production. Lovable alone has eight million users and a $6.6 billion valuation.
These tools aren’t the problem. The absence of a security layer between them and production is.
Common Vulnerabilities Found in AI-Generated Code
TL;DR: AI-generated, or vibe-coded, apps tend to fail in predictable ways: leaked secrets, broken authentication, missing input validation, vulnerable dependencies, and misconfigured infrastructure. The same patterns appear across platforms because they all train on the same insecure code.
When security researchers actually look at what’s getting shipped from AI coding tools, the same vulnerabilities surface again and again. Here’s where most of the damage clusters.
1. Hardcoded API Keys and Secrets
GitGuardian’s State of Secrets Sprawl report identified 28.76 million new secrets exposed in public GitHub commits in 2025, a 34% year-over-year increase. AI-assisted commits leaked secrets at roughly twice the baseline rate. AI coding tools generate working code that includes hardcoded credentials because the example code in their training data also included hardcoded credentials. Developers copy the pattern, replace the placeholder, and commit. The Cyber Express recently identified more than 5,000 GitHub repositories containing hardcoded OpenAI credentials and 3,000 production websites exposing API keys directly in client-side JavaScript.
2. Broken Authentication and Authorization
This is the headline failure mode of vibe coding. In May 2025, Lovable was found to have 170 out of 1,645 audited applications, more than 10%, with critical row-level security flaws that allowed anyone to access user names, emails, financial information, and API keys. The vulnerability earned a CVE designation. AI tools generate authentication code that looks correct but skips server-side enforcement, gets the access control logic inverted, or treats public/private settings as a UX problem rather than a security one.
3. Insecure Database Configurations
When AI tools wire up backends, they tend to default to overly permissive configurations. Missing Row Level Security on Supabase. Database credentials embedded directly in client code. Production databases without read-only constraints. The Moltbook incident exposed 1.5 million API keys due to missing Row Level Security, a single misconfiguration that the platform never enforced.
4. Vulnerable Open-Source Dependencies
Black Duck’s 2025 Open Source Security and Risk Analysis report found that 86% of commercial codebases contain vulnerable open-source components, and 81% contain high- or critical-risk vulnerabilities. The average application now contains 911 open-source dependencies. AI coding tools accelerate this by importing libraries without checking version, maintenance status, or known vulnerabilities. Black Duck noted that only 77% of dependencies could be identified via package manager scanning, suggesting the rest were introduced through AI coding assistants and other indirect means.
5. Missing Input Validation
Cross-site scripting (XSS) is the canonical example, and Veracode’s research found AI tools failed to defend against it in 86% of relevant code samples. Log injection clocked in at 88%. SQL injection, path traversal, command injection, all of it traces back to the same root cause: the AI doesn’t sanitize input by default unless explicitly prompted to.
6. Exposed Admin Routes and Debug Endpoints
The Tea app’s admin routes were left unlocked, exposing user data to anyone who found the endpoint. This is what happens when AI generates a working admin panel and nobody reviews whether it’s authenticated. Debug endpoints are a similar pattern, useful in development, dangerous in production, and rarely cleaned up automatically.
7. Overly Permissive Cloud Permissions
When AI tools provision cloud infrastructure, they tend to err toward “make it work” rather than “make it least-privilege.” Wide-open S3 buckets. IAM roles with admin policies. Service accounts with no scope. Once those defaults are deployed, fixing them later means refactoring whatever was built on top of them.
The throughline across all seven categories: these aren’t novel vulnerability classes. They’re the OWASP Top 10 with a velocity problem. AI didn’t invent insecure code, it just industrialized the production of it.
How Secure Code Review Works for Vibe-Coded Apps
TL;DR: Reviewing AI-generated applications requires the same layers as traditional review, plus an extra emphasis on architectural reasoning, exploitability validation, and inline developer feedback. Running a scanner isn’t enough.
Secure code review for vibe-coded, or AI-generated, applications follows the same general arc as traditional review, but a few stages take on outsized importance when the code wasn’t written by a person who understands it.
Reviewing Architecture and Business Logic
This is where most automated scanners fall down. A SAST tool can flag a SQL query that looks injectable, but it can’t tell you whether your tenant isolation logic actually isolates tenants. AI-generated code tends to produce architectures that look reasonable at the function level but break under adversarial reasoning. A reviewer, human or sufficiently capable agent, has to ask the questions the AI didn’t think to ask. Who can call this endpoint? What happens if they call it with someone else’s user ID? Is the authorization enforced server-side, or just hidden in the UI?
Identifying Exploitable Attack Paths
A list of vulnerabilities is not a list of risks. The Lovable incident is instructive: scanners would have flagged the missing Row Level Security as a misconfiguration, but the risk was that an unauthenticated request could pull every user’s email and API keys from the database. Modern secure code review focuses on the chain, not just the link. SafeHill’s approach builds attack paths from individual findings, then validates which ones are reachable in the running application.
Reviewing Authentication and API Security
Auth is the single most common failure point in vibe-coded apps. Review here means checking session management, token validation, API authorization, role enforcement, and password handling. It also means checking what isn’t enforced. AI-generated apps regularly skip server-side checks because the client-side check “looks like enough.”
Scanning Dependencies and Infrastructure
Software Composition Analysis catches vulnerable open-source packages. Infrastructure-as-Code scanning catches misconfigured cloud resources. Both are table-stakes for any modern review process, and both are particularly important for AI-generated apps because the AI imported dependencies and provisioned infrastructure that the developer probably didn’t audit.
Human Validation vs. AI-Only Scanning
This is the point worth pausing on. A pure AI scanner finds things that look suspicious. A pure human reviewer is too slow and expensive to scale to AI-velocity development. The teams getting this right combine deep static analysis with adversarial validation, then hand off the residual ambiguous findings to a human or expert AI agent for context. Security review is not “running a scanner.” Anyone telling you otherwise is selling you noise.
Automated Security Scanning vs. Human Secure Code Review
TL;DR: Automated scanners are essential for coverage but produce overwhelming noise. Human or AI-augmented review is essential for context but doesn’t scale alone. The teams winning at modern AppSec combine both, and prioritize validated findings over raw alert volume.
Automated security scanners are good at what they’re good at. They scan fast, run on every commit, and catch known-bad patterns reliably. They’re also notoriously noisy.
What Automated Scanners Are Good At
Pattern-based detection is genuinely useful. SAST tools find SQL injection, XSS, hardcoded secrets, dangerous function calls, and known-vulnerable dependency versions at scale. They’re cheap to run, easy to integrate into CI/CD, and provide a baseline of coverage no team can match manually.
Limitations of AI-Only Scanning
Here’s the catch. NIST’s SATE benchmarks have consistently shown that rule-based SAST tools produce false positive rates ranging from 30% to over 70% depending on the tool and codebase. A 2024 Ponemon Institute study found that application security teams spend over 25% of their time dealing with false positives. At scale, that’s thousands of developer hours per quarter spent triaging alerts that were never real vulnerabilities. Per SafeHill’s own Helix research, about 74% of scanner findings are false positives.
The problem compounds when scanners can’t reason about context. A SAST tool flags eval() whether the input is user-controlled or not. It flags hardcoded credentials whether they’re test fixtures or production secrets. It flags every dynamic SQL query whether it’s parameterized or not. The signal gets buried.
Why Context Matters
The difference between a finding and a risk is reachability. A SQL injection vulnerability in code that no user can reach is academic. A hardcoded API key in a test file that never ships is noise. A vulnerable dependency in a function that’s never called from user input is, at most, a low-priority cleanup task. Without context, every finding looks the same, and developers stop reading the report.
Why Exploitability Matters More Than Raw Findings
The shift in modern AppSec is from coverage to validation. Instead of producing a list of 500 possible vulnerabilities, the goal is producing a short list of confirmed exploitable paths. This is what changes the conversation with developers. “We found 500 issues, please prioritize” is a backlog. “We found three confirmed exploits, here’s the proof” is a roadmap.
This is where SafeHill’s Helix differs from traditional SAST. Helix layers AI-driven static analysis with autonomous exposure validation through Sentinel, which takes static findings and tests them against running applications using targeted attack playbooks. The output isn’t “here’s everything that looks risky.” It’s “here’s what’s actually exploitable, and here’s the proof.” For engineering teams swimming in scanner noise, that distinction is the difference between security review being useful and security review being ignored.
Do Startups Really Need Secure Code Review?
TL;DR: Yes. Attackers actively target early-stage startups precisely because they assume they aren’t being targeted, and the cost of fixing security issues climbs steeply once the product has customers, integrations, and compliance obligations attached.
There’s a predictable set of objections we hear from founders and early engineering leaders, and they’re worth addressing directly.
“We’re too early-stage.” The size of your company doesn’t matter to an automated bot scanning your subdomain for exposed admin endpoints. Bots don’t read TechCrunch, they scan everything. Early-stage startups are particularly attractive targets because they have valid AWS accounts, valid domain reputations, and valid user data, but rarely have a security team. A breach at this stage doesn’t just leak data, it ends fundraising conversations.
“We move too fast for security.” This is the objection that misunderstands how modern security review works. Slowing down for a multi-week formal audit isn’t realistic for a fast-moving team, and nobody serious is asking you to. Modern review fits inside the pull request workflow. Helix, for example, is embedded as a GitHub PR reviewer that scans new code and posts findings as inline comments on the exact lines that matter, before the merge. There’s no separate ticket queue, no PDF report, no context switch.
“We don’t have a big security team.” Most of our customers don’t either. That’s the point of platforms designed for this exact moment; automation provides the coverage, expert validation provides the judgment, and neither requires a full-time CISO on staff. This is precisely the gap continuous threat exposure management (CTEM) is designed to fill.
“The app is internal only.” Internal applications get breached through employee credential compromise, supply chain attacks, and lateral movement from external entry points. “Internal only” describes the intended audience. It doesn’t describe the actual attack surface, which includes anyone with VPN access, anyone whose credentials get phished, and any third-party tool integrated into the environment.
Why Attackers Increasingly Target Startups
The economics have shifted. Modern attackers run automated reconnaissance against the entire IPv4 space looking for low-hanging fruit. A small startup with a publicly accessible Supabase instance and weak access controls is the same target profile as a Fortune 500 with the same misconfiguration, except the startup is far less likely to have detection in place. Increasingly, the goal isn’t to compromise the startup itself. It’s to compromise the startup as a foothold to its enterprise customers, its cloud provider, or its supply chain.
Why Fixing Issues Early Is Cheaper
Beyond the bug-fix cost multipliers we covered earlier, there’s a compounding effect. A security decision baked into your architecture costs days. The same decision retrofitted after launch costs weeks. The same decision retrofitted after you have enterprise customers with SOC 2 questionnaires costs months. As SafeHill’s Secure by Design analysis notes, the cost of security decisions scales with how late they are made. What takes a day at the design stage takes a week at launch and a month after you have customers.
Best Practices for Securing AI-Generated Applications
TL;DR: The fundamentals haven’t changed: review before you ship, manage secrets properly, enforce least privilege, validate inputs, watch your dependencies, and add human or AI-augmented review before production. What’s changed is the velocity at which all of those things have to happen.
Here’s the actionable layer. None of these are revolutionary – the discipline is in actually doing them at AI-development speed.
- Never deploy AI-generated code without review. This is the single most important rule. Whether the review comes from a human, a security-focused AI agent like Helix, or both in combination, code that no one has examined for security issues should not reach production.
- Use least-privilege access controls by default. If your IAM role has admin permissions because “we’ll lock it down later,” you will not lock it down later. Set the floor and grant up from there.
- Store secrets in a secrets manager, not in code. Use Vault, AWS Secrets Manager, Doppler, or any of the alternatives. Add a pre-commit hook (Gitleaks, TruffleHog, ggshield) so secrets never reach version control. Then turn on GitHub Push Protection and run org-wide secret scanning periodically.
- Validate authentication and authorization flows server-side. Every endpoint. Every time. Client-side checks are UX, not security. AI-generated code regularly conflates the two, so this is one of the highest-yield review areas.
- Monitor your third-party dependencies continuously. Software Composition Analysis should run on every commit. Outdated dependencies don’t fix themselves, and 90% of audited codebases contain components more than four years out of date.
- Implement comprehensive logging and monitoring. You can’t respond to what you can’t see. At minimum: authentication events, authorization failures, admin actions, and data access patterns.
- Add human or AI-augmented review before production releases. Not every PR needs a security person reading it. Every PR should pass through a security review layer that flags findings inline and validates exploitability where it matters.
- Continuously test applications after deployment. Pentest cadence matters. Continuous validation through tools like SafeHill’s Sentinel matters more, because the threat landscape changes daily and your production environment changes with it.
The Future of Secure Code Review in the Age of AI Coding
TL;DR: AI-generated code is the new default, not a passing trend. Secure code review will increasingly mean AI scanning AI, with human expertise reserved for the architectural and strategic judgment that machines still can’t replicate. Speed and security stop being a tradeoff when the review layer is fast enough.
The trajectory is clear, even if the timing is uncertain. AI-generated applications will become the dominant form of new software development within the next several years. Gartner forecasts that 90% of enterprise software engineers will use AI coding assistants by 2028, up from less than 14% in early 2024. The question isn’t whether AI will write most of the code. It’s whether the security review process will keep pace.
A few things become true once that shift completes.
1. Security workflows have to evolve with AI-assisted development.
Manual review at human speed cannot scan AI output at AI speed. The review layer has to be as fast as the generation layer, which means more automation, more agentic AI security tooling, and a fundamentally different relationship between developers and security teams.
2. Startups need faster, continuous validation.
Annual or even quarterly pentests are becoming insufficient. The validation cadence has to match the development cadence, which is why the industry is moving toward continuous threat exposure management (CTEM) and autonomous exposure validation as standard practice rather than an enterprise luxury.
3. Human expertise still matters, but where it’s applied has to change.
Reviewing every line of AI-generated code by hand isn’t realistic. Designing the architecture, reasoning about adversarial scenarios, validating the toughest edge cases, that’s where human judgment still beats any automated tool. The shift is from line-level review to system-level review, with AI handling the volume.
4. Security should accelerate innovation, not slow it down.
The old model of security as a gate, where development pauses while security reviews run, is incompatible with AI-velocity development. The new model treats security as a continuous service that sits inside the developer workflow. Done right, the developer doesn’t slow down. They just stop shipping the bugs.
Secure Your AI-Generated Applications With Helix
The shift to vibe-coded software isn’t slowing down, and the gap between how fast code ships and how fast security reviews can run is widening every quarter. Most teams either ship without review and accept the risk, or they slow down and accept the velocity hit. Neither is sustainable.
There’s a third option.
SafeHill’s Helix is built specifically for the way modern software gets made. It scans AI- and human-generated code before it goes live, validates which findings are actually exploitable through Sentinel’s runtime testing, and posts inline comments on the exact lines of code that matter. No separate ticket queue. No PDF report nobody reads. No alert backlog full of false positives.
What that gives engineering teams:
- Stronger signal from static analysis. Helix catches harder-to-find issues that traditional tools miss, including logic flaws, multi-step injection paths, and auth bypasses.
- Runtime-confirmed exploitability. Sentinel tests static findings against the running application, so the issues that reach your developers are the ones worth fixing.
- Better security for fast-moving software. Helix is designed for AI-assisted, vibe-coded, and rapidly iterated codebases, aka the modern reality. Not the 2015 enterprise stack.
Move faster without guessing. Book a demo or see pricing to learn how Helix can fit into your team’s workflows.