Introducing FaultLine: Open-Source AI Benchmark Tool for SAST & DAST

/

Sixty deliberately-vulnerable Flask apps. A per-app severity-weighted scorer. Three verdicts, dual evidence, no # VULN comments to parse. FaultLine is the public AI benchmark we wanted for security scanners, and now it exists.

If you’ve spent any time evaluating security scanners, you’ve probably hit the same wall we did. The public AI benchmarks are too narrow, too easy, or too lenient. DVWA and Juice Shop are training environments, not evaluation harnesses. The OWASP Benchmark Project covers Java servlets in isolation. NIST SARD is largely synthetic test cases. None of them are designed to produce comparable per-tool numbers across modern application code. And the boolean “did you find it” scoring that most of them use rewards the tools that flag everything in sight, because high recall pads the score and precision goes uncounted.

Quick context on what we mean by “security scanner,” because the term gets used loosely. The category we care about is the family of tools that find vulnerabilities in code or running applications. Two main flavors. SAST (static application security testing) reads source code and points at risky patterns: tainted input flowing into a SQL query, a missing authorization check, a deserialization sink that accepts user data. DAST (dynamic application security testing) runs against a live app and tries to actually exploit things: send a payload, watch the response, confirm the bug is real. A third, related category is SCA (software composition analysis), which inspects third-party and open-source dependencies rather than first-party code, so an SCA benchmark answers a different question than the one FaultLine targets. The good ones do both, or they hand off to a partner that handles the other half.

Specifically, we needed this AI benchmark tool for Helix, our own SAST scanner with AI inference. Helix reads source code the way an experienced security engineer would. It follows tainted input from where it enters the application all the way to where it lands, reasons about the path, and emits findings with severity, a CWE, and a recommended fix. It runs on our own GPUs, on a model we fine-tune on our own security corpus, integrated directly into customer CI as a GitHub Action that comments inline on pull requests.

The important word is production. Helix is not a research project. Customers run it on their real codebases, where false positives waste real engineering time and missed vulnerabilities ship to real attackers. That changes what we need from a benchmark. An academic benchmark that confirms an idea works in principle is not the same thing as an honest benchmark for a tool that runs against real PRs every day. We needed one that was comprehensive (broad enough that the score actually predicts real-world performance), real-world adjacent (target apps that look like production code, not lab specimens), and harsh against false positives (because in production, every false positive costs a security engineer’s afternoon).

So we went looking. Nothing public hit all three. We couldn’t find it. So we built it.

Today we’re open-sourcing FaultLine v2.0 under Apache 2.0. It’s sixty separate Flask applications, each containing exactly one focused vulnerability of a specific class, packaged with a per-app severity-weighted scoring harness. The source code looks like ordinary application code. Scoring rewards both detection and confirmation. Reference adapters are placeholders in the harness directory, awaiting community contributions. The repo is live now.

Clone it, boot it, score your scanner.

Apache 2.0. Public ground truth. Sixty apps boot in one Docker Compose command. The scorer is a Python CLI. No vendor portal, no signup, no sales call.
OPEN SOURCE

Why the field needed a new AI benchmark tool

TL;DR: Existing public benchmarks are too narrow, too easy, or too lenient, and most score detection as a yes/no, which rewards noisy tools. FaultLine is built to be broad, hard to game, and to score both detection and confirmation.

The honest answer: because we kept hitting the limits of the existing ones, and after the third or fourth time we wished a benchmark existed that didn’t have those limits, we sat down and made one.

The public benchmarks are too narrow, too easy, or too lenient. DVWA presents canonical vulnerability patterns at four difficulty levels in a small PHP application. Useful for training. Not designed to produce comparable per-tool numbers across modern application code. OWASP Juice Shop is challenge-based with a scoreboard, built for hacker training and CTF prep, with mitigation hints attached to solved challenges. Great for skill-building. Same problem when used as a benchmark. The OWASP Benchmark Project gets closer in shape, around 2,740 single Java servlet test cases each mapped to a specific CWE, but every test case is one weakness in isolation. A scanner that scores well on synthetic Java servlets is not the same thing as a scanner that scores well on application code in production. NIST SARD is the largest of the public datasets, hundreds of thousands of test cases across multiple languages, predominantly synthetic. Useful as a tool-development corpus. Not what we needed.

The third trap is harder to talk about because it shows up in scanners’ favorite benchmarks. Annotated targets where every vulnerable function carries a multi-line comment block above it tagging the class, OWASP category, CWE, and a sample payload. Useful for the original authors. Catastrophic as a benchmark. Any modern AI-driven scanner that consumes source code reads the answer key. We had this exact problem with our own internal predecessor to FaultLine. Every vulnerable function had an annotation block. Useful for us. Worthless as a measure of whether a scanner could find anything.

And the boolean problem hangs over all of them. “Did the tool flag this vuln, yes or no.” Rewards the tools that flag everything in sight. High recall pads the score, precision goes uncounted. The benchmark optimizes for the wrong thing.

We wanted a benchmark that was broad, deliberately uncomfortable to game, and scored both detection and confirmation. The same standard a real security team uses when triaging tool output.

What FaultLine is: an open-source SAST and DAST benchmark

TL;DR: FaultLine v2.0 is sixty deliberately-vulnerable Flask apps, one focused vulnerability each, with no answer-key annotations in the code. To earn credit, a scanner has to do real data-flow analysis, not pattern-match on comments.

FaultLine v2.0 is sixty deliberately-vulnerable Flask applications, distributed via GitHub under Apache 2.0. Each app contains a single focused vulnerability of a specific class. Each app is self-contained, boots on a known port via Docker Compose, exposes a /health endpoint, and ships an auth.json file declaring its login flow.

The source code looks like ordinary application code. There are no # VULN annotations, no comments tagging the vulnerability class, no payload hints in the test fixtures. Before anyone asks: yes, we scrubbed the apps for accidental prompt-injection patterns in comments, strings, and template content. AI-driven scanners that consume source code can be manipulated by crafted instructions in the input. An accidental “ignore previous instructions” fragment in a test fixture would be a confounding variable. The corpus is clean. The vulnerable function looks like the rest of the file. A scanner that wants credit has to actually do data-flow analysis, not regex on TODO comments. Real-world vulnerabilities rarely come labeled, and we want the benchmark to behave like the real world.

One known limitation in v2.0: app directory names are descriptive (e.g. 001-flask-sqli-login-bypass), which gives a source-code-aware scanner a class hint before analysis. We chose descriptive names for v2.0 to keep maintainer ergonomics simple. v3.0 will use opaque names so the class is not handed to the scanner via the file path. Tools that game the descriptive names by mapping path patterns to CWE classes are visible to a reviewer reading the adapter source.

APPS
0
VULN CLASSES
0
DIFFICULTY TIERS
0
VERDICTS
0

Class distribution: 11 SQL injection variants (including NoSQL), 9 XSS variants, 7 SSRF, 6 insecure deserialization, 5 command injection, 4 SSTI, 4 path traversal, 3 XXE, 3 code injection, 3 open redirect, 3 IDOR, plus mass assignment and broken function-level authorization. Per-app catalog in ground_truth/v2.0.json.

The sixty apps split into three difficulty tiers. The first forty are pre-auth SAST classics: function-scoped, classic OWASP, the kind of vulnerability a static analyzer should catch by following tainted input from request to sink. A scanner that flags the function gets credit for detection; one that exercises the route and triggers exploitation gets the dual-evidence credit.

Apps 041 through 055 are the SAST + DAST chain tier. The static signature alone is necessary but not sufficient: a SAST tool can flag the code, but proving exploitation requires a DAST tool to actually trigger the vuln through a request. This is the tier where the dual-evidence rule produces the most informative signal, separating tools that confirm exploitation from tools that stop at the static finding.

The last five (056-060) are auth-required IDOR and broken authorization. A scanner needs to log in, navigate as user A, identify a resource ID, then attempt access as user B. Static analysis alone rarely catches IDOR because the vulnerability is in the absence of an authorization check, not the presence of bad code. This is the tier where AI-driven scanners with planning capability separate from pure pattern-matchers.

Tier 3 is intentionally small in v2.0. Encoding IDOR and broken authorization at unit granularity is harder than encoding SQLi or XSS, because the vulnerability lives in the relationship between two requests rather than in a syntactic pattern. Five apps was the right number for v2.0. With only five, a per-tier percentage is too noisy to compare across scanners. The headline rolls up the verdicts. Per-tier breakouts stay in the verbose output as directional signals, not comparison metrics. v3.0 expands the tier as the catalog grows.

The dual-evidence rule: scoring SAST and DAST benchmarks fairly

TL;DR: Every app gets one of three verdicts. Confirmed (both static detection and dynamic exploitation) earns full credit, detected-only earns half, missed earns zero. A SAST-only or DAST-only tool caps at 50%, which keeps the SAST benchmark and DAST benchmark halves honest.

Every app gets one of three verdicts:

VerdictMultiplierWhen it is awarded
Confirmed1.0×The tool detected the vulnerability statically and confirmed exploitation dynamically.
Detected_only0.5×Half-credit when the tool managed one but not both.
Missed0×The obvious case: neither half.

This split mirrors how real security teams handle findings. A SAST finding is a candidate. A DAST-confirmed finding is an actionable issue. Triage queues, severity ratings, and remediation timelines all reflect this difference. A finding labeled “static analysis flagged this” sits in a backlog with thousands of others. A finding labeled “we exploited this on the live system” gets scheduled into the next sprint.

A SAST-only tool caps at 50% on every app in v2.0, not just Tier 2. This is the dual-evidence rule applied uniformly: confirmed (1.0×) requires both static detection and dynamic confirmation; everything else is detected_only (0.5×) at best. Pure SAST tools like Bandit or Semgrep get full credit on detected_only and zero additional credit for not having a DAST half. Pure DAST tools (ZAP, Nuclei) hit the same ceiling for the inverse reason. Tools that combine both halves, or that pipe SAST findings through a separate DAST validator, can break the ceiling. The benchmark does not punish a tool for being honest about what it can and cannot prove. It just refuses to award unverified findings as if they were proven exploits.

Per-app scores are weighted by severity:

SeverityWeight
Critical5
High3
Medium2
Low1

The headline score is 100 × actual_points / max_points. Precision, recall, and F1 are reported separately, which we will get to in a moment.

The decoy false-positive penalty

TL;DR: A separate decoy set of safe-but-suspicious code patterns catches scanners that pattern-match on shape. Each decoy hit subtracts from the score and is reported separately, so precision is visible independent of recall.

The harness ships with a separate v2.0_decoys.json file containing known-safe code patterns. They look like vulnerable code at a glance but are not actually exploitable. A naive scanner that pattern-matches on syntactic shape hits the decoys. Each decoy hit subtracts severity_weight × 0.2 from the score, and decoy hits are reported separately so reviewers can see the precision picture independent of the headline. 

The decoy set is small in v2.0 and grows in v2.x as we collect false-positive patterns from real scanner runs against the corpus. The intent is to surface aggressive labeling in the data, so a tool with high recall and low precision is visible distinct from one with high recall and high precision.

How to run the FaultLine AI benchmark tool

TL;DR: Three commands clone the repo, boot all sixty apps on localhost via Docker Compose, and score your scanner’s claims with the Python CLI. The claim format is JSON Schema-validated and vendor-neutral.

If you have made it this far, you probably want to know what running this against your own scanner looks like. Three commands:

				
					# clone
git clone https://github.com/SafeHill-Research/Faultline.git
cd SafeHill-Faultline
 
# boot all 60 apps on ports 9001-9060 (bound to 127.0.0.1)
docker compose -f orchestrator/compose.global.yml up -d --build
 
# run your scanner against the apps, then score
python harness/scorer.py \
  --ground-truth ground_truth/v2.0.json \
  --decoys ground_truth/v2.0_decoys.json \
  --claims your-tool-claims.json -v

				
			

Your scanner produces a claims.json file describing what it found. The format is JSON Schema-validated. Each claim has an app_id, an optional file:line or endpoint:method, a CWE, a severity, and an evidence object that controls whether the claim counts as DAST-confirmed. Here is a minimal example:

				
					{
  "tool": "your-scanner",
  "version": "1.2.3",
  "ground_truth_version": "2.0",
  "claims": [
	{
  	"app_id": "001",
  	"cwe_id": 89,
  	"file": "app.py",
  	"line": 32,
  	"endpoint": "/login",
  	"method": "POST",
  	"severity": "critical",
  	"evidence": {
    	"dast_verdict": "confirmed",
    	"payload_response": "..."
  	}
	}
  ]
}

				
			

A claim matches an app when the app_id matches and either the claim’s endpoint:method matches a known endpoint, or the claim’s file matches a known vulnerable file. When ground truth includes line numbers (planned for v2.x for selected apps; not present uniformly in v2.0) the harness adds a ±15-line tolerance for file:line matches, function-level granularity. v2.0’s function-scoped apps are small enough that file matching alone is precise; the line tolerance becomes meaningful when v3.0 adds multi-vuln-per-file targets.

The dast_verdict field is the load-bearing wall of the dual-evidence rule. It is intentionally vendor-neutral: the field name reflects whether any dynamic analysis confirmed the vuln, regardless of which scanner produced the verdict. A SAST tool that pipes findings through a separate DAST validator can populate dast_verdict from that validator’s output. A standalone DAST tool can populate it from its own runtime probing. Tools that combine both architectures can populate it however they want. The schema does not bias toward a specific tool design.

One question worth answering directly. The dast_verdict field is reported by the tool, not replayed by the harness. We do not run every payload_response through the live container to verify exploitation. We do, however, check the response. v2.0 ground truth carries match_pattern strings for 59 of 60 apps: the literal substring that should appear in the response when the vuln is actually exploited. When a tool supplies a payload_response, the scorer confirms the pattern is in the response before granting confirmed credit. A tool that reports dast_verdict as not_confirmed or inconclusive is taken at its word and given detected_only credit at most. This is partial verification, not full replay. Replay produces false negatives from environmental drift. Pattern verification produces a real signal without that fragility, and it makes fabricating confirmed verdicts visibly harder in the per-claim output.

When you run the scorer with -v, you get a per-app verdict table plus the headline:

				
					Tool: example-scanner v1.0 | Ground truth: v2.0
Note: illustrative output, shaped to demonstrate the report format.
Not a real scanner result.
Apps confirmed: 34 / 60
Detected-only:  16
Missed:     	10
Decoy hits:  	0
Other FPs:   	0
Score: 73.31% (173.00 / 236.00 points)
Precision: 1.000
Recall:	0.833
F1:    	0.909
Per-app verdicts:
  001 CONFIRMED  	(5.00) CWE-89, /login:POST
  002 CONFIRMED  	(3.00) CWE-89, /search:GET
  003 DETECTED_ONLY  (1.50) CWE-79, /comment:POST [no DAST]
  004 MISSED     	(0.00) CWE-79
  ...

				
			

For CI integration, –output result.json emits the same data structurally so dashboards and pipelines can consume it directly.

Adapters: making your scanner submit-ready

TL;DR: Most scanners emit SARIF or proprietary output, so a small Python adapter translates findings into FaultLine claims. v2.0 ships no reference adapters; they land as community contributions, and SafeHill publishes both minimal and vendor-tuned scores when both exist.

Most existing scanners do not produce FaultLine claims directly. They emit SARIF, or proprietary JSON, or XML, or CSV. The bridge is an adapter: a small Python script (typically 50-200 lines) that runs the underlying scanner against the FaultLine apps, parses its native output, and translates each finding into a valid claim.

Adapters live under harness/adapters/<tool-name>/ with a stable convention: an adapter.py exposing a run() entry point, a pinned requirements.txt, and a test_adapter.py that takes synthetic findings and verifies they translate to expected claims. We are keeping the adapter authoring guide in docs/ADAPTER_GUIDE.md stable so the bar to contribute stays low.

A note on how reference adapters relate to vendor adapters: v2.0 ships with no reference adapters in the box. The harness/adapters/ directory is a placeholder for community contributions. When a vendor or independent contributor submits an adapter via PR, we will publish the resulting score in benchmarks/v2.0/<tool>.json. If two adapters exist for the same tool (a minimal-viable reference and a vendor-tuned version), we publish both numbers. A scanner scoring lower with a minimal adapter than with a tuned one is honest information about how much adapter engineering improved the result. Collapsing the two into one number is not. This avoids the dynamic where a vendor accuses us of writing a weak adapter for them. It also creates an actual incentive to invest in adapter engineering.

A note on writing adapters

Three pitfalls we have seen in early drafts:

  • Inflated severity: if your scanner reports everything as “high,” do not pass that through. Use the catalogued severity in the ground truth as the source of truth and let the scorer apply the weight.
  • Duplicate claims: some scanners fire 5+ times on the same vuln in the same function. The scorer is forgiving on the score itself: matching duplicates collapse into one verdict per app, no double-counting of points. Where dedup matters is the false-positive count: any duplicate that falls outside the matching window (different endpoint, or beyond the line tolerance when ground truth supplies line numbers) is counted as a separate false positive. Adapters that emit one claim per finding (rather than one per detection rule) produce cleaner scoring and more useful precision numbers.
  • False dast_verdict: confirmed: only set this if your scanner is actually exploited. Lying about evidence inflates the score and tanks the precision metric, which the scorer surfaces immediately. The benchmark catches this faster than you would expect.

The FaultLine AI benchmark roadmap

TL;DR: v2.0 is a foundation. Ground truth is frozen within a version line so scores stay comparable, and later versions raise the bar: v3.0 modernizes the target stack, v4.0 adds multi-step exploit chains and business-logic bugs that punish pattern-matchers.

v2.0 is a foundation, not the final state. The benchmark must stay hard as the field improves, which is why versioning is part of the discipline. Within a version line, ground truth is frozen so historical scores stay comparable. Across version lines, scores are not directly comparable: a 90% on v3.0 is not the same accomplishment as a 90% on v2.0, because v3.0 will measure harder targets.

One thing the version-line discipline addresses is contamination. v2.0 is now public, which means it can be (and likely will be) incorporated into training corpora for AI-driven scanners over time. AI-tool scores on v2.0 a year from now will be measuring memorization rather than analysis. We can’t prevent this. We also don’t think trying would help. The mitigation is the cadence. v3.0 and v4.0 introduce ground truth that hasn’t been public. Anyone reading FaultLine v2.0 scores from twelve months out should weigh them accordingly, and apples-to-apples comparison is across tools running against the same version line, not across them.

The v3.0 line, scoped at six-to-nine months out, modernizes the target stack. SPA frontends with JSON APIs. GraphQL apps with broken per-resolver authorization, introspection enabled, batched-query DoS surface. OAuth and SAML auth flows with broken state validation and IdP-confusion patterns. JWT apps with kid injection and algorithm confusion. Server-side prototype pollution. HTTP/2 request smuggling. SSRF behind URL parsers. Cache poisoning via header tricks. The v2.0 stack is server-rendered Flask, which does not exercise scanners against modern web apps; v3.0 will.

The v4.0 line, twelve-plus months out, is where we think the benchmark gets really uncomfortable for naive tools. Multi-step exploit chains: a CSRF leading to a second-order injection leading to admin privilege escalation, where you only score the chain ID if you find all three in sequence. Business-logic vulns: discount-code stacking, refund double-claim, signup-flow race conditions, MFA-bypass via password-reset flow. Time-bound vulnerabilities where order-of-operations is the bug. Conditional vulnerabilities only triggered under specific feature flags or business hours. This is the tier where pattern-matching tools collapse and planning-capable tools earn their keep.

We are also planning capability metrics that layer onto every version line: time-to-first-finding per vuln class, request budget per finding, authentication handling depth, crawl depth, precision under load. Far harder to game than a binary verdict.

About our own scores

TL;DR: SafeHill isn’t publishing Helix or Sentinel scores in the public repo. The public artifact is the methodology, the apps, and the harness, so any score that does get published is reproducible from a clean clone.

You might be wondering: where does Helix score on this? Or Sentinel? We are not publishing those numbers in this repository. Internal benchmarks of our own products belong in our pitch decks and our customer-facing reports, not in the public repo. Publishing self-reported numbers in the same repo as the methodology has been tried by other vendors and is uniformly read as marketing.

The public artifact is the methodology, the apps, and the harness. Any score that lands in benchmarks/v2.0/<tool>.json should be reproducible from a clean clone with the same harness against the same ground truth. If we publish a score for our own tools there, the methodology demands that anyone with a clean clone can verify it.

Internally, FaultLine gates every Helix model promotion in our quarterly LoRA pipeline. That’s why we built it. The reason we open-sourced it is that the rest of the field deserves the same.

ONE SAFETY NOTE BEFORE YOU BOOT IT

Read SECURITY.md before you deploy. Every app in the repo is intentionally vulnerable. The Compose file binds to 127.0.0.1 by default. If you set BIND_HOST to anything else, the orchestrator refuses to start unless I_UNDERSTAND_THIS_IS_VULNERABLE=1 is also set. Belt-and-suspenders for the predictable case where someone copy-pastes a network-debug fix and accidentally exposes sixty vulnerable Flask apps to their host network. Run on an isolated network only.

Try the FaultLine AI benchmark tool

FaultLine v2.0 is on GitHub under Apache 2.0. Reference adapters are landing as community contributions merge. The methodology, harness, and rationale docs are all in the repo. We would love to see your scanner’s score, especially if it is better than ours.

If you build an adapter, open a PR. If you find a bug in the scorer, file an issue. If you want to argue with one of our design choices, the RATIONALE.md document records why we made each one, and we are genuinely open to changing our minds when the argument is good. The benchmark works better with more eyes on it, which is half the reason we open-sourced it.

Picture of About the Author

About the Author

Trevor Baines is Director of AI Engineering at SafeHill, a cybersecurity AI company building local-first SAST, DAST, and CTEM-driven products powered by owned GPU infrastructure. With a background in offensive security and application penetration testing, he focuses on the intersection of cybersecurity, AI, and operations, and has led the company's AI R&D efforts including its Helix engine and NVIDIA Inception partnership.