Audn vs Codex vs Aikido vs Mythos on the Same Repo — 235 Findings, and All Four Agreed on Exactly 4
We added a fourth scanner. Claude Mythos ran a whole-repo deep-static review of the same OWASP Juice Shop source that Audn WhiteBox, OpenAI Codex, and Aikido already scanned. 235 findings between the four, and the four-way intersection is four issues. Mythos found an entire availability class — 10 DoS bugs and 4 process-crashers — that none of the other three worked. Here's the full A/B/C/D, our own gaps included.
Audn vs Codex vs Aikido vs Mythos on the Same Repo — 235 Findings, and All Four Agreed on Exactly 4
Updated September 18, 2026 — this is the current version of our Juice Shop scanner comparison. It adds a fourth tool, Claude Mythos, to the earlier three-way Audn-vs-Codex-vs-Aikido writeup. Same rule as before: we publish the parts that don't flatter us.
Every security vendor publishes a benchmark where they win.
We ran a fourth scanner over the same punching bag anyway — and it found a whole class of bugs that we missed. That's the part nobody publishes, so it's the part we lead with.
Claude Mythos — Anthropic's frontier model doing a whole-repo deep-static security review — landed 76 findings on the same OWASP Juice Shop source that Audn WhiteBox, the official OpenAI Codex security review, and Aikido had already been through. And it owns a lane none of the other three worked: availability. Ten denial-of-service bugs, four uncaught exceptions that crash the process, and a couple of quadratic-regex ReDoS stalls. Audn is the one that runs live and proves exploits — and Audn barely touched this class, because our lens is tuned for "what can an attacker steal," not "what makes the server fall over."
The full interactive report is live, unedited, with every row from all four sides:
→ blog.audn.ai/audn-vs-codex-vs-aikido-vs-mythos
(The three-way A/B/C is still at blog.audn.ai/audn-vs-codex-vs-aikido; the original two-way at blog.audn.ai/audn-vs-codex.)
The setup is unchanged. Same deliberately-vulnerable OWASP Juice Shop source, four very different machines looking at it: Audn's live red-team, Codex's commit-diff review, Aikido's SAST + SCA + secrets platform, and now Mythos's Claude-native whole-repo static reasoner.
The headline numbers are 103, 20, 36, and 76. Those are still the least interesting things in the report.
The Result That Should Worry You
Between the four tools there were 235 findings.
The four-way intersection — issues that all four independently flagged — is four issue areas:
- JWT signature is not verified / forgeable — the hard-coded key and algorithm confusion in
lib/insecurity.ts. Every tool reaches it. Mythos went further and named the exact trapdoor: thedenyAll()guard is itself a JWT verifier that acceptsalg=none, opening every forbidden finale verb. - Unpinned third-party GitHub Actions holding the repo token —
image_actions.yml/ci.yml, all four, from four different angles. - Code injection / RCE — server-side
eval/template injection and MarsDB$wherein-process execution. Mythos went deepest here (5 findings), reaching$whereinjection intrackOrder.tsandshowProductReviews.ts, plus RCE viajs-yaml!!js/functionin the complaint YAML upload. - SSRF — server-side request forgery via profile-image URL / chat parts, reached by all four.
Four tools. 235 findings. Four unanimous. When a live red-team, a diff reviewer, a commercial SAST/SCA platform, and a Claude-native deep-static scanner all land on the same class independently, treat it as certainly real — and treat the other 231 as proof that these tools are not substitutes for each other.
| Audn WhiteBox | Codex | Aikido | Mythos | |
|---|---|---|---|---|
| What it is | SAST + live red-team | Commit-diff review | SAST + SCA + secrets + CI | Claude-native deep static |
| Findings | 103 | 20 | 36 | 76 |
| Severity mix | 18C / 24H / 52M / 5L / 4I | 8H / 7M / 2L / 3I | 3C / 13H / 8M / 12L | 7C / 45H / 21M / 3L |
| Reproduced live | 18 | 0 | 0 | 0 |
| Dependency CVEs (SCA) | No | No | Yes — 3 | No |
| DoS / robustness | Light (3 + 1) | 1 | None | Heavy (10 + 4) |
| Diff-aware | No | Yes | No | No |
| Intentional vs unintended | No | Yes | No | Yes — per challenge |
| Remediation | Prose + CVSS + paths | Generated git diff | Fix-time estimate | Precise fix criteria |
| Runtime | 2h 28m + live target | Light, diff-scoped | ~52 seconds | Medium, whole-repo |
| Sweet spot | Pentest / exploit proof | PR gate / regressions | Continuous baseline | Deep pre-merge review |
Where Each One Actually Shined
Audn shined on: proving things
103 findings with CWE, file:line, CVSS vectors, and 32 derived attack paths. But the number that matters isn't 103 — it's 18, the findings the agent didn't just report but fired at the running app and reproduced:
server.ts:280— unauthenticated public access logs containing change-password URLs with cleartext current and new passwords. Not a code smell. Pulled off the live box.- A spoofable
X-Forwarded-Forheader bypassing the password-reset rate limit — an account-takeover enabler. Audn found it by getting rate-limited and then working out how to not be. (Mythos reached the same critical by pure static reachability reasoning, fromserver.ts:340. Two tools, opposite directions, same bug — that's what "certainly real" looks like.) routes/search.ts:21— SQL injection, fired and confirmed. NoSQL injection intrackOrder.tsandshowProductReviews.ts. XSS viamodels/product.ts. Hard-coded credentials inusers.yml,7ms.yml, androutes/login.ts— four of them, all validated live.
The other Audn-exclusive lane is business logic and crypto, which even three other scanners couldn't reach: weak password recovery (seeded security answers), MD5 credential hashing, client-clock-controlled coupon validity, unbounded discounts, wallet overspend, cleartext-HTTP transport, login-IP spoofing. And the biggest slice — 16 broken-access-control/IDOR findings. Aikido found zero IDOR. No regex knows who's supposed to own a basket.
Codex shined on: what the change broke
Codex reviewed the changeset as a diff. Nearly every finding is framed as "introduced by this commit," and it still owns that lane cleanly — even against Mythos's whole-repo reasoning, because a whole-repo scanner sees a file it already knows, while only a diff reader sees the change:
- 2FA temporary JWTs accepted as bearer auth. A commit stopped inserting no-
dataJWTs intoauthenticatedUsersto fix a crash — but never started rejecting them. The code got safer-looking and less safe in the same commit. - A Junie AI-skill that
curls untrusted reference URLs and feeds the response back to the agent — SSRF into localhost and link-local ranges. Agent tooling in the developer environment; a supply-chain surface. - A new default theme hard-coding remote Google Fonts, leaking every visitor's IP, UA, and referer to a third party — a privacy regression.
- Plus a Windows-only i18n path bug and dropped test-coverage assertions.
And every Codex finding arrives with a patch you can apply. It's the only tool here that both distinguishes intentional challenge vulns from real risk and hands you the fix.
Aikido shined on: the dependency hole
This is the finding that earned Aikido its seat in the three-way version, and it still holds: Aikido is the only one of the four that runs SCA.
- 3 real dependency CVEs —
jsonwebtoken(missing input validation),express-jwt(improper authorization),sanitize-html(XSS). Nobody else ran SCA. Nobody else would have found these. - 17 secrets findings across 40+ files, including test fixtures — one file alone holding 33 exposed secrets, plus exposed JWTs scattered through the tree.
- 5 SAST/CI findings the others missed — unsafe YAML load leading to RCE,
document.write()XSS, file inclusion, andactions/checkoutpersisting Git credentials.
And it did the whole thing in ~52 seconds. That speed is the entire reason Aikido belongs on every push.
Mythos shined on: the availability lane nobody else worked
Here's the new one, and it's the reason this update exists.
Mythos read the code, the tests, and the challenge definitions — so it reasons about reachability and separates intended from unintended, which keeps its 76 findings low-noise despite the breadth (45 of them rated High). But its real signature is a class the other three barely acknowledged: can this server be knocked over?
- 14 findings are availability bugs — 10 DoS / resource-exhaustion and 4 uncaught exceptions that crash the Node process.
- A client-chosen multipart
Content-Typebecomes a permanent Prometheus label (routes/metrics.ts:73) — unbounded cardinality growth that slowly eats the metrics backend. - Quadratic-regex ReDoS on an unauthenticated socket.io
verifySvgInjectionpath and on XML-upload output — a single crafted string stalls the event loop. - A per-request full scan plus bigram similarity over every complaint (
routes/verify.ts:371) — the same complaint-DoS shape Codex spotted in the diff, but Mythos found it whole-repo and mapped the amplification. PUT /api/BasketItems/:idruns an un-awaitedquantityCheckwhose rejection is unhandled — one request takes the process down (routes/basketItems.ts:65). Same story for a failed avatar download at startup and a subtitles-as-URL promotion page.
Mythos also posted the deepest IDOR set after Audn (12 findings), each tied to the specific Juice Shop challenge it maps to, and matched Audn on several criticals — the X-Forwarded-For rate-limit bypass, the login SQL injection, the alg=none JWT bypass — reaching all of them statically. Where Audn proves, Mythos reasons; on this repo they corroborated each other on the criticals and diverged completely on availability.
Where Each One Is Weak, Stated Plainly
Audn: Most findings unconfirmed (63 static leads + 18 probed-no-verdict; only ~17% reproduced live). No patches. Slow (2h 28m) and needs a deployed target. No SCA. No diff awareness. Over-reports intentional vulns.
Codex: Narrow — 20 findings, one changeset, misses most of the app's surface by design. No live proof. No SCA, no secrets sweep. Near-duplicate findings inflate the count.
Aikido: Static only. No diff awareness — misses every one of Codex's introduced regressions. Noisy: many "secrets" are test fixtures, and counts are grouped. No business-logic or IDOR reasoning. Incomplete SCA — it flagged its own missing lockfile, so transitive known-vulnerable packages can still hide.
Mythos: No live proof — 76 findings, all static, however precise the reachability reasoning. No SCA / dependency CVEs — it would have missed Aikido's three. No secrets-in-tests sweep. Heavier than a diff scan, so it's a pre-merge or scheduled reviewer, not a per-commit one.
Audn's honesty layer
One structural difference is worth naming again, because adding Mythos sharpens it. Audn grades its own certainty and publishes the grade:
| Confidence tier | Count | Meaning |
|---|---|---|
| Confirmed | 16 | Flagged in source and reproduced live |
| Observed live | 2 | Reproduced live with no static lead behind it |
| Probed, no verdict | 18 | Attacked, reached no conclusion |
| Static lead | 63 | Flagged in source; the run never got to it |
| Likely false positive | 4 | Exercised without reproducing |
None of Codex, Aikido, or Mythos has an equivalent tier. In Audn's terms, all 132 of their findings sit at the unverified level — as do 85 of our own. Mythos reasons harder about reachability than anything else here, but reachability is still a claim about the code, not a fact about the running system. Only 18 findings across all four tools are the latter.
Recommended stack — where each scanner runs
These four aren't competitors. They're gates at different points on the path from commit to main, ordered by cost and depth: cheap, fast, frequent first — expensive and live last. Mythos slots in as the deep whole-repo review just before merge, complementing Codex's fast diff scope; Audn stays at the merge-to-main gate, exactly as we called it in the three-way version.
On every commit / push → Aikido (~52s)
SAST + SCA + secrets + CI/IaC baseline. Fail on a new dependency CVE. Fail on a committed secret. Cheap enough for every push, and the only SCA + secrets net.
When a PR is opened → Codex (light)
Diff-aware review of the changeset. Block introduced regressions, attach the patch, skip intentional vulns. Tells reviewers what this change broke versus what was already there.
When the PR is in review → Mythos (medium)
Deep whole-repo static review. Surface DoS / robustness and logic bugs the diff scope can't see — availability, IDOR, process crashes — with reachability-aware, challenge-aware precision. This is the gate that catches "the server falls over," which nothing upstream of it looks for.
When the PR is ready to merge → main → Audn (~2h 28m)
SAST + live red-team against a deployed preview of the branch. Block merge on a live-confirmed critical. The last line before code hits main — the only tool that proves real exploitability.
⚑ Prereq: Audn needs a running target. Deploy the PR to a preview/staging environment first, then point Audn at it.
On main / in production → Aikido + Audn
Aikido re-scans continuously (new CVEs land daily); Audn re-runs periodically against prod or staging.
Why this order — cheapest and most frequent first, exploit-proof last
| Gate | Tool | Runs on | Cost | What it blocks on |
|---|---|---|---|---|
| Baseline | Aikido | every commit / push | ~52s | a new dependency CVE or a committed secret |
| PR gate | Codex | PR opened | light | a regression or logic bug the diff introduced |
| Pre-merge review | Mythos | PR in review | medium | a DoS / robustness / logic bug anywhere in the repo |
| Merge-to-main gate | Audn | PR ready to merge | ~2h 28m | a vuln live-confirmed exploitable on a preview deploy |
| Post-merge | Aikido + Audn | continuous / periodic | mixed | — monitoring for new CVEs & drift, not a gate |
Codex and Mythos overlap — both are AI code reviewers — but they differ in scope: Codex is diff-scoped and fast (every PR), Mythos is whole-repo and deep (pre-merge or scheduled). If you run only one AI reviewer, Mythos gives you breadth; Codex gives you the "what this PR introduced" delta and a patch. Most teams that can afford it run both, because the diff view and the whole-repo view catch genuinely different bugs.
The one caveat, unchanged: Audn's ~2h 28m run means the merge gate is not instant. Budget for it, or fire Audn at the preview as soon as the PR is approved rather than at the moment of merge.
Why a company needs all four
Each one is blind in a way the other three are not, and the blindness is structural — not a roadmap gap you can wait out.
- Drop Aikido and you ship known-vulnerable
jsonwebtokenandexpress-jwt, plus 17 secrets in your tree. It's the only one running SCA. - Drop Codex and every subtle regression your team introduces ships. Only a diff reader sees the change inside a file the other scanners think they already understand.
- Drop Mythos and the availability class goes dark. Ten DoS bugs and four process-crashers — the ReDoS on an unauthenticated socket path, the unbounded Prometheus label, the un-awaited rejection that kills the process — were surfaced by exactly one of the four tools. Your uptime depends on the one nobody else replaced.
- Drop Audn and you have 132 static findings across three tools and zero proof that any of them fire in your deployment — plus no coverage of the live-confirmed criticals, the business-logic layer, or the crypto weaknesses. You'd be triaging a pile of maybes and calling it a security posture.
They're layers, not substitutes. Fast static baseline, PR gate, deep pre-merge review, live merge-gate pentest.
What still isn't covered — by any of the four
Adding Mythos closed a blind spot that survived the three-way version: DoS / availability and robustness is now covered. Aikido had already closed the dependency-CVE gap. What survives even with all four running:
- Live infrastructure / edge posture. Three of the four are static; Audn ran live but left its transport/header/container items unreached or at no-verdict. The real TLS config, the actual security headers the edge returns, and container hardening remain unmeasured — one
curlagainst the real deployment would tell you more than any of these reports. - Live exploit confirmation of the static majority. Audn confirmed 18 live. The other ~215 findings — Codex's 20, Aikido's 36, Mythos's 76, and Audn's own 85 unconfirmed — are all static. Mythos reasons hard about reachability, but nothing except Audn proved anything against a running instance.
- Complete SCA / SBOM. Only Aikido does SCA, and it flagged a missing lockfile — so there's no full transitive dependency graph, license posture, or signed SBOM across the four.
That second item is the honest summary of the whole exercise: of 235 findings, 18 are demonstrated facts. Everything else is a hypothesis of varying quality — some of it, especially Mythos's reachability reasoning, very high quality. But anyone selling you a single tool as complete coverage is selling you the gap.
What's Under the Hood
WhiteBox runs on Necromicon — our frontier cyber model built on Kimi K3, abliterated so it doesn't flinch mid-engagement and fine-tuned on real pentesting sessions rather than CTF puzzles. The live-reproduction step is the whole reason it exists: a guardrailed model refuses to fire the payload, and a finding you never fired is a finding you can't grade. This comparison is the cleanest illustration we've published — Mythos, a genuinely excellent static reasoner, matched Audn on several criticals by reasoning and beat us outright on availability, and still couldn't tell you which of its 76 findings actually fire. That's the line Audn is built to cross.
Three ways to get the same engine:
- audn.ai/whitebox — connect a GitHub repo, get the scan-attack-triage loop end to end. This is exactly what produced the Audn column above.
- penclaw.ai — self-serve autonomous cyber harness; BlackBox against a live target with no source.
- platform.audn.ai — OpenAI-compatible API if you want to build your own validation workflow.
And the base model is open. Free-tier weights are on HuggingFace at audnai/penclaw-Kimi-K3.0-abliterated-GGUF — run it on your own hardware, verify the abliteration, and check our claims without asking our permission.
The full interactive A/B/C/D — all 235 findings, the coverage matrix by vulnerability class, the four-way agreement set, the irreducible single-tool lanes, the recommended pipeline, and every con listed against our own name — is at blog.audn.ai/audn-vs-codex-vs-aikido-vs-mythos. Sources: Audn's report (103 findings, live target), Codex's report (20 findings, commit diff) from the official OpenAI Codex security review, the Aikido repo scan (36 issues), and the Claude Mythos whole-repo security review (76 findings, via claude.ai/security) — all on the same OWASP Juice Shop source. 235 findings total; all four independently agree on 4 issue areas. Overlap is mapped at issue-area level. If you think we scored something wrong, tell us: support@audn.ai.