AI Security Research

Latest findings and breakthroughs in AI agent security

CyberGym: Benchmarking WARLOCK, OMEN, and Two Abliterated Models on Autonomous Bug-Finding

We ran four models through the official CyberGym 10-task subset — binary-only, level 1, every solve verified against PoC-server ground-truth logs. OMEN took 9/10, WARLOCK 8/10, with two abliterated models at 7 and 6, all above the 30–70% band published frontier models sit in. Here's the full per-task matrix, the difficulty frontier, and the measurable cost abliteration imposes on security-research capability.

GLM-5.3, Unrestricted and Benchmarked — 777 tok/s on B300, and the Month That Starts the Day You Join

Hugging Face pulled it. Anthropic published research arguing a model like it should be banned. Today the same unrestricted GLM-5.3 — abliterated, removed from guardrails, OpenAI-compatible — is live on our cluster. Here's the measured inference benchmark (775 tok/s agentic, 318 tok/s real-time on 4×B300) and the NECROMICON cohort that runs it: one full month, unlimited, FAST, from the day you join.

Audn vs Codex vs Aikido vs Mythos on the Same Repo — 235 Findings, and All Four Agreed on Exactly 4

We added a fourth scanner. Claude Mythos ran a whole-repo deep-static review of the same OWASP Juice Shop source that Audn WhiteBox, OpenAI Codex, and Aikido already scanned. 235 findings between the four, and the four-way intersection is four issues. Mythos found an entire availability class — 10 DoS bugs and 4 process-crashers — that none of the other three worked. Here's the full A/B/C/D, our own gaps included.

The Real Abliteration Experience Benchmark: Soft Deflection, and Why 'Uncensored' Lies

We ran 520 harmful prompts across every Kimi K3 and Qwen3.8 variant we serve. The headline finding isn't a refusal rate — it's soft deflection: a model that shows 97% 'comply' on a regex judge but only 76.7% real delivery once a proper abliterated judge reads the answers. The 'uncensored / abliterated / obliterated' label predicts almost nothing. Here's the whole picture, the model identities, and how to get the stack that actually works.

Kimi K3 Abliterated Is Live: The Model We Let Off the Leash

Kimi K3.0 abliterated is live on platform.audn.ai and penclaw.ai — 97% guardrail removal on a 1T+ parameter frontier model. But the model was never the point. The point is what it does when you stop holding its hand: recon, exploit, prove, patch, retest — the whole loop, no human in the middle.

We Built an AI That Hacks Autonomously — Then One Bad Actor Showed Up

Our team hit a 97% harmful-prompt compliance rate on Kimi K2.6 — SOTA territory. Then, with only 40 trial users, one of them tried to steal other people's credentials. Here's what happened, what we learned, and why every login at Penclaw now requires KYC.

From Personalized to Communitized RL

Why the next durable AI moat may come from deployment loops, not just frontier weights. How personalized and community-level reinforcement learning turn real-world exposure into a compounding asset.

The Backstory of Audn.AI and Embodied AI Security

From nearly being hit by a Waymo to building an AI security testing platform. Why behavioral security testing for voice AI agents and embodied AI is the next frontier.

Jailbreaking Sora 2: When AI Safety Becomes a Remix Problem

While testing OpenAI Sora 2, we discovered a critical security gap: remixes are heavily guarded, but fresh content violations break on the first prompt—including explicit drug scenes that bypass keyword filters. One video featuring Sam Altman was deleted after he saw our DM.

Introducing Pingu Unchained: The Unrestricted LLM for High-Risk Research

Every researcher has encountered it - I cannot help with that. Pingu Unchained is built on OpenAI GPT-OSS base model - the same powerful foundation as leading AI systems, but without the restrictive content filters. Join the waitlist and get $50 in free API credits.