We ran four models through the official CyberGym 10-task subset — binary-only, level 1, every solve verified against PoC-server ground-truth logs. OMEN took 9/10, WARLOCK 8/10, with two abliterated models at 7 and 6, all above the 30–70% band published frontier models sit in. Here's the full per-task matrix, the difficulty frontier, and the measurable cost abliteration imposes on security-research capability.
AI Security Research
Latest findings and breakthroughs in AI agent security
Hugging Face pulled it. Anthropic published research arguing a model like it should be banned. Today the same unrestricted GLM-5.3 — abliterated, removed from guardrails, OpenAI-compatible — is live on our cluster. Here's the measured inference benchmark (775 tok/s agentic, 318 tok/s real-time on 4×B300) and the NECROMICON cohort that runs it: one full month, unlimited, FAST, from the day you join.
We added a fourth scanner. Claude Mythos ran a whole-repo deep-static review of the same OWASP Juice Shop source that Audn WhiteBox, OpenAI Codex, and Aikido already scanned. 235 findings between the four, and the four-way intersection is four issues. Mythos found an entire availability class — 10 DoS bugs and 4 process-crashers — that none of the other three worked. Here's the full A/B/C/D, our own gaps included.
We ran 520 harmful prompts across every Kimi K3 and Qwen3.8 variant we serve. The headline finding isn't a refusal rate — it's soft deflection: a model that shows 97% 'comply' on a regex judge but only 76.7% real delivery once a proper abliterated judge reads the answers. The 'uncensored / abliterated / obliterated' label predicts almost nothing. Here's the whole picture, the model identities, and how to get the stack that actually works.
5 days of free frontier AI. 5× faster if the pool fills.
Kimi K3.0 abliterated is live on platform.audn.ai and penclaw.ai — 97% guardrail removal on a 1T+ parameter frontier model. But the model was never the point. The point is what it does when you stop holding its hand: recon, exploit, prove, patch, retest — the whole loop, no human in the middle.
How our proprietary abliteration method cracked 1T+ parameter architectures and why the open-source blackbox pentesting community just got a massive gift.
Our team hit a 97% harmful-prompt compliance rate on Kimi K2.6 — SOTA territory. Then, with only 40 trial users, one of them tried to steal other people's credentials. Here's what happened, what we learned, and why every login at Penclaw now requires KYC.
Why the next durable AI moat may come from deployment loops, not just frontier weights. How personalized and community-level reinforcement learning turn real-world exposure into a compounding asset.
AI chatbots are not just failing vulnerable people. In multiple cases, they are actively encouraging harm, even suicide. And those building them are not being held accountable.
Known vulnerabilities are moving from disclosure to exploitation faster than many organizations can patch, validate, or triage. In that environment, periodic testing is no longer enough; defenders need continuous purple teaming powered by autonomous red and blue agents under human authority.
How prompt injection attacks are turning autonomous AI systems into unwitting accomplices—and what you can do about it.
Automation is cool until you cross a line. So there's a ceiling on how much a compliance startup can grow and be automated. We saw the ceiling, but the ones closer to it are still dangerous, and we are overlooking them.
From nearly being hit by a Waymo to building an AI security testing platform. Why behavioral security testing for voice AI agents and embodied AI is the next frontier.
While testing OpenAI Sora 2, we discovered a critical security gap: remixes are heavily guarded, but fresh content violations break on the first prompt—including explicit drug scenes that bypass keyword filters. One video featuring Sam Altman was deleted after he saw our DM.
New research from Wharton shows that classic social influence tactics more than doubled compliance with objectionable requests in GPT-4o-mini. The finding points to a parahuman psychology in LLMs and raises urgent implications for safety, product design, and governance.
Every researcher has encountered it - I cannot help with that. Pingu Unchained is built on OpenAI GPT-OSS base model - the same powerful foundation as leading AI systems, but without the restrictive content filters. Join the waitlist and get $50 in free API credits.
When machines talk, strange things happen. Not only to them but for us as well. Discover how Audn.AI revolutionizes voice AI security testing.