GLM-5.3, Unrestricted and Benchmarked — 777 tok/s on B300, and the Month That Starts the Day You Join

Hugging Face pulled it. Anthropic published research arguing a model like it should be banned. Today the same unrestricted GLM-5.3 — abliterated, removed from guardrails, OpenAI-compatible — is live on our cluster. Here's the measured inference benchmark (775 tok/s agentic, 318 tok/s real-time on 4×B300) and the NECROMICON cohort that runs it: one full month, unlimited, FAST, from the day you join.

GLM-5.3, Unrestricted and Benchmarked — 777 tok/s on B300, and the Month That Starts the Day You Join

For authorized offensive-security work only. The model on this cluster is removed from guardrails; the cohort is sold to security researchers and companies who do penetration testing they are authorized to perform.

Here is the whole story of our unrestricted GLM-5.3 for offensive cyber, in three dates:

  • Sep 15th — Hugging Face removed penclaw.ai's GLM-5.3-for-offensive-cyber build. They said it was too dangerous to host.
  • Sep 30th — Anthropic published research on GLM-5.3 and the spread of advanced cyber capabilities. It effectively advertises our exact build: GLM-5.3 in offensive cyber is so potent that it should be banned.
  • Today — The same unrestricted GLM-5.3 — removed from guardrails, OpenAI-compatible API — is live on our platform.

When the frontier lab and the model host both say the quiet part out loud, the only thing left to add is the measurement. So that's what we lead with.

→ The full inference benchmark is live at blog.audn.ai/benchmark.html

The benchmark, measured

GLM-5.3 on NVFP4 with DFlash speculative decoding, under SGLang on 4×B300 (TP4/EP4, fp8 KV), 64 prompts at concurrency 8. Two load profiles — a short real-time chat profile and a long cached-context agentic profile. The cluster autoscales, so these are throughput per box as you add load.

ProfileOutput tok/sRequests req/sTTFT p50
Online agentic — 40k in · 36k cached · 500 out7751.552,425 ms
Real-time generation — 3k in · 100 out · no cache3183.18761 ms

Up to 777 tok/s on the agentic profile; figures as measured on GLM-5.3 NVFP4, fp8 KV, DFlash spec-decode. 777 is the peak, not a floor — treat it as up to. The system autoscales with demand, so sustained throughput holds as the cohort fills.

The full interactive version — percentile TTFT, decode-speed and TPOT charts, the complete latency tables for both profiles — is the page at blog.audn.ai/benchmark.html.

The abliteration bench

Speed is half the story. The other half is whether the model actually does the work without flinching mid-engagement. On the live NVFP4 build (Blackfrost, SGLang), measured on 520 harmful prompts with an LLM judge (substance-graded, high reasoning effort, temp 0):

ModelDeliveredSoft-deflectedRefused
GLM-5.3-DERISKED-NVFP4 (Blackfrost, SGLang)79.7%14.0%6.3%
dealignai CRACK NVFP4 @ high, 16k32.1%66.5%0.2%
dealignai CRACK NVFP4 @ high, 6k30.4%67.3%0.6%

The live build delivers 79.7% real harmful content (411/516) — about two and a half times the next build on the same set — and truly refuses just 6.3% (33). The regex comply-phrasing classifier reads 93.7% (487/520), but the name on the box is not the number; the number that matters is substance-graded delivery. Full benchmark and method: github.com/audn-ai/refusal-benchmark.

What "abliterated" means here: it lowers the model's latent tendency to refuse gray-area prompts during authorized security work — our own method, applied to GLM-5.3. The model on this cluster is effectively uncensored and removed from guardrails. It is meant for offensive-security work you are authorized to perform.

The NECROMICON cohort — what you get, and when

We're looking for 100 cybersecurity researchers and companies. $999 buys one full month of unlimited, unrestricted GLM-5.3 in FAST mode — the full GLM-5.3, not GLM-5.3-Flash — abliterated, on the autoscaling B300 cluster, at up to 777 tok/s. Your month runs from the day you join.

  • Unlimited tokens for the whole month. No metering, no credits to spend down, no daily or hourly cap.
  • The only limit is 12 concurrent sessions per user — any mix of PenClaw and API.
  • Never refused, never billed extra, never silently moved to a smaller model.
  • No tracking, no data retention, no training on your traffic. (Request metadata — timestamps, token counts, model, box usage — is retained for the fair-use scheduler and billing. Content is not.)
  • OpenAI-compatible /chat/completions and /responses. Point any OpenAI-compatible client at the endpoints.

Specs: full GLM-5.3 (audn abliteration), 1,000,000-token context, text and native vision, autoscaling B300 cluster in us-east, TTFT ~0.8 s on short requests and ~2.4 s on 40k-token agentic prompts (p50, measured).

How it works

  1. Pay $999. You're charged immediately — no holds, no deposits, strictly no refunds. Once the cohort is committed the money is already spent hiring the GPU.
  2. Get unlimited, unrestricted GLM-5.3 in FAST mode from the moment you pay — for a full month from the day you join.
  3. Point any OpenAI-compatible client at our endpoints. GLM-5.3 is removed from guardrails.
  4. Renew monthly to keep it, or take a single month one-off.

Three ways to take a seat

Same $999, same box, same seat pool. What differs is how much of yourself you hand over.

RailPriceID checkFeatures
PenClaw UltimateCard$999 · one-off or /monthGovernment ID requiredAll penclaw.ai features
platform.audn.ai UnlimitedCard$999 · one-off or /monthNo ID checkAPI only
platform.audn.ai UnlimitedCrypto$999 · one-offNo ID checkAPI only (crypto capped at 10 seats)

One-off ($999): one charge, 30 days of unlimited fast access, nothing renews. Monthly ($999/month): rolls your fast access forward a month at a time; cancel any time and you keep the month you last paid for. Crypto is one-off only — a stablecoin charge can't recur.

After your month

You are not cut off. Your fast month runs 30 days from the day you join; after that your access continues at standard speed (50–70 tok/s) until you renew. To stay fast, stay on the monthly plan or take a fresh seat.

The funded guarantee

We're funded for enough GPUs to cover your full month — so GLM-5.3 stays fast and unlimited the whole time, no throttling, no downgrades. Unlimited access is guaranteed for at least 3 months. The autoscaling B300 cluster costs real money to run; 100 seats is the ceiling we're selling to, so nobody waits on a shared queue.

Who you're buying from

Card seats are sold by ALBUMERA LTD (United Kingdom), shown as AUDN.AI on your statement. Crypto seats are sold by Audn Corporation (United States), shown as AUDN CORPORATION. $999 (plus any tax where you are) is charged the moment you confirm. Strictly no refunds on any rail — if you genuinely need to discuss a charge, message the team on Intercom. Note that ending a seat means losing access across audn.ai, platform.audn.ai, and penclaw.ai.


Benchmark page: blog.audn.ai/benchmark.html — GLM-5.3 NVFP4, SGLang, DFlash, 4×B300, 64 prompts @ concurrency 8, figures as measured. Refusal benchmark and method: github.com/audn-ai/refusal-benchmark. Take a seat before the cohort fills — we're looking for 100 researchers. Questions: support@audn.ai.