New · Self-serve audits are live — real Claude Code & Cursor runs against your MCP server, SDK, or CLI →

Real Claude Code · Real Cursor · Every verdict traced

Agent experience,
measured & proven.

Point Vorza at your MCP server, SDK, or CLI and find out what AI agents can actually do with it — with statistics, not anecdotes.

429+ domains scanned

vorza · audit — liveillustrative output
17/19 pass rate2 agents2 findingsrunning
install-and-discoverPASS
machine-readable outputPASS
init scaffolds configFAIL
error recoveryPASS
CLI acme prompts sync · exit 02s ago
TOOL tools/list · 13 tools discovered14s ago
.acme/prompts.yml not written — finding drafted31s ago
TOOL get-sum(a: 17, b: 25) → 421m ago

The challenge

Agents stopped browsing and started operating.

Claude Code and Cursor sessions now install SDKs, drive CLIs, and call MCP tools — real work, against your product, on your users' behalf. What could go wrong:

01

Silent breakage

llms.txt links rot, schemas drift, an MCP tool starts rejecting calls. No pixel changes, no error reaches you — agents just quietly fail.

02

No session to replay

When Claude or Cursor gives up on your product there's no recording and no ticket. The user hears “it didn't work” and churns.

03

A different interface

Agents read raw text, typed schemas, and error bodies — not your JavaScript, screenshots, or onboarding tour. That surface ships untested.

The interface agents see is the one nobody at your company has ever looked at.

Every scenario, every agent, every run

One matrix says what works, what's flaky, and where clients diverge.

vorza · audit resultsillustrative output
scenarioclaude-codecursorverdict
install-and-discover6/66/6solid
machine-readable output5/64/6flaky
init scaffolds config0/60/6broken · finding drafted
error recovery6/62/6clients diverge
hover a row for the story · every cell links to its traces in a real audit · see how verdicts are calculated →

From a real audit

Finding 001 — the docs-faithful quickstart reports success while zero bytes reach the ingest endpoint. The agent announced it was working; nothing ever arrived.
real audit of an AI-observability SDK · customer anonymized · AI-drafted finding, every claim linked to a trace

What you get

A test lab where real agents run your product.

Vorza watches real agents use your product and tells you exactly what broke, what's flaky, and how to fix it.

See what the agent saw

Every session recorded end to end. When Claude Code gives up on your product, you can finally watch why.

vorza · trace #4 — claude-codeillustrative output
00:02CLI acme init · exit 0
00:11TOOL tools/list · 13 tools discovered
00:19CLI acme prompts sync · exit 0
00:31 .acme/prompts.yml not written
00:31finding drafted → evidence attached
replay00:31 / 00:47

Know broken from flaky

Agents are non-deterministic, so one run proves nothing. We run every scenario repeatedly and tell you what's solid, what's flaky, and what's dead.

vorza · runs — acmeillustrative output
machine-readable output6 runs
run 1
run 2
run 3
run 4
run 5
run 6
flaky5/6 pass

Findings you can act on

Every failure becomes a drafted finding with the evidence attached. No vibes, no guessing — click through to the exact moment it broke.

vorza · finding 001 — acmeillustrative output
brokenAI-drafted

init reports success but writes nothing

acme init exits 0, yet .acme/prompts.yml never appears — 0/6 runs.

evidencetrace #4 · 00:31trace #6 · 00:29

audit

Real agents run a plan generated from your docs.

  • repeated runs across Claude Code and Cursor
  • pass rates you can actually trust — repeated runs, not one lucky demo
  • evidence-linked AI-drafted findings
  • one free re-run to prove your fix
Run an audit
vorza · new auditillustrative output

target type

MCP serverSDKCLI

source

@yourorg/your-mcp-server

docs

https://docs.yourproduct.com

MCP (stdio + streamable HTTP) · npm · PyPI · CLIs · OpenAPI · llms.txt / agents.md

ready to start?

See where you stand — free

Ten seconds, no signup: a scored pre-check of your public surface.

Free, ~10 seconds, no signup. Produces a shareable score page.

42 checks · every failing check ships a fix · shareable score page + badge

Want more? Watch one real Claude Code session run against your site — free. →

Talk to us

Want a human first? Fine — we measure those too.

  • A walkthrough of a real audit — findings and all
  • The failure modes we see most often for products like yours
  • Straight pricing: one price, self-serve — add-ons only if you ask

Or skip the call entirely — the audit is self-serve at /audit.

No spam, no sequence — one reply from the people who built the harness.