New · Self-serve audits are live — real Claude Code & Cursor runs against your MCP server, SDK, or CLI →

Accuracy

The number behind the guarantee.

A finding is only worth trusting if it's usually right. We measure that against a labeled benchmark — real defects, known answers — and report how often Vorza catches them (recall) and how often a reported finding is a true defect (precision), per class of defect.

Benchmark in progress

The first labeled benchmark run hasn't published yet. This page shows measured recall and precision the moment it does — we don't put a number here we haven't earned. Placeholder — replaced by B's benchmark artifact (recall/precision per defect class) once the first labeled benchmark run completes. The /accuracy page renders whatever this file contains; no code changes needed to publish real numbers.

How verdicts are calculated →