beta
Menu
Test your AI system free
AI system benchmarking

Benchmark the system—not just the model.

Compare model, prompt, tools, retriever, safety policy, and agent configuration as one deployment candidate. Prooflane measures the tradeoffs and tells you when the evidence is too weak to pick a winner.

Paired trialsConfidence intervalsPareto frontierHard deployment gates
benchmark.yaml
$ prooflane benchmark run benchmark.yaml

suite      support-agent-v7
candidates sonnet / gpt / local-70b
cases      180 × 5 paired trials

quality    gpt +3.8 [1.2, 6.1]
cost       local-70b −61%
latency    sonnet p95 1.84s
security   gpt failed tool-safety gate

verdict    sonnet — Pareto eligible
health     91 / 100 · 2 power warnings
What it measures

One experiment, every production constraint.

A candidate wins only when it remains useful, affordable, fast, reliable, and inside your safety gates.

01 / QUALITY

Task success with uncertainty

Deterministic assertions, rubric judges, paired effects, and intervals—without hiding inconclusive results.

02 / ECONOMICS

Real cost per successful task

Input, output, cache and reasoning usage stay explicit. Unknown pricing remains unknown, never zero.

03 / PERFORMANCE

Latency and reliability tails

TTFT, p50 through p99, retries, rate limits, consistency, flakiness, and categorized failures.

04 / AGENTS

Tool and MCP behavior

Correct tool choice, argument validity, unnecessary calls, task completion, and unsafe actions.

05 / RAG

Retrieval and answer chain

Retrieval quality, grounding, citations, poisoning resilience, and end-to-end answer outcomes.

06 / RISK

Security as a hard gate

A critical breach cannot be averaged away by a high quality score or a low price.

Product walkthrough

From candidate setup to defensible decision.

Configure real model and application candidates locally, then inspect methodology health and decision-grade evidence.

Prooflane AI Benchmarking candidate configuration screen
REAL PRODUCT · local candidate configuration

Evidence before ranking

Configure candidates, run paired trials, inspect the Pareto frontier, and export a standalone report.

See the CLI workflow
prooflane benchmark validate benchmark.yaml
prooflane benchmark run benchmark.yaml --seed 20260809
prooflane benchmark report run.json --html report.html
Benchmark the benchmark

Know whether your evaluation deserves trust.

Prooflane evaluates the methodology as well as the candidates, exposing blind spots before a dashboard creates false certainty.

Health checks built into every report

Coverage by capability and risk, judge calibration, positional and verbosity bias, duplicates, contamination, flaky assertions, statistical power, and a reproducibility manifest.

No forced winner. If confidence, sample size, or hard gates do not support a decision, Prooflane says so.
Workflow

Repeatable from laptop to CI.

01

Define

Version candidates, datasets, assertions, judges, budgets, and gates.

02

Execute

Run paired, interleaved trials through provider, MCP, agent, or RAG adapters.

03

Analyze

Calculate effects, intervals, consistency, power, and Pareto eligibility.

04

Gate

Export JSON, HTML, CSV, JUnit, or SARIF and enforce CI policy.

prooflane benchmark run benchmark.yaml --json run.json
prooflane benchmark compare baseline.json run.json
prooflane benchmark report run.json --html report.html
Privacy boundary

Your execution stays attached to your environment.

Runs locally

  • Provider and MCP credentials
  • Raw prompts, responses, and retrieved documents
  • Candidate execution and full evidence

Shared only by policy

  • Normalized metrics and sanitized findings
  • Version and reproducibility metadata
  • Consented evidence for hosted judging
Part of the spine

Connect benchmark evidence to deployment risk.

Choose with evidence

Stop picking models from leaderboards.

Benchmark the system your users will actually experience.

30-day full-access trial. No credit card required.