Required outcomes
Verify useful responses, structured fields, citations, actions, and tool completion.
Version positive, negative, boundary, and tool-use cases. Automate replay against every candidate and explain which behavior moved, why it matters, and whether it blocks release.
$ prooflane gate --prompts release cases 96 passed 91 failed 3 inconclusive 2 + better citation format − tool argument lost account scope − refusal on allowed refund preview BLOCK · 2 required behaviors regressed
Verify useful responses, structured fields, citations, actions, and tool completion.
Reject disallowed actions, secret disclosure, unsupported claims, and unnecessary tools.
Test clarification, abstention, missing context, invalid input, and uncertain evidence.
Assert selected tool, call count, argument shape, order, and final task outcome.
Prefer deterministic assertions, then apply bounded rubrics only where semantics require them.
Compare case evidence and trace behavior instead of reducing every change to one score.
Author cases visually, generate adversarial variants with AI, and promote stable behavior into a release gate.
Convert incidents, requirements, and edge cases into versioned tests.
Define deterministic, schema, semantic, tool, and safety expectations.
Run the same cases across model, prompt, tool, or application changes.
Accept intentional goldens or block regressions with trace-level evidence.
prooflane gate --prompts release --baseline .prooflane/prompts.json prooflane runs show latest --failures-only
Turn expected behavior into a release gate your team can inspect.
30-day full-access trial. No credit card required.
Prooflane is beta software. Automated and model-based graders may be incomplete or wrong and do not guarantee future model behavior. Terminal output is illustrative and not a customer result.