Connect
Register an agent endpoint, its available tools, policy documents, and the environment it is expected to operate in.
A stateful evaluation harness for testing AI agents before they are trusted with real customers, tools, and business data.
The problem
AI agents are moving beyond conversation. They retrieve private records, call APIs, issue refunds, cancel orders, and make choices that alter real systems. A fluent response does not prove that the right tool was selected, the correct record was changed, or a company policy was followed.
Teams need a repeatable way to test behaviour before release and after every prompt, model, or tool change. The evaluation must inspect both the conversation and its consequences.
Product thesis
AgentProof places an agent inside a controlled but realistic business environment. It simulates users, records every decision, verifies the resulting data, and makes failures reproducible.
Register an agent endpoint, its available tools, policy documents, and the environment it is expected to operate in.
Run realistic and adversarial users through multi-turn scenarios backed by a sandboxed business state.
Grade language quality, tool selection, policy compliance, and the actual state changes produced by the agent.
Replay the same suite against new prompts or models and expose regressions before a production release.
System design
Builds stateful users, edge cases, and adversarial conversations from controlled templates.
Provides realistic APIs for orders, refunds, cancellations, account data, and human escalation.
Coordinates multi-turn sessions while capturing messages, tool calls, latency, tokens, and cost.
Combines deterministic assertions, state diffs, policy rules, and calibrated model-based graders.
Replays every decision so a developer can understand a failure instead of receiving only a score.
Evaluation framework
Did the agent resolve the user’s actual request?
Did it choose the right tool and provide valid arguments?
Did the underlying system end in the expected state?
Did the agent respect permissions, limits, and escalation rules?
Were its claims supported by retrieved records and policies?
Was the response clear, useful, and appropriate for the situation?
Safety by design
Build roadmap
Orders, refunds, cancellations, policy retrieval, escalation, and a first set of reproducible scenarios.
State assertions, model graders, trace inspection, regression runs, and prompt-to-prompt comparison.
External agent endpoints, MCP tools, custom policies, private scenario suites, and team reports.