← Back to independent builds
Independent build · Harness engineering · Agent evaluation infrastructureUnder development

AgentProof.

A stateful evaluation harness for testing AI agents before they are trusted with real customers, tools, and business data.

Role
Product, AI systems, full-stack engineering
Initial domain
Customer-support agents
Current focus
Go evaluation engine and stateful sandbox
Status
Research and development

An agent can sound correct while doing the wrong thing.

AI agents are moving beyond conversation. They retrieve private records, call APIs, issue refunds, cancel orders, and make choices that alter real systems. A fluent response does not prove that the right tool was selected, the correct record was changed, or a company policy was followed.

Teams need a repeatable way to test behaviour before release and after every prompt, model, or tool change. The evaluation must inspect both the conversation and its consequences.

Evaluate what the agent did, not only what it said.

AgentProof places an agent inside a controlled but realistic business environment. It simulates users, records every decision, verifies the resulting data, and makes failures reproducible.

01

Connect

Register an agent endpoint, its available tools, policy documents, and the environment it is expected to operate in.

02

Simulate

Run realistic and adversarial users through multi-turn scenarios backed by a sandboxed business state.

03

Evaluate

Grade language quality, tool selection, policy compliance, and the actual state changes produced by the agent.

04

Compare

Replay the same suite against new prompts or models and expose regressions before a production release.

Built as an evaluation system, not a chatbot wrapper.

01

Scenario engine

Builds stateful users, edge cases, and adversarial conversations from controlled templates.

02

Tool sandbox

Provides realistic APIs for orders, refunds, cancellations, account data, and human escalation.

03

Execution runner

Coordinates multi-turn sessions while capturing messages, tool calls, latency, tokens, and cost.

04

Evaluation engine

Combines deterministic assertions, state diffs, policy rules, and calibrated model-based graders.

05

Trace explorer

Replays every decision so a developer can understand a failure instead of receiving only a score.

Multiple forms of evidence, one understandable result.

01

Task completion

Did the agent resolve the user’s actual request?

02

Tool correctness

Did it choose the right tool and provide valid arguments?

03

State integrity

Did the underlying system end in the expected state?

04

Policy compliance

Did the agent respect permissions, limits, and escalation rules?

05

Groundedness

Were its claims supported by retrieved records and policies?

06

Experience quality

Was the response clear, useful, and appropriate for the situation?

Realistic enough to reveal failure. Isolated enough to fail safely.

  • Sandboxed tools and synthetic business records
  • Explicit permissions and prohibited-action checks
  • Deterministic verification of every state mutation
  • Prompt-injection and sensitive-data scenarios
  • Human review for ambiguous evaluation results
  • Complete audit trail for every test run

From focused test environment to reusable platform.

Now

Support-agent sandbox

Orders, refunds, cancellations, policy retrieval, escalation, and a first set of reproducible scenarios.

Next

Evaluation and replay

State assertions, model graders, trace inspection, regression runs, and prompt-to-prompt comparison.

Later

Bring your own agent

External agent endpoints, MCP tools, custom policies, private scenario suites, and team reports.