Funded evaluations

Funded evaluations.

If you build an agent runtime or governance tool and want it examined in depth, you can fund an independent evaluation as a research package. Each package has a fixed scope and fee agreed up front; the fee does not depend on the result. We publish our findings unedited, including negative ones, and never as an endorsement. We take on packages selectively, where the work advances our research.

What we examine

A runtime or governance tool makes a claim about a chain of events. We check each link in that chain, and whether the record of one link actually depends on the link before it.

  1. Policy decision
  2. Authorisation
  3. Execution
  4. Observation
  5. Durable evidence
  6. Independent verification
01 · The record

Does the record show what happened?

Whether each recorded decision, approval and action can be tied to the one before it, or whether the record only asserts that it happened.

02 · Enforcement

Runtime enforcement and bypass

Whether a deny or an approval requirement actually stops the action at runtime, and whether there are routes around it that the record does not show.

03 · Integrity

Tamper, replay and forgery

Whether altered, replayed or fabricated evidence is detected, or accepted as genuine by the tool's own verification.

04 · Reproducibility

From a fixed commit

Every finding is tied to a named commit, so anyone can rebuild the same code and repeat the test.

05 · Isolation

Network-less execution

The code under evaluation runs in an isolated environment with no network access.

What an evaluation can find

Two examples from our evaluation work, anonymised. Neither tool nor its developer is named.

“Completed” for actions that never ran

A gateway reported actions as completed that it had never performed. Its record looked complete, but the outcome it recorded was not evidence that anything had run.

A forgery that passed verification

An evidence digest accepted a forged record. An entry was altered and re-hashed, and the tool's own verification reported it as genuine.

Separate from AgenticBench

AgenticBench is separate: it never accepts funding from the agents it tests, and coding agents on the benchmark are not eligible for funded evaluations.

Declared interests

Declared interest: we maintain the open-source HookBus project and the AgentHook standard. We sell no agent-governance product.

We have committed to transferring the AgentHook standard to an independent foundation. Our full list of declared interests is on the Trust page.

Enquire

Tell us what you build, what its record is meant to show, and what you would like examined.

partnerships@agenticthinking.uk