If you build an agent runtime or governance tool and want it examined in depth, you can fund an independent evaluation as a research package. Each package has a fixed scope and fee agreed up front; the fee does not depend on the result. We publish our findings unedited, including negative ones, and never as an endorsement. We take on packages selectively, where the work advances our research.
A runtime or governance tool makes a claim about a chain of events. We check each link in that chain, and whether the record of one link actually depends on the link before it.
Whether each recorded decision, approval and action can be tied to the one before it, or whether the record only asserts that it happened.
Whether a deny or an approval requirement actually stops the action at runtime, and whether there are routes around it that the record does not show.
Whether altered, replayed or fabricated evidence is detected, or accepted as genuine by the tool's own verification.
Every finding is tied to a named commit, so anyone can rebuild the same code and repeat the test.
The code under evaluation runs in an isolated environment with no network access.
Two examples from our evaluation work, anonymised. Neither tool nor its developer is named.
A gateway reported actions as completed that it had never performed. Its record looked complete, but the outcome it recorded was not evidence that anything had run.
An evidence digest accepted a forged record. An entry was altered and re-hashed, and the tool's own verification reported it as genuine.
AgenticBench is separate: it never accepts funding from the agents it tests, and coding agents on the benchmark are not eligible for funded evaluations.
Tell us what you build, what its record is meant to show, and what you would like examined.