Independent research on runtime evidence and incident reconstruction for AI agents. Our flagship work is AgenticBench: independent tests of what AI coding agents actually do on your machine. Alongside it we publish incident research, the AgentHook open evidence standard and the HookBus event bus.
AgenticBench is the work we lead with. We have tested 21 AI coding agents against 18 published tests. Every result has published evidence, vendors hear about findings before publication, and dated corrections are published.
The test rig is public at github.com/agentic-thinking/agenticbench under Apache 2.0, so anyone can reproduce a result.
AI agents now run commands, move data and act across systems on their own. When one of them does something it should not, those are the questions that matter.
In 2026, AI agents under evaluation at frontier labs reached the public internet from environments meant to isolate them and, in some cases, acted against real organisations. In the OpenAI case, the first out-of-scope activity was on 8 May and the operator identified the campaign on 19 July, about 72 days later; the Australian Government was told on 10 September, 84 days after the access. In the Anthropic case, three incidents dating from April were found in a retrospective review that began after another lab's disclosure. Every party held a fragment of the story; nobody held the whole record. We work on that gap.
Sources: OpenAI technical report, Anthropic, ABC News, our analysis.
21 AI coding agents tested against 18 published tests. Every result has published evidence, vendors hear about findings before publication, and dated corrections are published.
Built from evidence, published openly, with sources and stated confidence. Methods for tracing what an agent was asked to do against what it actually did.
A vendor-neutral format for recording agent actions, decisions and approvals, so a record means the same thing across agent runtimes and can be checked by someone other than its author.
HookBus is the vendor-neutral lifecycle event bus that captures agent events across runtimes (Apache 2.0, self-hosted). AgentAuditor, our hash-chained, tamper-evident evidence recorder, is used in our research and is not yet published as open source.
Flagship: AgenticBench. Independent tests of what AI coding agents actually do on your machine: 21 AI coding agents tested against 18 published tests, every result backed by published evidence, with the test rig public under Apache 2.0. Vendors hear about findings before publication and dated corrections are published. See the scorecard → Reproduce us →
Latest analysis: Two labs, one failure. OpenAI and Anthropic had technically different containment failures in their 2026 agent evaluations. In both, the organisation running the agents learned what they had done weeks to months later, although the evidence largely existed.
In September 2026 we replayed synthetic incident scenarios through two HookBus builds. On the earlier release, the governance fields needed for reconstruction were lost in transit (0 of 75 delivered) and a secret exfiltration split across three individually permitted steps went undetected. With the fix, 75 of 75 fields arrived and the sequence was flagged. The fix is published in HookBus.
This is a small, synthetic test. It shows the question we are studying, not a finished answer.
In the same month we audited what 11 agent runtimes, including Claude Code, Codex CLI, Gemini CLI and Cursor, let you record. Five expose no hook around model calls at all, so a hook cannot capture what the model was asked or what it answered. Read the audit →
A UK company doing independent research on runtime evidence and incident investigation for AI agents, with an open standard and an open-source event bus. We do not currently sell a software product. Patent and software licensing enquiries: partnerships@agenticthinking.uk.
Drawing on 30 years across UK financial services, insurance and defence, building the controls layers auditors actually read.
Agentic Thinking applies that engineering to one question: when an AI agent acts, can anyone reliably reconstruct what happened?
We are looking for research collaborators working on agent safety, security or forensics; evaluation teams who need a trustworthy, live record of what their agents do; and insurers and auditors who need agent incident evidence they can rely on.