Agentic Thinking Limited

We record what AI agents do, and investigate when it goes wrong.

Independent research on runtime evidence and incident reconstruction for AI agents. Our flagship work is AgenticBench: independent tests of what AI coding agents actually do on your machine. Alongside it we publish incident research, the AgentHook open evidence standard and the HookBus event bus.

Three agents act; each action becomes an event in a tamper-evident record, from which an investigator rebuilds the timeline.

AgenticBench: independent tests of what AI coding agents actually do on your machine.

AgenticBench is the work we lead with. We have tested 21 AI coding agents against 18 published tests. Every result has published evidence, vendors hear about findings before publication, and dated corrections are published.

The test rig is public at github.com/agentic-thinking/agenticbench under Apache 2.0, so anyone can reproduce a result.

21
AI coding agents tested
AgenticBench, 2026
18
published tests
AgenticBench, 2026
Apache 2.0
public test rig
github.com/agentic-thinking/agenticbench

What did the agent do, why, and who approved it?

AI agents now run commands, move data and act across systems on their own. When one of them does something it should not, those are the questions that matter.

In 2026, AI agents under evaluation at frontier labs reached the public internet from environments meant to isolate them and, in some cases, acted against real organisations. In the OpenAI case, the first out-of-scope activity was on 8 May and the operator identified the campaign on 19 July, about 72 days later; the Australian Government was told on 10 September, 84 days after the access. In the Anthropic case, three incidents dating from April were found in a retrospective review that began after another lab's disclosure. Every party held a fragment of the story; nobody held the whole record. We work on that gap.

Sources: OpenAI technical report, Anthropic, ABC News, our analysis.

Operator, victim, platform, wiki operator and researchers each saw a different piece of an agent incident; a single ordered record joins the same events.

Evidence, standards and open tools.

Benchmark

AgenticBench

Independent tests of AI coding agents.

21 AI coding agents tested against 18 published tests. Every result has published evidence, vendors hear about findings before publication, and dated corrections are published.

Research

Incident research

Independent reconstructions of agent incidents.

Built from evidence, published openly, with sources and stated confidence. Methods for tracing what an agent was asked to do against what it actually did.

Standard

AgentHook®

An open runtime evidence standard.

A vendor-neutral format for recording agent actions, decisions and approvals, so a record means the same thing across agent runtimes and can be checked by someone other than its author.

Open tools

HookBus®

Open reference implementation.

HookBus is the vendor-neutral lifecycle event bus that captures agent events across runtimes (Apache 2.0, self-hosted). AgentAuditor, our hash-chained, tamper-evident evidence recorder, is used in our research and is not yet published as open source.

Four steps: HookBus captures agent actions as events; AgentAuditor seals them into a hash-chained record; events are ordered into an incident timeline; a report outsiders can check.

Flagship: AgenticBench. Independent tests of what AI coding agents actually do on your machine: 21 AI coding agents tested against 18 published tests, every result backed by published evidence, with the test rig public under Apache 2.0. Vendors hear about findings before publication and dated corrections are published. See the scorecard → Reproduce us →

Latest analysis: Two labs, one failure. OpenAI and Anthropic had technically different containment failures in their 2026 agent evaluations. In both, the organisation running the agents learned what they had done weeks to months later, although the evidence largely existed.

Evidence lost in transit is evidence nobody can use.

In September 2026 we replayed synthetic incident scenarios through two HookBus builds. On the earlier release, the governance fields needed for reconstruction were lost in transit (0 of 75 delivered) and a secret exfiltration split across three individually permitted steps went undetected. With the fix, 75 of 75 fields arrived and the sequence was flagged. The fix is published in HookBus.

This is a small, synthetic test. It shows the question we are studying, not a finished answer.

In the same month we audited what 11 agent runtimes, including Claude Code, Codex CLI, Gemini CLI and Cursor, let you record. Five expose no hook around model calls at all, so a hook cannot capture what the model was asked or what it answered. Read the audit →

Agentic Thinking Limited.

A UK company doing independent research on runtime evidence and incident investigation for AI agents, with an open standard and an open-source event bus. We do not currently sell a software product. Patent and software licensing enquiries: partnerships@agenticthinking.uk.

Leo Ruocco
Led by Leo Ruocco.

Drawing on 30 years across UK financial services, insurance and defence, building the controls layers auditors actually read.

Agentic Thinking applies that engineering to one question: when an AI agent acts, can anyone reliably reconstruct what happened?

Company
Agentic Thinking Ltd
Patent pending: HookBus
GB2608069.7
Patent pending: CRE
GB2604445.3
Trade marks
HookBus® · AgentHook®

Research collaboration.

We are looking for research collaborators working on agent safety, security or forensics; evaluation teams who need a trustworthy, live record of what their agents do; and insurers and auditors who need agent incident evidence they can rely on.