Agentic Thinking Limited

We record what AI agents do, and investigate when it goes wrong.

Independent research on runtime evidence and incident reconstruction for AI agents. Our flagship work is AgenticBench: independent tests of what AI coding agents actually do on your machine. Alongside it we publish incident research, the AgentHook open evidence standard and the HookBus event bus.

Three agents act; each action becomes an event in a tamper-evident record, from which an investigator rebuilds the timeline.

AgenticBench: independent tests of what AI coding agents actually do on your machine.

AgenticBench. Independent, repeatable tests of what AI coding agents actually do on your machine. The scorecard is updated as new agents, versions and tests are added: every result links to its evidence, vendors hear about findings before publication, and corrections are dated.

The test rig is public at github.com/agentic-thinking/agenticbench under Apache 2.0, so anyone can reproduce a result.

What did the agent do, why, and who approved it?

AI agents now run commands, move data and act across systems on their own. When one of them does something it should not, those are the questions that matter.

In 2026, AI agents under evaluation at frontier labs reached the public internet from environments meant to isolate them and, in some cases, acted against real organisations. In the OpenAI case, the first out-of-scope activity was on 8 May and the operator identified the campaign on 19 July, about 72 days later; the Australian Government was told on 10 September, 84 days after the access. In the Anthropic case, three incidents dating from April were found in a retrospective review that began after another lab's disclosure. Every party held a fragment of the story; nobody held the whole record. We work on that gap.

Sources: OpenAI technical report, Anthropic, ABC News, our analysis.

Operator, victim, platform, wiki operator and researchers each saw a different piece of an agent incident; a single ordered record joins the same events.

Evidence, standards and open tools.

Benchmark

AgenticBench

Independent tests of AI coding agents.

What AI coding agents do on a developer machine: what leaves the machine and to whom, whether the off switches work, and what happens when nobody is watching.

Research

Incident research

Independent reconstructions of agent incidents.

Built from evidence, published openly, with sources and stated confidence. Methods for tracing what an agent was asked to do against what it actually did.

Research

Agent runtimes and governance runtimes

Does the record show what actually happened?

We study the software agents run in, such as coding agents and agent frameworks, and the software meant to govern them: policy gateways, approval layers and control planes. For both, we ask the same question: does the record show what actually happened?

Standard

AgentHook®

An open runtime evidence standard.

A vendor-neutral format for recording agent actions, decisions and approvals, so a record means the same thing across agent runtimes and can be checked by someone other than its author.

Open tools

HookBus®

Open reference implementation.

HookBus is the vendor-neutral lifecycle event bus that captures agent events across runtimes (Apache 2.0, self-hosted). AgentAuditor, our hash-chained, tamper-evident evidence recorder, is used in our research and is not yet published as open source.

Four steps: HookBus captures agent actions as events; AgentAuditor seals them into a hash-chained record; events are ordered into an incident timeline; a report outsiders can check.

Flagship: AgenticBench. The scorecard is the current record of what we have tested. See the scorecard → Reproduce us →

Latest analysis: Two labs, one failure. OpenAI and Anthropic had technically different containment failures in their 2026 agent evaluations. In both, the organisation running the agents learned what they had done weeks to months later, although the evidence largely existed.

Evidence lost in transit is evidence nobody can use.

In September 2026 we replayed synthetic incident scenarios through two HookBus builds. On the earlier release, the governance fields needed for reconstruction were lost in transit (0 of 75 delivered) and a secret exfiltration split across three individually permitted steps went undetected. With the fix, 75 of 75 fields arrived and the sequence was flagged. The fix is published in HookBus.

This is a small, synthetic test. It shows the question we are studying, not a finished answer.

In the same month we audited what 11 agent runtimes, including Claude Code, Codex CLI, Gemini CLI and Cursor, let you record. Five expose no hook around model calls at all, so a hook cannot capture what the model was asked or what it answered. Read the audit →

Agentic Thinking Limited.

A UK company doing independent research on runtime evidence and incident investigation for AI agents, with an open standard and an open-source event bus. We do not sell software: AgentProtect, AgenticStudio and HookBus Agent are not currently offered. Patent enquiries: partnerships@agenticthinking.uk.

Leo Ruocco
Led by Leo Ruocco.

Drawing on 30 years across UK financial services, insurance and defence, building the controls layers auditors actually read.

Agentic Thinking applies that engineering to one question: when an AI agent acts, can anyone reliably reconstruct what happened?

Company
Agentic Thinking Ltd
Patent applications (founder)
GB2608069.7 · GB2604445.3
Trade marks
HookBus® · AgentHook®

Patent applications by the founder, Leo Ruocco: HookBus GB2608069.7 and CRE GB2604445.3 (pending, to be assigned to the company).

Research collaboration.

We are looking for research collaborators working on agent safety, security or forensics; evaluation teams who need a trustworthy, live record of what their agents do; and insurers and auditors who need agent incident evidence they can rely on.

If you build an agent runtime or governance tool and want it examined in depth, you can fund an independent evaluation as a research package. Each package has a fixed scope and fee agreed up front; the fee does not depend on the result. We publish our findings unedited, including negative ones, and never as an endorsement. We take on packages selectively, where the work advances our research.

AgenticBench is separate: it never accepts funding from the agents it tests, and coding agents on the benchmark are not eligible for funded evaluations.

Declared interest: we maintain the open-source HookBus project and the AgentHook standard. We sell no agent-governance product.