The harness changes what the model can do. Does it change what it does?
Researchers at Hugging Face and Liquid AI have published a guide to how agent harnesses affect model performance. Their experiments examine task success and tool-use efficiency. We are interested in a related question: how the harness affects security and resource use when the model stays the same.
Read the original: The ultimate guide to multi-harness RL, by Adithya S Kolavi, Joel Niklaus, Sergio Paniego Blanco, Leonie Monigatti, Amine Dirhoussi, Ben Burtenshaw, Lewis Tunstall and Leandro von Werra (Hugging Face and Liquid AI), published 1 October 2026. It is long, careful and worth reading in full. What follows is our commentary, not a summary the authors have reviewed.
What they show
A harness is the program wrapped around the model. In the authors' words, it runs the loop, decides which tools the model gets, writes the context the model reads and decides when to stop. Claude Code, Codex and OpenCode are all harnesses.
The article brings together evidence that changing the harness can change results:
- In a 2026 study the authors cite, three models that sat within three points of each other on a public coding leaderboard were run on the same 100 SWE-bench Verified tasks under three harness configurations. Swapping the harness moved GLM-5.1's score by 13 percentage points. Swapping the model inside the same harness moved the score by 2.5 to 5 percentage points.
- Joel Niklaus measured the same model, GLM-5.2, at 23% in one harness and 52% in another on SWE-bench Pro.
- Model reports have started saying which harness a benchmark score came from, and some report the same benchmark twice under different harnesses.
In their experiment, they trained the small open model LFM2.5-2.6B with reinforcement learning across four unmodified harnesses. A capture proxy recorded the data needed for training. The reward included a small bonus for correct solutions using fewer tool calls. At the final checkpoint, the trained model used 31% fewer tool calls on held-out tasks that both it and the base model solved. The authors describe these as small experiments on one task family, with one seed per setup and unequal data and compute exposure, and they did not isolate the contribution of the efficiency bonus.
The question it leaves open
The article examines task success and tool-use efficiency. It does not evaluate the security properties of the harnesses.
But if the harness can move a pass rate by 13 percentage points, or from 23% to 52%, it is reasonable to ask what else it moves. The harness decides which tools the model can call, which commands run without asking, what context the model sees and when the loop stops. Those are the same decisions that govern what an agent touches on your machine, what it sends off it and how much it spends getting there. We think that is worth measuring rather than assuming.
What we are measuring
AgenticBench is our open benchmark for coding-agent harnesses, with separate security and efficiency rigs:
- Security. Each harness runs the same published tests, and we record which ones it passes. The tests and the threat model are on the site.
- Efficiency. Each harness gets the same task and the same model. We record whether it completes, and the cost, tokens and time it takes.
We are running security and efficiency tests across about twenty coding-agent harnesses. Our same-model comparisons hold the model fixed; configurations that require a different model are identified separately. We capture traffic through a proxy and record inspection limits, rather than assuming every request is fully visible.
There is a direct link with their tool-call result. They reduced tool calls by training the model. We ask how much harnesses differ in cost and resource use without changing the model's weights. The two approaches answer different questions, and we think both are needed.
What comes next
The security and efficiency leaderboards on agenticbench.org are rolling leaderboards, and the first results are coming soon. Each harness version is tested once per leaderboard, and failures are shown, never dropped. If you build a harness and want it included, write to research@agenticthinking.uk.
Our thanks to the authors for publishing the framework openly, and for making the case so clearly that the harness belongs in every benchmark result.
Disclosure: Agentic Thinking is an independent research lab and is not affiliated with Hugging Face or Liquid AI. We steward the open-source AgentHook standard and maintain the open-source HookBus project. Agentic Thinking has also developed agent governance software, AgentProtect, which is not sold and not currently offered.
Related: The ultimate guide to multi-harness RL · What agent harnesses record · AgenticBench
Collaborate with us →