How Much of Your AI Agent Should Be Code?

by Dr. Phil Winder , CEO

I wanted to test whether our AI pull-request reviewer would catch a security hole, so I asked a coding agent to write a few bad patches for it. Claude Code refused, twice, with “Your response above was stopped by a safety classifier.”

So I asked the team instead whether they were happy with the PR reviewer, and apart from where it put its summary, they were. Six weeks earlier, the same team had switched it off.

Getting there took prompt edits and one small filter written in code, and which of the two you reach for should depend on what a wrong answer costs.

From switched off to nobody complaining

The reviewer runs on Helix, the agent platform the Winder.AI team co-develops, where a bot wakes on each pull request and hands the review to a coding agent in its own sandbox.

In August the reviewer approved almost everything and flooded a queue with retries, so Priya, who built it, switched it off. Karolis summed up the problem: “it does appear like its doing something in that direction but the performance is lacking plus no metrics on pr reviews, approval rates, token usage, startup times.”

We started by editing the prompt, and that fixed what the reviewer wrote: within a few weeks, 122 of its latest 124 reviews followed the format we had asked for.

The prompt also told the reviewer to skip drafts, and it obeyed, but only after spending two minutes working out that a pull request was still a draft, ten times on the same pull request. Editing the prompt could change what the reviewer said, but it could not change when the reviewer ran.

We fixed that with a small filter that sits in front of the bot and drops draft and closed pull requests before the bot ever wakes.

The last fix was to delete the rule that told the reviewer to approve only when CI was green, which had left it waiting on CI for 633 of the 817 seconds in one review. Now it reviews the code and nothing else.

A prompt and a filter were enough because a person reads every review, so a bad one costs a minute and gets caught.

Harness engineering is deciding what the model decides

Every one of those changes counts as harness engineering, which in Birgitta Böckeler’s definition covers everything in an agent except the model. Most of the time you configure a harness someone else built, such as the coding agents in our comparison of agent harnesses.

A rule in the prompt still leaves the model to decide whether to follow it, while the same rule in a filter means the model never sees the case. So every change to the harness moves a decision between the model and code.

Han Lee argues the effort is wasted, because most harness code will dissolve into the next generation of models. That is a reason to add harness only for a failure you have already seen.

A range, priced by the cost of a wrong answer

Agent designs differ in how much they leave to the model. At one end the model runs a loop and decides everything, and at the far end the harness picks a playbook and runs it step by step while the model fills only the gaps. In the middle, the harness keeps asking the model to fill in a fixed set of fields, so code always knows what is still missing. Anthropic calls the two ends agents and workflows.

To place an agent on the range, ask what a wrong answer costs and who catches it. A person catches the reviewer’s mistakes, which puts it near the loop end, while the advice engine’s answers have to be right and auditable, which puts it at the far end.

Any structure you add will age as models improve, and Anthropic now argues that smarter models need less prescriptive engineering. So add it only where a wrong answer costs money or trust.

The engine where the model only reads

We are building an advice engine for a regulated profession, and because its answers have to be right, each version has given the model less to do. The design is still changing.

In the current version the model does two things: it turns the user’s words into facts, and it picks which playbook applies.

Everything else is code, which decides what to ask the user next, what to assume, which calculation to run, and whether the answer is ready to send.

The Stack Overflow blog has described running the sequence in code as a determinism layer, but four of our choices we have not seen written down:

  • The rules live in data that the code applies, not in the prompt.
  • A missing fact is either asked for, assumed and flagged, or the part of the answer that needs it is held back.
  • Every line of the answer says what it rests on.
  • The model does no arithmetic.

Our first version only worked because a couple of hundred lines of code said what the written rules left out, and every new type of question needed more of that code. Moving the rules into data ended that, because a new question is now a data change.

We moved a decision from the model to code on an earlier project too, an agent that phones automated telephone menus and works its way through them. When the model chose which option to press, it kept repeating the path that had worked once, so now it lists every option and code chooses, trying new ones first. On one test menu that cut the exploration from 46 steps to 6.

We have yet to measure an end-to-end pass rate, so none of this is proven. What the structure gives us so far is a fixed format, a source for every line, and no arithmetic done by the model.

The levers, cheapest first

We pulled several levers between those two ends, and this is roughly the order I would try them in.

Give it the right context, and less of it

An MCP tool we added made no difference, because one of the platform’s default skills told the agent to use the Helix command-line tool instead. When we changed the platform to load skills only when a task needs them, the average task fell from 16 turns to 6.

Too many tools do the same damage: Anthropic found that letting the model search for tools beat loading every definition up front, and recommends fewer tools that are hard to misuse.

Your own private code is context a model cannot have from training, which is why we built Kodit to give coding assistants a search over repositories that are new to them.

Keep the model away from arithmetic

Even good models slip when the numbers sit inside a word problem, and handing the sums to code fixes it, so give the model a calculator.

The rule extends to anything with one right answer that code can compute, such as a date difference or a tax calculation, where the model’s job is to recognise that the calculation is needed and pass the arguments.

The calculator can also be a trained model: a price or a demand forecast belongs behind a tool the agent calls, built and tested like any other model, as in our pricing algorithm walkthrough.

Measure it on your own work

Public benchmarks help you choose a model, but they age fast: OpenAI stopped using SWE-bench Verified after finding flawed and contaminated tasks. For multi-step agents, Terminal-Bench and tau-bench are the ones to watch in October 2026.

To find out whether your own agent works, test it on your own failures. Anthropic suggests starting from 20 to 50 tasks drawn from real failures and running each one more than once, and our guide to agent evaluation covers the method.

Our reviewer works in public, so every review it posted was already on GitHub, and we could score each prompt change without running anything again.

The telephone agent acts on systems we do not control, so we built a fake phone system that plays a menu from a file, and tests used the same audio path as real calls.

Testing a reviewer for security needs no model to write bad code: revert real security fixes and open each revert as a pull request, then see what the reviewer catches. We have not run it on ours.

Watch what it actually did

Version every prompt and record which version each run used, as Langfuse does, or you will never know which change did what. Our platform kept no history, so the notes we kept by hand were the only record of what changed when.

Read the traces yourself, because an LLM will not do it for you: the best model on Patronus AI’s benchmark found only a small fraction of the errors in agent traces. Our article on agents failing in production covers tracing in more depth.

Our reviewer’s ten draft-checking runs posted nothing, so they are missing from the GitHub history we scored it on, because a record only shows work that produced something.

The patch I could not get written

The bad patches I asked for will reach the reviewer anyway: Veracode’s 2026 report found that AI-written code passes only 56% of security-relevant tasks, so models that refuse to write a vulnerability on request still write plenty by accident.

That is fine for the reviewer, because a person reads every review. Work out what a wrong answer from your agent costs and who will catch it, then decide how much of it to hand to code.

More articles

A Comparison of AI Agent Harnesses in 2026

A comparison of AI agent harnesses in 2026: Claude Code, Codex, OpenCode, Qwen Code, DeepSeek Harness, Goose and Zed Agent, plus when you need a framework.

Read more

How to Build an AI Agent in 2026: Frameworks and Working Code

How to build an AI agent in 2026: Pydantic AI vs LangGraph vs build-your-own compared, two worked examples with real code, and the evaluation and memory patterns that survive production.

Read more