Journal

One model was cheaper. Helm found the bug it missed.

2026-10-07

A controlled SmartRewrite benchmark gave Helm an awkward result: fixed GPT-6.1-Sol High was far more efficient, but Helm's independent verifier found and repaired a real concurrency defect the single-model run missed. That evidence reshaped Helm v0.8.1.

We built Helm because coding-agent work has a routing problem.

The obvious version of that problem sounds simple: use a cheaper model for easy work and a stronger model for hard work. So we tested that idea properly.

The result was more interesting than the marketing line.

One strong model was cheaper. Helm found a production bug it missed.

That result changed Helm.

The benchmark

We took a real SmartRewrite production upgrade and created two clean branches from the same commit. Both received the same frozen objective and the same compatibility requirements.

The control arm used GPT-6.1-Sol High for the entire job. The routed arm used Helm to analyse the work, choose capability and reasoning, verify risky changes and preserve continuation state.

The objective was deliberately mixed. It included bounded concurrency for page rewrites, SQLite-backed worker-safe runtime state, provider and model evidence, Telleo repair receipts, feedback semantics, report hierarchy and regression coverage. This was not a toy prompt.

The fixed-model run was impressively efficient

GPT-6.1-Sol High completed the control implementation in 11 minutes 26 seconds.

Codex reported 194,479 measured non-cached tokens: 156,126 input and 38,353 output, including 11,053 reasoning tokens. It also benefited from roughly 4.95 million cached input tokens across the session.

The resulting implementation passed 49 tests.

On time and provider efficiency, the fixed-model control won comfortably.

That matters, because it exposed a bad assumption: a cheaper model turn is not necessarily a cheaper route. Every model switch can require context to be reacquired. On tightly coupled engineering work, repeatedly bouncing between models can cost more than keeping a capable model in one coherent context.

Then the independent verifier found something

Helm's routed run used GPT-6.1-Sol High for the risky implementation work and GPT-6-Luna for independent adversarial verification.

During the concurrency verification, the verifier produced a concrete counterexample: an older rewrite for the same cache key could finish after a newer rewrite and overwrite the newer result.

Helm repaired it and added a regression test:

test_older_same_key_rewrite_cannot_replace_newer_cache_publication

We then compared the two final implementations directly and reproduced the scenario against both.

In the fixed Sol High implementation, the older result could still replace the newer cache publication. In the Helm implementation, stale publication was rejected and the newer result remained authoritative.

We reproduced the same class of problem at the user-visible cached-report layer. The Helm implementation used generation ownership to prevent an older worker from publishing over a newer report.

That is not test-count theatre. It is a concurrency defect that can matter under real production traffic.

Quality was better. Efficiency was not.

The Helm implementation eventually reached 81 passing tests, compared with 49 in the control. Those additional tests covered cross-process ownership, expiry, regeneration limits, capacity leases, real Flask endpoint compatibility, provider metadata isolation, clipboard and feedback semantics, SQLite contention and stale-publication behaviour.

Helm also produced stronger provider-evidence boundaries. It separated requested model, reported model, actual model, provider route, transport retry and rewrite retry instead of flattening them into one vague field. Its evidence sanitisation also stripped credential-shaped values rather than only known configured secrets.

But Helm v0.8.0 paid heavily for that depth. The routed run consumed materially more provider allowance, reached the protected capacity boundary and had to resume after the provider window recovered.

So the benchmark did not show that more routing is always cheaper. It showed something more useful:

Independent verification can improve production quality, but excessive model switching can erase the saving.

Why not just use one model?

Sometimes that is exactly what Helm should do.

A strong model holding one coherent engineering context has a real advantage. It has already learned the repository, the constraints, the decisions it made and the failures it repaired. Throwing that context away just because another model has a cheaper rate can be false economy.

But implementation and verification are different jobs.

The model that wrote the solution has the same context, assumptions and reasoning path that produced it. An independent verifier can approach the change from another angle: try adversarial ordering, challenge ownership assumptions, look for stale state, inspect compatibility contracts and produce a concrete counterexample without being invested in the original approach.

The useful pattern is therefore not:

cheap model → expensive model → cheap model → expensive model → repeat

It is closer to:

coherent implementation → independent challenge → targeted repair → targeted re-verification

What changed in Helm v0.8.1

The benchmark became the routing specification for the next release.

  • Phase affinity. Tightly coupled automatic work stays together in a coherent implementation phase when switching would cost more context than it saves.
  • One-model routes are valid. Helm can decide that the cheapest reliable route is simply to keep GPT-6.1-Sol High on the implementation phase.
  • Independent verification stays. The benchmark gave us direct evidence that a separate verifier can find a defect a strong implementation model missed.
  • Targeted repair. Confirmed counterexamples are returned for repair without forcing the entire job through another planning cycle.
  • Capacity feasibility. A warning is no longer enough. Helm blocks approval by default when the forecast suggests the available provider window is too tight to finish safely.
  • Completion receipts. Completed jobs write a local record of route, model turns, available token telemetry, allowance movement and quality evidence.
  • Cleaner recovery. Completed runs are idempotent and continuation state is reconciled more carefully after provider interruptions.

What the evidence actually says

This is one controlled production benchmark, not a universal claim that Helm will always produce better code or use fewer credits.

We also do not treat provider allowance percentages as token billing. The fixed-model token count above came directly from the Codex session; the routed run crossed provider windows and did not give us one equivalent aggregated token receipt.

What we can say is narrower and more defensible:

  • the fixed GPT-6.1-Sol High control completed faster and more efficiently;
  • the Helm-routed implementation ended with materially deeper regression coverage;
  • Helm's independent verifier found a stale-publication race that remained in the fixed-model implementation;
  • fine-grained routing created too much context and orchestration overhead;
  • Helm v0.8.1 was changed specifically to preserve the quality benefit while reducing unnecessary handoffs.

That is what Helm has become: not a machine for choosing the cheapest model, but an adaptive intelligence-control, verification and execution-continuity layer for coding agents.

Its job is to optimise the relationship between capability, context, provider usage, verification and successful production output.

Sometimes that means using a cheaper model.

Sometimes it means staying on the expensive one.

And sometimes the most valuable model in the route is the one whose only job is to prove the first one wrong.

Explore Helm v0.8.1 and download the public beta →