Contents · 1 / 9
01TL;DR 02Method 03The local result 04Cost 05The hosted baseline 06What happens next 07Corrections 08Cite this 09Footnotes
Jackbench · publication one

Harness choice moved a local model 19.4 points

TL;DR

Harness choice moved Qwen 3.6 27B by 19.4 points on the same 144 tasks: 103/144 in Pi against 75/144 in OpenCode (95% CI 10.4 to 28.5, exact McNemar p = 6.2e-05), at zero provider cost. In the hosted baseline (7 models, 2 harnesses, 1,008 attempts, 153 USD scored), one effect survived Holm correction: GPT-5.6 Sol, again +19.4 points on Pi (adjusted p = 0.009). K3 tied on success across harnesses whilst costing 5.62× more on OpenCode.

Method

Each model ran the same 144 tasks in each harness. Grading is deterministic: a task passes its checks or it doesn't, and every attempt writes a sealed receipt at completion. The analysis plan was registered before the runs, and every comparison reported here carries a Holm correction. Anything outside the registered comparison set is labelled descriptive.1

Receipts

Every attempt writes a sealed receipt. The tables below come from those receipts, not from notes taken along the way.

The local result

Qwen 3.6 27B runs on hardware the studio owns, so the provider cost for this arm was zero.2 The same model, on the same tasks, produced very different results depending on the harness it sat in.

Table 1 · Qwen 3.6 27B, 144 tasks per harness
HarnessSolvedRateΔ vs OpenCodep
Pi 103/144 71.5% +19.4 6.2e-05
OpenCode 75/144 52.1% ref ref
Claude Code (descriptive) 102/144 70.8% · ·

Exact McNemar, paired. 95% CI on the gap: 10.4–28.5. Claude Code sat outside the registered comparison set; its rate is descriptive. Provider cost for the local runs: 0 USD.

Download the results (CSV)
Pi
71.5%
Claude Code (descr.)
70.8%
OpenCode
52.1%
Figure 1 as a table
Pi71.5%
Claude Code (descriptive)70.8%
OpenCode52.1%
Figure 1 · Solve rate by harness, Qwen 3.6 27B. Same numbers as table 1.

The gap is 19.4 points and the interval doesn't get near zero. For a model that costs nothing to run, that's a lot of performance to gain or throw away on harness choice alone.

Cost

K3 succeeded at near-identical rates in both harnesses, whilst spending 5.62× more and taking 3.22× the median duration on OpenCode. Judged on success rate alone the two harnesses look interchangeable for K3, which is exactly why the study records cost and duration as well.

The hosted baseline

The hosted arm ran 7 models in 2 harnesses, 1,008 attempts in total, for a scored effective sample of 153 USD. One harness effect survived Holm correction: GPT-5.6 Sol gained 19.4 points on Pi (adjusted p = 0.009). No other harness effect survived. The best cell in the grid was Sol on Pi at 80.6%.

[Full per-model grid and its chart land here when the receipts export is in the repo; same CSS-bar pattern as figure 1, no charting runtime.]

What happens next

A bridge run is underway under a 500 USD ceiling, extending the baseline before the confirmatory gates. Roughly 1,000 USD has gone into the programme so far. If you want to fund more attempts, the donate page sets out exactly what they cost.

Corrections and limitations

No corrections recorded. If something in here is wrong, tell me. Fixes get logged in this section with a date.

Known limits: one local model, two harnesses, one task set. The hosted baseline covers 7 models in 2 harnesses; it doesn't license claims about harnesses it never ran.

Cite this

Tyler, J. (2026). Harness choice moved a local model 19.4 points. Cold Anvil Studios. https://coldanvil.com/research/jackbench-harness-choice/

@misc{tyler2026harness,
  author       = {Tyler, Jack},
  title        = {Harness choice moved a local model 19.4 points},
  year         = {2026},
  publisher    = {Cold Anvil Studios},
  howpublished = {\url{https://coldanvil.com/research/jackbench-harness-choice/}}
}

  1. Descriptive means reported without a significance claim. Claude Code's 70.8% in table 1 is the example.
  2. Zero provider cost covers inference. Electricity is real, but it isn't metered per attempt.