Contents · 1 / 9
Harness choice moved a local model 19.4 points
Harness choice moved Qwen 3.6 27B by 19.4 points on the same 144 tasks: 103/144 in Pi against 75/144 in OpenCode (95% CI 10.4 to 28.5, exact McNemar p = 6.2e-05), at zero provider cost. In the hosted baseline (7 models, 2 harnesses, 1,008 attempts, 153 USD scored), one effect survived Holm correction: GPT-5.6 Sol, again +19.4 points on Pi (adjusted p = 0.009). K3 tied on success across harnesses whilst costing 5.62× more on OpenCode.
Method
Each model ran the same 144 tasks in each harness. Grading is deterministic: a task passes its checks or it doesn't, and every attempt writes a sealed receipt at completion. The analysis plan was registered before the runs, and every comparison reported here carries a Holm correction. Anything outside the registered comparison set is labelled descriptive.1
Every attempt writes a sealed receipt. The tables below come from those receipts, not from notes taken along the way.
The local result
Qwen 3.6 27B runs on hardware the studio owns, so the provider cost for this arm was zero.2 The same model, on the same tasks, produced very different results depending on the harness it sat in.
| Harness | Solved | Rate | Δ vs OpenCode | p |
|---|---|---|---|---|
| Pi | 103/144 | 71.5% | +19.4 | 6.2e-05 |
| OpenCode | 75/144 | 52.1% | ref | ref |
| Claude Code (descriptive) | 102/144 | 70.8% | · | · |
Exact McNemar, paired. 95% CI on the gap: 10.4–28.5. Claude Code sat outside the registered comparison set; its rate is descriptive. Provider cost for the local runs: 0 USD.
Download the results (CSV)Figure 1 as a table
| Pi | 71.5% |
| Claude Code (descriptive) | 70.8% |
| OpenCode | 52.1% |
The gap is 19.4 points and the interval doesn't get near zero. For a model that costs nothing to run, that's a lot of performance to gain or throw away on harness choice alone.
Cost
K3 succeeded at near-identical rates in both harnesses, whilst spending 5.62× more and taking 3.22× the median duration on OpenCode. Judged on success rate alone the two harnesses look interchangeable for K3, which is exactly why the study records cost and duration as well.
The hosted baseline
The hosted arm ran 7 models in 2 harnesses, 1,008 attempts in total, for a scored effective sample of 153 USD. One harness effect survived Holm correction: GPT-5.6 Sol gained 19.4 points on Pi (adjusted p = 0.009). No other harness effect survived. The best cell in the grid was Sol on Pi at 80.6%.
[Full per-model grid and its chart land here when the receipts export is in the repo; same CSS-bar pattern as figure 1, no charting runtime.]
What happens next
A bridge run is underway under a 500 USD ceiling, extending the baseline before the confirmatory gates. Roughly 1,000 USD has gone into the programme so far. If you want to fund more attempts, the donate page sets out exactly what they cost.
Corrections and limitations
Known limits: one local model, two harnesses, one task set. The hosted baseline covers 7 models in 2 harnesses; it doesn't license claims about harnesses it never ran.
Cite this
Tyler, J. (2026). Harness choice moved a local model 19.4 points. Cold Anvil Studios. https://coldanvil.com/research/jackbench-harness-choice/
@misc{tyler2026harness,
author = {Tyler, Jack},
title = {Harness choice moved a local model 19.4 points},
year = {2026},
publisher = {Cold Anvil Studios},
howpublished = {\url{https://coldanvil.com/research/jackbench-harness-choice/}}
}