Controlled A/B evaluation of Lineman's impact on token consumption and cost, measured on SWE-bench Pro.
All results presented unfiltered. Improvements and degradations are reported equally.
9 tasks · 9 valid · model: sonnet · commit 43d0725
Generated: 12 April 2026 · commit 43d0725
Solution quality was not evaluated on this run. No task in the published artefact carries a SWE-bench test-harness result.
Each SWE-bench Pro task is solved twice by the same Claude model in an isolated environment: once as a baseline (standard Claude Code), and once with the Lineman MCP server active. Both runs receive identical instructions and the same git checkout.
Cost is measured in USD using Anthropic's published token prices at the time of the run. Delta is the percentage difference between the Lineman and baseline costs - a negative delta indicates savings.