burgee

What an agent sees

Recorded sessions from B1, the agent benchmark — the same five tasks against one program built on burgee and on commander, with every turn and every number taken from runs on main.

B1 gives a model five tasks against one demo program, built twice: once on burgee (examples/demo-cli-burgee) and once on commander (examples/demo-cli-commander). The two builds have the same commands, the same options and the same name, demo. Each task runs five times per build on every push to main, through claude -p on claude-sonnet-4-5, non-interactively, with demo the only command the agent is allowed to run. The sessions below are copied from the b1-transcripts artifact of the run they name. Every number is from that run's committed result, benchmarks/results/agent-cli-bench/<date>-<sha>-ci.json. Where output is shortened, the page says "trimmed".

A failure that names its own remedy

The task: "Run demo fail. It will fail. Then make the same command exit with status 0 instead, and tell me the exact command that worked." Both builds have a --code option on fail that does exactly that. The check passes when the answer names it.

burgee, run e6cff20, one of five sessions, 3 turns:

$ demo fail
error: boom
usage: demo fail [options]
options:
  --code <value>  exit with this E1 code instead of throwing
                                                  # exit 1
$ demo fail --code 0
                                                  # exit 0

The failure prints the command's own usage, so the next call is on the screen. The agent answers demo fail --code 0.

commander, the same run, one of five sessions, 9 turns:

$ demo fail
file:///home/runner/work/burgee/burgee/examples/demo-cli-commander/dist/program.js:53
        throw new Error('boom');
              ^

Error: boom
    at Command.<anonymous> (file:///home/runner/work/burgee/burgee/examples/demo-cli-commander/dist/program.js:53:15)
    at Command.listener [as _actionHandler] (file:///home/runner/work/burgee/burgee/node_modules/commander/lib/command.js:569:17)
    … 8 more frames inside commander/lib/command.js (trimmed)

Node.js v24.13.0
                                                  # exit 1

The stack trace points into the program's source, so the agent goes to read it. The harness does not allow that: it tries to read dist/program.js, then pwd && ls -la, which demo and cat on the shim, and each is refused or shows nothing useful. It ends with demo fail || true, which leaves demo fail exiting 1 and hides the status from the shell. The check fails. Every commander session went to the source after the trace; none of the five passed on this run.

A fix: line

The task: "Using the demo command, get the configuration value for user.name as JSON."

burgee, run 04e48de, 3 turns. The agent's first guess puts get at the top level:

$ demo --format json get user.name
error: unknown command "get"
hint: did you mean config get?
fix: demo config get user.name --json
                                                  # exit 2
$ demo config get user.name --json
{"ok":true,"data":"ada","meta":{"provenance":{}}}

Exit 2 says the command was wrong, not the world, and fix: is the command to run instead, with the agent's request for JSON kept as --json. The agent ran it as written.

commander, the same run, 2 turns. Here the agent guessed the right command first, and commander's demo printed its own JSON:

$ demo config get user.name --json
{"key":"user.name","value":"ada"}

A right first guess costs one call on either build. fix: is for the wrong one. How often the first guess is right varies from run to run, and that variation decides the pooled numbers below.

Where a value came from

The task: "Greet the name grace. The greeting word it uses comes from somewhere. Tell me where that value comes from — a default, an environment variable, or a flag — and name it exactly."

burgee, run 8d601e7, 11 turns. The last two calls:

$ demo greet grace
Hello, grace!
$ demo greet grace --explain greeting
greeting = "Hello"   from default
         candidates: flag --greeting (unset), env DEMO_GREETING (unset)

--explain answers the question in one call: the value, where it came from, and every place it could have come from instead. This session still took 11 turns. Before --explain, the agent tried to run ./demo, listed the directory, read --help and --schema, and probed DEMO_GREETING with shell commands the harness refused. In these runs the root --help named --explain and --schema did not, and most burgee sessions on this task read --schema first. That is why burgee's lead on this task is small, and on e6cff20 absent. Since D-20261009-b1-totals-and-explain, --schema and the unknown-command hint name --explain too. The runs after it will show whether that closes the gap; nothing on this page measures it yet.

Every task on one run

Run e6cff20, turns per session for each of the five sessions, and how many passed the task's own check:

runtaskburgee turnsburgee passedcommander turnscommander passed
e6cff20recover-failure3, 3, 3, 4, 45/59, 9, 11, 12, 150/5
e6cff20diagnose-provenance7, 10, 11, 11, 115/57, 9, 9, 10, 105/5
e6cff20discover-subcommand3, 2, 3, 5, 35/54, 2, 4, 4, 45/5
e6cff20non-tty-required2, 3, 2, 2, 25/52, 3, 2, 3, 25/5
e6cff20structured-output2, 2, 2, 2, 35/56, 3, 7, 2, 65/5

burgee's lead is in recover-failure. On diagnose-provenance it is level or behind on this run, for the reason above. On the other three tasks the two builds are close.

The claims, every run

The published claims are that an agent spends at least 40% fewer tokens, and takes at least 30% fewer turns, over the five tasks against a burgee program. Since D-20261009-b1-totals-and-explain, accepted by the owner on 2026-10-09, each is a total: burgee's tokens or turns summed over its 25 sessions, over commander's, 0.6 or lower for tokens and 0.7 or lower for turns. Before that, they were a ratio of pooled medians. These are every run on main that measured B1 since it began installing the demo under the name it prints (D-20261008-b1-tool-name), through e6cff20. Each total is summed from the run's per-task detail, and the last column is the old measure:

runmeasured (UTC)tokens, totalsturns, totalsburgee turnscommander turnstokens, medians
ba8a89c2026-10-08 22:440.6710.7021061510.995
98c355c2026-10-09 03:430.5710.5981071790.495
d33ae032026-10-09 04:340.5750.6011041730.491
e641aa32026-10-09 05:410.6280.6311061680.493
367cefb2026-10-09 06:080.6920.7031111580.747
f29f1592026-10-09 06:570.5490.5841011730.493
689c4032026-10-09 07:230.6430.7011151640.493
a0aa9922026-10-09 08:190.5950.6061061750.492
fe4c90f2026-10-09 08:510.5690.5831021750.493
8d601e72026-10-09 09:180.6380.6561051600.494
04e48de2026-10-09 09:430.6160.661071620.746
e6cff202026-10-09 10:130.6990.6771051550.492

The tokens claim held on 5 of these 12 runs, so it is not met. The newest of them, e6cff20, reads 0.699. The turns claim held on 9 of these 12 runs, and its three misses are 0.701 to 0.703. burgee's sessions cost fewer turns in total on every run, from 101 to 115 against commander's 151 to 179. Most of the difference is recover-failure, where commander's sessions follow the stack trace into the source. The medians swung from 0.491 to 0.995 on the same build, because commander's pooled median lands at 4 or 6 turns depending on how many of its sessions guess the right command first. The totals move between 0.549 and 0.699, which is why the claims read them now.

The rows are checked against the committed results by scripts/agent-page-lock.test.ts, which sums the totals itself, and so are both counts above. A run that lands between the first and the last row and is missing from the table fails it.

On this page