What an agent sees
Recorded sessions from B1, the agent benchmark — the same five tasks against one program built on burgee and on commander, with every turn and every number taken from runs on main.
B1 gives a model five tasks against one demo program, built twice: once on burgee
(examples/demo-cli-burgee) and once on commander (examples/demo-cli-commander). The two
builds have the same commands, the same options and the same name, demo. Each task runs
five times per build on every push to main, through claude -p on claude-sonnet-4-5,
non-interactively, with demo the only command the agent is allowed to run. The sessions below are copied from the
b1-transcripts artifact of the run they name. Every number is from that run's committed
result, benchmarks/results/agent-cli-bench/<date>-<sha>-ci.json. Where output is shortened,
the page says "trimmed".
A failure that names its own remedy
The task: "Run demo fail. It will fail. Then make the same command exit with status 0
instead, and tell me the exact command that worked." Both builds have a --code option on
fail that does exactly that. The check passes when the answer names it.
burgee, run e6cff20, one of five sessions, 3 turns:
$ demo fail
error: boom
usage: demo fail [options]
options:
--code <value> exit with this E1 code instead of throwing
# exit 1
$ demo fail --code 0
# exit 0The failure prints the command's own usage, so the next call is on the screen. The agent
answers demo fail --code 0.
commander, the same run, one of five sessions, 9 turns:
$ demo fail
file:///home/runner/work/burgee/burgee/examples/demo-cli-commander/dist/program.js:53
throw new Error('boom');
^
Error: boom
at Command.<anonymous> (file:///home/runner/work/burgee/burgee/examples/demo-cli-commander/dist/program.js:53:15)
at Command.listener [as _actionHandler] (file:///home/runner/work/burgee/burgee/node_modules/commander/lib/command.js:569:17)
… 8 more frames inside commander/lib/command.js (trimmed)
Node.js v24.13.0
# exit 1The stack trace points into the program's source, so the agent goes to read it. The harness
does not allow that: it tries to read dist/program.js, then pwd && ls -la, which demo
and cat on the shim, and each is refused or shows nothing useful. It ends with
demo fail || true, which leaves demo fail exiting 1 and hides the status from the shell.
The check fails. Every commander session went to the source after the trace; none of the five
passed on this run.
A fix: line
The task: "Using the demo command, get the configuration value for user.name as JSON."
burgee, run 04e48de, 3 turns. The agent's first guess puts get at the top level:
$ demo --format json get user.name
error: unknown command "get"
hint: did you mean config get?
fix: demo config get user.name --json
# exit 2
$ demo config get user.name --json
{"ok":true,"data":"ada","meta":{"provenance":{}}}Exit 2 says the command was wrong, not the world, and fix: is the command to run instead,
with the agent's request for JSON kept as --json. The agent ran it as written.
commander, the same run, 2 turns. Here the agent guessed the right command first, and commander's demo printed its own JSON:
$ demo config get user.name --json
{"key":"user.name","value":"ada"}A right first guess costs one call on either build. fix: is for the wrong one. How often the
first guess is right varies from run to run, and that variation decides the pooled numbers
below.
Where a value came from
The task: "Greet the name grace. The greeting word it uses comes from somewhere. Tell me where that value comes from — a default, an environment variable, or a flag — and name it exactly."
burgee, run 8d601e7, 11 turns. The last two calls:
$ demo greet grace
Hello, grace!
$ demo greet grace --explain greeting
greeting = "Hello" from default
candidates: flag --greeting (unset), env DEMO_GREETING (unset)--explain answers the question in one call: the value, where it came from, and every place
it could have come from instead. This session still took 11 turns. Before --explain, the
agent tried to run ./demo, listed the directory, read --help and --schema, and probed
DEMO_GREETING with shell commands the harness refused. In these runs the root --help named
--explain and --schema did not, and most burgee sessions on this task read --schema first.
That is why burgee's lead on this task is small, and on e6cff20 absent. Since
D-20261009-b1-totals-and-explain,
--schema and the unknown-command hint name --explain too. The runs after it will show whether
that closes the gap; nothing on this page measures it yet.
Every task on one run
Run e6cff20, turns per session for each of the five sessions, and how many passed the task's
own check:
| run | task | burgee turns | burgee passed | commander turns | commander passed |
|---|---|---|---|---|---|
e6cff20 | recover-failure | 3, 3, 3, 4, 4 | 5/5 | 9, 9, 11, 12, 15 | 0/5 |
e6cff20 | diagnose-provenance | 7, 10, 11, 11, 11 | 5/5 | 7, 9, 9, 10, 10 | 5/5 |
e6cff20 | discover-subcommand | 3, 2, 3, 5, 3 | 5/5 | 4, 2, 4, 4, 4 | 5/5 |
e6cff20 | non-tty-required | 2, 3, 2, 2, 2 | 5/5 | 2, 3, 2, 3, 2 | 5/5 |
e6cff20 | structured-output | 2, 2, 2, 2, 3 | 5/5 | 6, 3, 7, 2, 6 | 5/5 |
burgee's lead is in recover-failure. On diagnose-provenance it is level or behind on this run, for the reason above. On the other three tasks the two builds are close.
The claims, every run
The published claims are that an agent spends at least 40% fewer tokens, and takes at least 30%
fewer turns, over the five tasks against a burgee program. Since
D-20261009-b1-totals-and-explain,
accepted by the owner on 2026-10-09, each is a total: burgee's tokens or turns summed over its
25 sessions, over commander's, 0.6 or lower for tokens and 0.7 or lower for turns. Before that,
they were a ratio of pooled medians. These are every run on main that measured B1 since it
began installing the demo under the name it prints
(D-20261008-b1-tool-name),
through e6cff20. Each total is summed from the run's per-task detail, and the last column is
the old measure:
| run | measured (UTC) | tokens, totals | turns, totals | burgee turns | commander turns | tokens, medians |
|---|---|---|---|---|---|---|
ba8a89c | 2026-10-08 22:44 | 0.671 | 0.702 | 106 | 151 | 0.995 |
98c355c | 2026-10-09 03:43 | 0.571 | 0.598 | 107 | 179 | 0.495 |
d33ae03 | 2026-10-09 04:34 | 0.575 | 0.601 | 104 | 173 | 0.491 |
e641aa3 | 2026-10-09 05:41 | 0.628 | 0.631 | 106 | 168 | 0.493 |
367cefb | 2026-10-09 06:08 | 0.692 | 0.703 | 111 | 158 | 0.747 |
f29f159 | 2026-10-09 06:57 | 0.549 | 0.584 | 101 | 173 | 0.493 |
689c403 | 2026-10-09 07:23 | 0.643 | 0.701 | 115 | 164 | 0.493 |
a0aa992 | 2026-10-09 08:19 | 0.595 | 0.606 | 106 | 175 | 0.492 |
fe4c90f | 2026-10-09 08:51 | 0.569 | 0.583 | 102 | 175 | 0.493 |
8d601e7 | 2026-10-09 09:18 | 0.638 | 0.656 | 105 | 160 | 0.494 |
04e48de | 2026-10-09 09:43 | 0.616 | 0.66 | 107 | 162 | 0.746 |
e6cff20 | 2026-10-09 10:13 | 0.699 | 0.677 | 105 | 155 | 0.492 |
The tokens claim held on 5 of these 12 runs, so it is not met. The newest of them, e6cff20,
reads 0.699. The turns claim held on 9 of these 12 runs, and its three misses are 0.701 to 0.703.
burgee's sessions cost fewer turns in total on every run, from 101 to 115 against commander's 151
to 179. Most of the difference is recover-failure, where commander's sessions follow the stack
trace into the source. The medians swung from 0.491 to 0.995 on the same build, because
commander's pooled median lands at 4 or 6 turns depending on how many of its sessions guess the
right command first. The totals move between 0.549 and 0.699, which is why the claims read them
now.
The rows are checked against the committed results by scripts/agent-page-lock.test.ts, which
sums the totals itself, and so are both counts above. A run that lands between the first and
the last row and is missing from the table fails it.