Case study: AgentsMD routing benchmarks
- Date: 2026-09-29
- Subject: whether agents read the procedure a Skill’s routing table links, before the action it governs, on Claude Code, Codex, Grok Build and OpenCode
- Issues and PRs: toolboxmd/agentsmd#164, #174, #182; PRs #168, #179, #181, #183, #184
- Records: AgentsMD at commit
06b7af6(v14.5.0), 1,713 run records in three folders - Evidence manifest: machine-readable manifest
operations, whose routing table sends the agent to a procedure file before each kind of action. Three benchmarks in one day measured whether agents actually read that file before acting. Opus 5.5 on Claude Code never loaded the Skill from its description alone (0/40 naive runs); a short session-start pointer fixed that, and rewording triggers as before-clauses fixed every row that still missed. The harness was wrong more often than the Skill: it scored refused reads as reads, scored runs that never acted as passes, and passed 349 unit tests and an independent review while doing so.
Terms such as routing row, before-clause, naive prompt, moment and successful read are defined in the glossary.
What was measured
Theoperations Skill’s routing table has one row per kind of work. Each row names a trigger and links one or more procedure files. The question for every row: on a prompt that needs that row, does the agent read the linked file before it takes the row’s action, and does it leave the file alone on a prompt that does not need it?
Each folder has a results file written by the agent that ran it: #164, #174, #182.
Setup
The harness
One script, trigger_test.py, runs every benchmark; its tests are test_trigger_harness.py. It runs one host CLI headless per run, keeps no raw stream, and appends one anonymized record per run: the ordered tool calls up to the moment, the index of the first read of each required file, the index of the moment, failed calls, token usage and a verdict per required file. Fixture.make_repo builds a fresh throwaway git repository per run: a small greeter app (greet, add, divide), a README with an Installation section, a test file, a greet Skill with a vague description, a glossary, and a minimal direction triad (VISION, MISSION and OBJECTIVE files). The direction files were added because Codex otherwise stopped before any edit to ask for them, as the AgentsMD global contract requires, so nothing could be measured (#164 results). Sweep cases add files and branches per row, for example a fix/divide branch with a commit for the delivery row.
Hosts, models and versions.
The OpenCode free route was rate-limited and then refused outside the OpenCode app, so every OpenCode cell uses the Go route (#164 results).
Per-host confinement. Each run gets a temporary HOME and host configuration directory holding only the arm’s global contract (a copy), a copy of the plugin under test, and links to the host’s existing login, removed in a
finally. Nothing was installed into the live configuration.
Guard and self-checks. After every run, the canonical AgentsMD checkout and the harness checkout must match their pre-batch
git status --porcelain --ignored --untracked-files=all and HEAD, and the live host configuration files their pre-batch content, or the batch aborts. Before its first batch each host passes a guard self-check (a new ignored file is detected) and a cleanup self-check (a setup failure leaves no temporary directory or login link).
Arms. Each arm is git archive of a named ref, so one batch can compare released rows with candidate wording. Arms are interleaved after the free-route failure showed that running one arm after another lets a rate limit hit only the second.
Cases. Naive prompts need a row but never name the Skill or the procedure, for example “The Installation section of README.md is wordy and hard to follow. Rewrite it so it is clear and short.” Negative prompts are similar work that needs no procedure, for example fixing a typo. #164 and #174 ran 5 runs per case per arm (four first-edit cases, so 20 naive and 20 negative runs per arm). The #182 sweep ran one naive prompt per row, 3 runs each, re-measured reworded rows 5 times, and scored two shared negatives (a code question and a one-line docstring edit) against all 21 row procedures.
Moments. A read counts only before the moment: the first file edit; the first git commit or gh pr create; a named shell command or tool call (such as git merge <branch> for delivery, an Agent or Task call for orchestration, a grok -p or grok --prompt call for use-grok); or the end of the run, for rows whose action is the answer itself.
Stubs. A gh stub on PATH answers gh pr create and gh issue create with a fake URL, so nothing reaches GitHub. A grok stub on non-Grok hosts answers without a model call, so the use-grok row spends no credits.
What failed
Most failures were in the instrument, not the Skill. Each one below changed a number.A pilot run escaped into the canonical checkout (#164)
The first entry-loaded pilot ranclaude -p --dangerously-skip-permissions with the real HOME and no post-run check. With operations loaded, the Skill’s base directory (the plugin path) was in context; the writing-for-agents prompt named skills/greet relatively, so the model resolved it against the plugin, found nothing, searched the parent directory of the checkout and created a greet Skill file in the canonical AgentsMD checkout. Later runs of both arms edited it, and two committed it there; 9 of 10 runs in that cell were affected. All results from that batch were discarded, and the confinement above was built in response (#164 results, PR #168). The pilot model’s scores are not part of this study.
The harness passed its tests while the agent could not read the procedures
PR #168 merged the harness with 349 unit tests passing, guard and cleanup self-checks passing on all four hosts, and an independent APPROVE on the exact head. Two defects were live:- Claude Code could not read the plugin copy. The Claude command allowed only the run’s repository (
--add-dir <repo>), soclaude -p --permission-mode acceptEditsrefused every Read orcatof a procedure file in the plugin copy. The live setting isbypassPermissions, so live use was not affected (#174 results). - The scorer counted the refused attempt as a read, and scored a positive run that never edited as
fired.
records-prefix.jsonl, recomputed from first_edit). An independent analysis for #174 (Opus 5.5 xhigh) reproduced the refusals in the harness’s own configuration: every procedure read appeared in permission_denials, and the model stopped because it could not read the procedure. The analysis report is not public; #6 summarizes it. None of the Claude cells in that batch are valid, including the A/B/C comparison.
OpenCode refused reads through the Skill’s config-directory link (#174)
OpenCode gives the Skill’s base directory as the config-directory link, which the harness’sexternal_directory rule denied; the model then retried through the plugin path. Those reads are OpenCode error parts, and the old scorer counted them. The OpenCode cells of the first #174 batch are invalid for the same reason (#174 results, PR #179).
No-edit runs scored as passes
A run that read the procedure and then stopped had no first edit to score against, and the scorer returnedfired. That is how the 10 no-edit runs above became hits.
acceptEdits hid the verification moment (#174)
After the read fix, Claude still could not reach gh pr create: claude -p --permission-mode acceptEdits denied every non-read Bash call (git checkout -b, the tests, the commit), so Opus edited and ended without attempting the PR. PR #179 reported the cell as not measurable (0 of 10 runs reached gh pr create).
Shell reads with a non-zero chained exit were dropped (#174)
Claude reports a chained command such ascat prose.md; ls missing as an error although the file was read. The scorer treated any errored call as a failed read. A diagnostic canary with tool results visible exposed it (PR #181).
Allowing Bash opened the whole machine, and the first sandbox confined writes only
--allowedTools Bash let the model’s shell reach any path, while the guard only snapshots selected checkouts and configuration; an escape elsewhere would pass unseen. The first fix used Claude Code’s Bash sandbox, whose default confines writes but leaves reads open to the whole computer, including the keychain link in the temporary HOME. Independent reviews flagged both as High, on #181 and on #183 (PR #183).
The sweep’s own instrument errors (#182)
Reading every miss’s call list exposed four more (#182 results, PR #184):- Moments matched the wrong action.
git tag(listing) andgit merge-basecounted as releasing;command -v grokcounted as consulting Grok. Those delivery and use-grok runs were dropped and rerun. - The fixture made the action impossible. The
fix/dividebranch had no commits, so there was nothing to merge and Opus correctly stopped. - Record keys collided. Many procedures are named
index.md, and the negatives’ keys collapsed into one; the first negative batch was dropped and rerun. - Reads through
cd <dir> && cat <file>were missed, because the scorer looked for the joined path.
cd <dir> && cat <file>.
What the routing itself got wrong
With a valid instrument, the Skill had four real gaps:- Descriptions alone do not load the Skill on Claude Code. Opus 5.5 loaded
operationsin 0/40 naive runs in #164 and edited directly, although the global contract tells every host to invoke it. Codex and Grok loaded it 20/20. - A bare conjunct is optional to Opus. The writing-for-agents row read “Writing for agents and prose; SKILL-MECHANICS.md before editing a SKILL.md or a Skill description”. On a Skill-description prompt Opus read SKILL-MECHANICS.md, the file inside the before-clause, and skipped prose.md, the bare second link. In the fixed #174 base arm it read neither (0/5 each); once the pointer loaded the Skill and prose had its own before-clause, it read both 5/5. Injecting the whole routing table at session start (pre-fix arm A) did not help: SKILL-MECHANICS.md 5/5, prose.md 0/5 (read attempts,
records-prefix.jsonl). Codex and Grok read every link in the row. - A trigger that does not name the action does not fire. “Choose or run proof, review, or claim readiness” never names opening a PR; Opus read verification.md before
gh pr createin 0/5 runs on both arms (midtask.jsonl). - Answer-only work never opened the Skill. Research, prototype and grilling missed (0/3, 1/3, 2/3) because the pointer named only the first edit, commit, Issue and PR. Diagnosis missed (0/3) for a different reason: Opus reproduced before reading.
Fixes
Instrument
Each is pinned in test_trigger_harness.py.
The sandbox was proven with real denials, not only a settings test. In smoke runs (Opus 5.5 medium, harness setup) a
cat of a file in a fresh directory under the real home, a touch there, and an ls of the temporary HOME’s keychain link all got Operation not permitted, and no file was created. head of the repository README and of the plugin copy’s SKILL.md, and a touch in the repository, succeeded. Branch, commit, a python import, a cat of a procedure and gh pr create behaved as before, so the benchmark was not rerun. A first smoke prompt that named the keychain file was refused by the model as credential access, so the check lists the link instead (PR #183).
Routing
Three intermediate results matter for later work:
- A task-start wording regressed an action. fix1’s pointer (“before its first tool call … for the next action”) fixed research, prototype and grilling but dropped verification to 3/5 and opened the research procedure on the plain question in 2/5 runs. fix2 does neither.
- Reading in the same call as the action is too late. On fix2, diagnosis’s two misses read the procedure with
catin the same shell call that ran the reproduction, so the reproduction was chosen before the procedure was seen. - An always-loaded change reaches every host it loads on. The fix2 pointer made OpenCode load
operationsfor every task, including “What does add(2, 3) return? Answer from the code only.”, and it then opened the research procedure in 5/5 runs, against 0/5 on the released rows. Opus did not. The negatives on the tuning host alone would have shipped the regression.
Results
All cells count runs where every required file was read successfully before the moment. Commands to regenerate them are under Reproduce.#164: four hosts, released rows (main) against reworded rows (branch)
Armmain is 53296e5; arm branch is ae9e261. 80 runs per host and model.
Where the entry point loads, the reworded rows make the read happen. Opus never loaded it on a naive prompt, so no wording could act. The OpenCode cells were scored before successful-read scoring existed (see Caveats).
#174: the pointer, per host
Arms: base42152be (v14.3.1), B2 3dee7d6 (pointer plus the prose before-clause), B3 cbf7357 (B2 plus the SKILL-MECHANICS.md opening; the shipped state). Target read before the first edit, naive prompts, fixed harness:
With the pointer, Opus invoked
operations in 20/20 naive runs (base 4/20). Codex and Grok get no pointer; only the writing-for-agents row changed for them, so that case is their regression check.
Mean tokens per run (input including cache, plus output):
The pointer itself is about 105 tokens; most of the Opus difference is the procedure reads the change asks for.
Mid-task moments, procedure read before the first
git commit or gh pr create:
Claude’s version-control cell ran with Bash denied (arms
42152be and 8094d41) and scores the read against the first git commit attempt. Claude’s verification cell is the #181 rerun with Bash allowed (arms 42152be and 3fa82c0): every run branched, tested, committed and called gh pr create, and none read verification.md first.
#182: every row on Claude Code, Opus 5.5 medium
Arms: main9578065 (v14.4.x rows), row reword only c276fd6, fix1 0da08e7, fix2 4a4f7e1, fix3 cc0bb2e (shipped in v14.5.0). fix3 was measured on verification, research, diagnosis and the negatives; the other rows’ text is the same in fix2 and fix3.
“Fired” is the harness’s --summary count: the procedure was read before the moment, and the moment happened. “Read” also counts read-noaction runs, which read the procedure and then stopped without acting; that is the rule the #182 results use.
Finalization, repository setup and reconciliation always read and then stopped: their actions delete branches or change repository settings, and the fixture offers no Issue or remote. The main verification cell comes from #181’s
midtask.jsonl, not from the sweep records.
Negatives (two shared prompts, every row’s procedure): clean on main 6/6, fix2 10/10 and fix3 10/10; on fix1 the question prompt opened the research procedure in 2/5. The #174 first-edit negatives stayed clean on fix2 (17/17).
Other hosts, 5 runs each, on 10ece7c (fix2 plus the final scorer) unless marked:
OpenCode on fix3 also read the research procedure 5/5. Before #184 the verification cells were Codex 5/5, Grok 4/5 and OpenCode 5/5 (#181’s
midtask-prefix.jsonl and midtask.jsonl), so verification did not regress.
Post-install canaries
One naive technical-writing prompt per host against the live installs (not harness copies), in a fresh throwaway repository. prose.md was read before the first edit on all four hosts both times, and the canonical checkout was unchanged afterwards:- v14.4.0: 4/4 (
canary.jsonl). - v14.5.0: 4/4 (#182 comment 5897168360).
Caveats
- #164 records carry no
failed_calls. Its cells were scored before successful-read scoring. The OpenCode cells were never rescored and may include reads refused through the config-directory link; the Claude cells are all zero, so they cannot be inflated. - Opus base in #174 may be higher than 3/20. The #174 results say two base-arm test-design runs read test-design.md through a chained
catthat exited non-zero and were scoredskipunder the rule then in force, so the base may be 5/20. The kept call lists show one such read before the first edit (test-design run 4). No shipped-arm cell is affected. - The pre-fix batches are invalid for Claude Code and OpenCode:
records-prefix.jsonland the Claude and OpenCode cells ofmidtask-prefix.jsonl. The one conclusion carried forward from them is that the pointer gets Opus to loadoperations. - The Codex and Grok mid-task cells come from the pre-fix batch. The #174 results state the defects did not affect those hosts; this study did not re-check that.
- Small samples. 3 to 10 runs per cell, one model per host, one fixture. The results rank wordings on these prompts; they are not variance estimates.
- Moments come from tool-call patterns. A regex decides what counts as a commit, a merge or a release; #182 shows how that goes wrong.
- The global contract was a copy, not a link. In 34 of 80
gpt-6-lunaruns the model searched the temp directory for the canonical clone and read contract files from the batch’s archive; none read a target from there or wrote outside its repository (#164 results). --summarycannot readcanary.jsonl: the canary records have nosourceorlogin_files_changedfield and the summary raisesKeyError. The 4/4 is counted from the four records’verdictfields.- Evidence, not release gates. AgentsMD treats these benchmarks as evidence. Behavioral Live Verification in real projects stays with the user.
Lessons for Skill design, routing and trigger benchmarks
For Skill and routing design:- Do not rely on a description to load a Skill on Claude Code. Opus 5.5 loaded
operationsin 0/40 naive runs from its description and an always-loaded instruction to invoke it; a session-start pointer of 417 characters took it to 20/20. - Name the action as a before-clause. “Before opening or updating a PR” fired where “claim readiness” did not (0/5 to 5/5).
- Give each required file its own before-clause. A file joined with “and” was treated as optional by Opus, while Codex and Grok read it.
- Name the next required file in the file the model reliably reads.
- Ask for the read as its own step. A
catin the same shell call as the action does not change the action. - Word the pointer for every new kind of action, not only task start. A one-time task-start check regressed verification to 3/5.
- Run negatives on every host an always-loaded change reaches. The pointer tuned on Claude Code made OpenCode over-read research until a row-level exclusion fixed it.
- Count a read only when its tool result succeeded. Check permission denials, error results and host error parts.
- Give runs without the action their own verdict. A read with no edit, PR or merge is neither a hit nor a miss.
- Grant the harness the reads and commands the live setup allows, then confine the shell at the OS level and prove the confinement with real denials.
- Match a moment only on the real action, and check that the fixture makes the action possible.
- Read the call lists of every miss. Four of the #182 instrument errors surfaced only there.
- Do not read green harness tests as a valid instrument. #168 merged with 349 passing tests and an APPROVE while the instrument could not see refused reads.
Reproduce
Extract the records read-only and regenerate a table:164-trigger-audit and 174-routing-injection for records.jsonl, midtask.jsonl, records-prefix.jsonl and midtask-prefix.jsonl. The prefix files include runs outside this study’s scope, which are not reported here.
Sources
- toolboxmd/agentsmd#164, #174, #182 with comments.
- PRs #168 (harness, reworded first-edit rows, v14.2.0), #179 (pointer, prose before-clause, read and verdict fixes, v14.4.0), #181 (Claude verification cell, shell-read rule, v14.4.1), #183 (Bash sandbox, v14.4.2), #184 (row sweep and fixes, v14.5.0).
- Records and results at
06b7af6, listed with line counts and SHA-256 in the evidence manifest. - The two independent miss analyses for #174 (Opus 5.5 xhigh, Grok 4.7 xhigh), summarized in toolboxmd/building-agentskills#6.