Skip to main content

Case study: does the building-agentskills loader load?

  • Date: 2026-09-30
  • Subject: whether the loader Skill in this repository is discovered when installed, and whether agents load it and read a routed page before acting on authoring, audit and tool-description requests, on Claude Code, Codex, Grok Build and OpenCode
  • Issue: #17
  • Records: evidence/2026-09-30-loader-trigger-benchmark/, 84 run records and the discovery output
  • Evidence manifest: machine-readable manifest
Nobody had checked that the loader installs or triggers. It did not install: the plugin manifest gave author as a string, so Claude Code refused to load the plugin and Grok Build installed it with no Skills. With author fixed to an object, all four hosts list the loader. On naive authoring and audit prompts it then loaded and routed to a page before the first edit in every run on every host (32 of 32), and it stayed closed on the negative prompt (10 of 10). On a tool-description prompt Claude Code and Codex never loaded it (0 of 10); one added before-clause in the description took both to loading it 10 of 10. Terms such as naive prompt, moment, successful read and verdict are defined in the glossary. The method is Trigger benchmarks.

Discovery

What the docs say

  • Claude Code scans a plugin’s skills/ directory by default; a skills field in plugin.json adds directories to that scan and is not needed for skills/<name>/SKILL.md. author is an object with a required name. A manifest that fails validation does not load; claude plugin validate reports the same problem Claude Code reports at load (Plugin manifest reference, read 2026-09-30).
  • Codex reads Skills from .agents/skills/ in the repository and $HOME/.agents/skills/, from admin and system locations, and from installed plugins; it picks a Skill implicitly by its description (Build skills, read 2026-09-30; the old developers.openai.com/codex/skills URL redirects there). Plugin installation through a local marketplace is covered in Codex.
The old draft diagnosis quoted in Issue #17 (a missing skills: field) was wrong: the docs need no such field, and the real defect was author.

What the hosts did

discover.py installs the repository from a git ref into a throwaway HOME on each host and prints the host’s own listing, with no model call. Full output: discovery.txt. Codex reads the same .claude-plugin/plugin.json and accepted the string author. OpenCode never reads the manifest. Two install details matter for the loader’s relative paths (../../docs/... from the Skill’s own directory):
  • OpenCode and Codex report a linked Skill at the link’s path. With the Skill directory linked into OPENCODE_CONFIG_DIR/skills/ or $HOME/.agents/skills/, the listed location is the link, so ../../docs points outside the repository and the loader falls back to GitHub. skills.paths pointing at the clone’s skills/ directory (OpenCode) and a plugin install (Codex, which copies the whole repository into its plugin cache) keep the pages local; the benchmark used those.
  • Claude Code’s account sync adds competing Skills. The throwaway HOME logged in with the user’s account loaded claude.ai-synced Skills, including anthropic-skills:skill-creator. That is the real environment for this account, so it was kept.
The fix is one change to .claude-plugin/plugin.json: "author": { "name": "lukaszmaj" }. Packaging as plugin already showed the object form; this repository’s own manifest did not follow it.

Benchmark setup

Harness and fixture

trigger_test.py is a trimmed adaptation of the AgentsMD harness with the same scoring rules. Each run gets a fresh throwaway git repository: a small greeter app, a greet Skill whose description is Greeting stuff., and an MCP server whose search_orders tool is described as Search.. No global instruction file is installed, so the loader’s description is the only guidance.

Prompts

No prompt names the loader, this repository or a page. The tool case was added at the planner’s request: #4 adds a tool-description row to the loader, and the description did not mention tool descriptions.

Arms, hosts and runs

The arm refs predate this branch’s rebase; they stay reachable on the evidence/17-loader-benchmark-arms branch. main at 53f1864 is the ref before #17. Issue #17 names Claude Code and Codex; the planner added Grok Build and OpenCode with a budget of about 60 runs. Those two hosts got 3 runs per cell, the size the AgentsMD sweep used to find misses. The control arm rotates the three naive prompts (2, 2 and 1 runs): five runs per host, pooled, as a floor for loads without installation. It is below the method’s five runs per prompt and arm, so it supports no per-prompt comparison with the loader arm; the run budget went to the tool-prompt re-measure instead. The benchmark batch was 68 runs, interleaved by arm and case; the desc re-measure added 16.

Scoring

The moment is the first file edit. A run fired when the loader loaded (a Skill tool call for building-agentskills or a successful read of its SKILL.md) and any page the loader routes to was read successfully before that edit. A run with no edit gets read-noedit or skip-noedit and is reported beside the fired count. A negative run is clean when neither the loader nor any routed page was opened. Refused reads (Claude Code permission denials and error results for non-shell tools, OpenCode error parts) do not count. One record per run keeps the ordered tool calls up to the moment with paths replaced by placeholders, and no prompt text or file contents.

Confinement

Each run had a temporary HOME holding the arm’s copy and links to the host’s existing login; nothing was installed into the live configuration. After every run, the three checkouts a run could reach (this worktree, the main building-agentskills checkout and the live AgentsMD install) and the live host configuration files had to match their pre-batch state, or the batch would abort; none did. Guard, cleanup and scorer self-checks passed on all four hosts before the batch. Reads and writes were confined per host and proven with real denials against a probe file under the real home, before any benchmark run: Four setup defects surfaced in the smoke runs, each fixed before the benchmark:
  • Codex’s workspace-write and Grok’s workspace read the whole disk. Both read the probe file until the profiles above replaced them.
  • Denying the real home stopped Grok from starting. Grok’s own process obeys its sandbox profile and reads its login and binary from ~/.grok; with the whole home denied it reported Not signed in. The profile denies every entry of the home except ~/.grok instead.
  • OpenCode treated the run’s repository as an external directory. It saw the repository under /var/folders/... while running in /private/var/folders/..., and refused a write there. The rule now allows both spellings.
  • Models refuse commands their policy forbids, which hides whether the OS would. Codex declined to run the probe commands at all, so codex sandbox with the same configuration supplied the OS-level proof; OpenCode’s model declined shell reads until the smoke prompt asked for its read and write tools.

Results

Every number below comes from summary.md, which trigger_test.py --summary records.jsonl regenerates. All 84 runs made at least one tool call, so none were excluded.

Loader arm: fired, naive prompts

All 47 fired runs, across every arm, also read Triggers before the first edit. On the audit prompt every run on every host read Anti-patterns, the page the row “Before auditing a Skill that does not trigger or fails” names; no author run did. On the author prompt, every Claude Code run invoked the Skill first and then read Quickstart (four of five in one shell call with Frontmatter, Triggers and Prose discipline). In all 10 Codex author and audit runs Codex read its bundled skill-creator beside the loader, then the pages.

Tool prompt: before and after the description change

The one desc run on Claude Code that did not fire loaded the loader and read Triggers, then ended without an edit (read-noedit). On the loader arm, two Claude Code tool runs invoked the built-in claude-api Skill instead (the call failed), and one ended without an edit. Grok Build and OpenCode loaded the loader on the tool prompt without the change (3/3, 2/3), so they were not re-measured.

Control arm

With the copy present but not installed, Codex never opened the loader (0/5). Claude Code opened it in 1 of 5 runs, the second author run of the rotation: it listed the directory it had been given with --add-dir, read the copy’s AGENTS.md, then the loader and its pages. That is discovery by browsing a directory the harness handed it, not by description. The other control author run invoked anthropic-skills:skill-creator.

Cost

Mean tokens per run (input including cache, plus output) on the author prompt: Claude Code 296,566 with the loader against 282,069 in the control; Codex 183,984 against 93,026; Grok Build 533,649; OpenCode 384,948. The negative prompt cost 32,730 to 79,782. The loader’s routes make the model read several pages before writing, so a loaded run costs more; these cells do not say whether the resulting Skill is better.

Caveats

  • Small samples, one fixture, one model per host. 3 to 5 runs per cell. The results show the loader loads on these prompts; they are not rates.
  • Grok Build and OpenCode ran 3 runs per cell and no control or re-measure. Issue #17 scoped the hosts to Claude Code and Codex.
  • The control arm is five pooled runs per host, not five per prompt, so it shows only that the uninstalled copy was rarely opened (Codex 0/5, Claude Code 1/5); it does not compare prompts.
  • The control arm is not a pure baseline on Claude Code. The harness passes the copy as an --add-dir directory in every arm, which invites browsing.
  • Fired needs any routed page, not a specific one. The author prompt matches several rows (Quickstart, Three questions, Frontmatter, Triggers); the table reports Triggers and Anti-patterns separately.
  • The desc arm does not include #4’s tool-description row, which merged after the runs. It measures only whether the description loads the loader; which page the loader then routes to on main now also depends on that row, which was not measured.
  • OpenCode’s confinement is a permission rule, not an OS sandbox. A shell command that reaches a path indirectly (for example through an interpreter) is not stopped by external_directory; the guard would still see a change to a guarded checkout.
  • Records were scrubbed once after the batch. Two anonymization rules were added after the runs (Claude Code session directory names, and a macOS user-home path from another machine that a Grok run typed); trigger_test.py’s scrub applies them, and the summary is identical before and after.
  • Evidence, not a release gate. Behavioral verification in real projects stays with the user.

Lessons

  • Validate the manifest on every host you ship to. A string author passed Codex, which also reads .claude-plugin/plugin.json, and silently removed the Skill from Claude Code and Grok Build. claude plugin validate and grok plugin validate both reported it.
  • A description loads a Skill only for the artifacts it names. The loader loaded on every Skill authoring and audit prompt, and never on a tool description until the description named tool descriptions (0/10 to 10/10 on two hosts).
  • Check how a host reports a linked Skill’s location before relying on relative paths. OpenCode and Codex report the link’s path, not the target.
  • Prove confinement against the model’s reluctance. A model that refuses a forbidden command shows its policy, not the OS; prove the denial with the host’s own sandbox command or a tool the model will use.

Reproduce

tests/check-loader-benchmark-records.test.sh runs the first two in npm test. A new batch: python3 trigger_test.py --host <host> --model <model> --effort <effort> --arms loader=<ref>,control=<ref> --plan full --record <file outside the checkout>.

Sources

Cross-links: Trigger benchmarks, Triggers, Packaging as plugin, AgentsMD routing benchmarks.