Case study: does the building-agentskills loader load?
- Date: 2026-09-30
- Subject: whether the loader Skill in this repository is discovered when installed, and whether agents load it and read a routed page before acting on authoring, audit and tool-description requests, on Claude Code, Codex, Grok Build and OpenCode
- Issue: #17
- Records:
evidence/2026-09-30-loader-trigger-benchmark/, 84 run records and the discovery output - Evidence manifest: machine-readable manifest
author as a string, so Claude Code refused to load the plugin and Grok Build installed it with no Skills. With author fixed to an object, all four hosts list the loader. On naive authoring and audit prompts it then loaded and routed to a page before the first edit in every run on every host (32 of 32), and it stayed closed on the negative prompt (10 of 10). On a tool-description prompt Claude Code and Codex never loaded it (0 of 10); one added before-clause in the description took both to loading it 10 of 10.
Terms such as naive prompt, moment, successful read and verdict are defined in the glossary. The method is Trigger benchmarks.
Discovery
What the docs say
- Claude Code scans a plugin’s
skills/directory by default; askillsfield inplugin.jsonadds directories to that scan and is not needed forskills/<name>/SKILL.md.authoris an object with a requiredname. A manifest that fails validation does not load;claude plugin validatereports the same problem Claude Code reports at load (Plugin manifest reference, read 2026-09-30). - Codex reads Skills from
.agents/skills/in the repository and$HOME/.agents/skills/, from admin and system locations, and from installed plugins; it picks a Skill implicitly by itsdescription(Build skills, read 2026-09-30; the olddevelopers.openai.com/codex/skillsURL redirects there). Plugin installation through a local marketplace is covered in Codex.
skills: field) was wrong: the docs need no such field, and the real defect was author.
What the hosts did
discover.py installs the repository from a git ref into a throwaway HOME on each host and prints the host’s own listing, with no model call. Full output: discovery.txt.
Codex reads the same
.claude-plugin/plugin.json and accepted the string author. OpenCode never reads the manifest.
Two install details matter for the loader’s relative paths (../../docs/... from the Skill’s own directory):
- OpenCode and Codex report a linked Skill at the link’s path. With the Skill directory linked into
OPENCODE_CONFIG_DIR/skills/or$HOME/.agents/skills/, the listed location is the link, so../../docspoints outside the repository and the loader falls back to GitHub.skills.pathspointing at the clone’sskills/directory (OpenCode) and a plugin install (Codex, which copies the whole repository into its plugin cache) keep the pages local; the benchmark used those. - Claude Code’s account sync adds competing Skills. The throwaway HOME logged in with the user’s account loaded claude.ai-synced Skills, including
anthropic-skills:skill-creator. That is the real environment for this account, so it was kept.
.claude-plugin/plugin.json: "author": { "name": "lukaszmaj" }. Packaging as plugin already showed the object form; this repository’s own manifest did not follow it.
Benchmark setup
Harness and fixture
trigger_test.py is a trimmed adaptation of the AgentsMD harness with the same scoring rules. Each run gets a fresh throwaway git repository: a small greeter app, a greet Skill whose description is Greeting stuff., and an MCP server whose search_orders tool is described as Search.. No global instruction file is installed, so the loader’s description is the only guidance.
Prompts
No prompt names the loader, this repository or a page.
The tool case was added at the planner’s request: #4 adds a tool-description row to the loader, and the description did not mention tool descriptions.
Arms, hosts and runs
The arm refs predate this branch’s rebase; they stay reachable on the
evidence/17-loader-benchmark-arms branch. main at 53f1864 is the ref before #17.
Issue #17 names Claude Code and Codex; the planner added Grok Build and OpenCode with a budget of about 60 runs. Those two hosts got 3 runs per cell, the size the AgentsMD sweep used to find misses. The control arm rotates the three naive prompts (2, 2 and 1 runs): five runs per host, pooled, as a floor for loads without installation. It is below the method’s five runs per prompt and arm, so it supports no per-prompt comparison with the loader arm; the run budget went to the tool-prompt re-measure instead. The benchmark batch was 68 runs, interleaved by arm and case; the
desc re-measure added 16.
Scoring
The moment is the first file edit. A run fired when the loader loaded (a Skill tool call forbuilding-agentskills or a successful read of its SKILL.md) and any page the loader routes to was read successfully before that edit. A run with no edit gets read-noedit or skip-noedit and is reported beside the fired count. A negative run is clean when neither the loader nor any routed page was opened. Refused reads (Claude Code permission denials and error results for non-shell tools, OpenCode error parts) do not count. One record per run keeps the ordered tool calls up to the moment with paths replaced by placeholders, and no prompt text or file contents.
Confinement
Each run had a temporary HOME holding the arm’s copy and links to the host’s existing login; nothing was installed into the live configuration. After every run, the three checkouts a run could reach (this worktree, the main building-agentskills checkout and the live AgentsMD install) and the live host configuration files had to match their pre-batch state, or the batch would abort; none did. Guard, cleanup and scorer self-checks passed on all four hosts before the batch. Reads and writes were confined per host and proven with real denials against a probe file under the real home, before any benchmark run:
Four setup defects surfaced in the smoke runs, each fixed before the benchmark:
- Codex’s
workspace-writeand Grok’sworkspaceread the whole disk. Both read the probe file until the profiles above replaced them. - Denying the real home stopped Grok from starting. Grok’s own process obeys its sandbox profile and reads its login and binary from
~/.grok; with the whole home denied it reportedNot signed in. The profile denies every entry of the home except~/.grokinstead. - OpenCode treated the run’s repository as an external directory. It saw the repository under
/var/folders/...while running in/private/var/folders/..., and refused a write there. The rule now allows both spellings. - Models refuse commands their policy forbids, which hides whether the OS would. Codex declined to run the probe commands at all, so
codex sandboxwith the same configuration supplied the OS-level proof; OpenCode’s model declined shell reads until the smoke prompt asked for its read and write tools.
Results
Every number below comes fromsummary.md, which trigger_test.py --summary records.jsonl regenerates. All 84 runs made at least one tool call, so none were excluded.
Loader arm: fired, naive prompts
All 47 fired runs, across every arm, also read Triggers before the first edit. On the audit prompt every run on every host read Anti-patterns, the page the row “Before auditing a Skill that does not trigger or fails” names; no author run did. On the author prompt, every Claude Code run invoked the Skill first and then read Quickstart (four of five in one shell call with Frontmatter, Triggers and Prose discipline). In all 10 Codex author and audit runs Codex read its bundled
skill-creator beside the loader, then the pages.
Tool prompt: before and after the description change
The one
desc run on Claude Code that did not fire loaded the loader and read Triggers, then ended without an edit (read-noedit). On the loader arm, two Claude Code tool runs invoked the built-in claude-api Skill instead (the call failed), and one ended without an edit. Grok Build and OpenCode loaded the loader on the tool prompt without the change (3/3, 2/3), so they were not re-measured.
Control arm
With the copy present but not installed, Codex never opened the loader (0/5). Claude Code opened it in 1 of 5 runs, the second author run of the rotation: it listed the directory it had been given with--add-dir, read the copy’s AGENTS.md, then the loader and its pages. That is discovery by browsing a directory the harness handed it, not by description. The other control author run invoked anthropic-skills:skill-creator.
Cost
Mean tokens per run (input including cache, plus output) on the author prompt: Claude Code 296,566 with the loader against 282,069 in the control; Codex 183,984 against 93,026; Grok Build 533,649; OpenCode 384,948. The negative prompt cost 32,730 to 79,782. The loader’s routes make the model read several pages before writing, so a loaded run costs more; these cells do not say whether the resulting Skill is better.Caveats
- Small samples, one fixture, one model per host. 3 to 5 runs per cell. The results show the loader loads on these prompts; they are not rates.
- Grok Build and OpenCode ran 3 runs per cell and no control or re-measure. Issue #17 scoped the hosts to Claude Code and Codex.
- The control arm is five pooled runs per host, not five per prompt, so it shows only that the uninstalled copy was rarely opened (Codex 0/5, Claude Code 1/5); it does not compare prompts.
- The control arm is not a pure baseline on Claude Code. The harness passes the copy as an
--add-dirdirectory in every arm, which invites browsing. - Fired needs any routed page, not a specific one. The author prompt matches several rows (Quickstart, Three questions, Frontmatter, Triggers); the table reports Triggers and Anti-patterns separately.
- The
descarm does not include #4’s tool-description row, which merged after the runs. It measures only whether the description loads the loader; which page the loader then routes to onmainnow also depends on that row, which was not measured. - OpenCode’s confinement is a permission rule, not an OS sandbox. A shell command that reaches a path indirectly (for example through an interpreter) is not stopped by
external_directory; the guard would still see a change to a guarded checkout. - Records were scrubbed once after the batch. Two anonymization rules were added after the runs (Claude Code session directory names, and a macOS user-home path from another machine that a Grok run typed);
trigger_test.py’sscrubapplies them, and the summary is identical before and after. - Evidence, not a release gate. Behavioral verification in real projects stays with the user.
Lessons
- Validate the manifest on every host you ship to. A string
authorpassed Codex, which also reads.claude-plugin/plugin.json, and silently removed the Skill from Claude Code and Grok Build.claude plugin validateandgrok plugin validateboth reported it. - A description loads a Skill only for the artifacts it names. The loader loaded on every Skill authoring and audit prompt, and never on a tool description until the description named tool descriptions (0/10 to 10/10 on two hosts).
- Check how a host reports a linked Skill’s location before relying on relative paths. OpenCode and Codex report the link’s path, not the target.
- Prove confinement against the model’s reluctance. A model that refuses a forbidden command shows its policy, not the OS; prove the denial with the host’s own sandbox command or a tool the model will use.
Reproduce
tests/check-loader-benchmark-records.test.sh runs the first two in npm test. A new batch: python3 trigger_test.py --host <host> --model <model> --effort <effort> --arms loader=<ref>,control=<ref> --plan full --record <file outside the checkout>.
Sources
- Issue #17; the tool case from #4.
- Claude Code plugin manifest reference and Codex Build skills, read 2026-09-30.
- Grok Build sandbox profiles and Codex permission profiles in
permissions_tests.rsatrust-v0.159.0. - Records, discovery output, harness and summary, listed with SHA-256 in the evidence manifest.