Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/rstackjs/agent-skills/rstack-skill-evaluatornpx skills add rstackjs/agent-skills --skill rstack-skill-evaluatorgit clone --depth 1 https://github.com/rstackjs/agent-skillsWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00048 | $0.01090 |
| Opus 5 | $0.00024 | $0.00545 |
| Sonnet 5 | $0.00010 | $0.00218 |
| Haiku 4.5 | $0.00005 | $0.00109 |
Grade A, and why
rstack-skill-evaluator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 70 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Rstack Skill Evaluator
A repo-specific compatibility layer on top of skill-creator. Reuse its Test / Improve / Benchmark concepts, JSON schemas, grading guidance, and eval viewer, but select the executor for the current environment.
Select the executor
- When the user requests Codex or
codexis the available CLI, read references/codex-cli.md and follow it. Its instructions override Claude-specific commands and subagent mechanics inskill-creator. - Otherwise, follow
skill-creatordirectly.
Do not invoke skill-creator/scripts/run_eval.py, run_loop.py, or improve_description.py in Codex mode. Those scripts shell out to claude -p and test Claude-specific skill discovery. Provider-neutral utilities such as quick_validate.py, aggregate_benchmark.py, and eval-viewer/generate_review.py can be reused after validating the installed dependency version.
Targeting a skill
If the user hasn't named a target, ask. Skills live under skills/ (production) and .agents/skills/ (internal-only).
Before editing an existing skill, snapshot it so the next iteration can compare the candidate against the previous version. Keep the snapshot and all raw run data outside tracked artifact paths.
Minimum eval rules
Use these rules for a basic eval unless the user requests a larger benchmark:
- Define at least two realistic cases: one representative workflow and one boundary, failure, or constraint case. Prefer a third case when the skill has multiple distinct modes.
- Give each case 2-5 outcome-focused assertions that can be verified from files, command results, or other durable evidence. Do not reward an agent merely for saying it succeeded.
- Run every case as a matched pair on fresh, identical fixture copies:
with_skillandwithout_skill. When improving an existing skill, also compare the candidate against the snapshotted previous version when that is the more useful baseline. - Keep the task prompt and runtime controls identical across configurations. The only intended difference is access to the target skill. Do not expose assertions, expected grader decisions, or another run's outputs to the executor.
- Use a fresh session for every run. Pin and record the CLI version, model, sandbox, approval, network, and relevant config. Never reuse a mutated working copy.
- Grade both configurations with the same checks. Prefer deterministic scripts for objective assertions; use an independent grader only for semantic checks, and require concrete evidence for every pass.
- Treat CLI crashes, timeouts, missing fixtures, and auth failures as harness failures, not skill failures. Fix or clearly report the harness problem before drawing skill conclusions.
- One run per configuration is a smoke eval. Use at least three repetitions before making claims about reliability, variance, token cost, or wall-time improvements.
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 70 lines · 48 tokens per session scan A 2bf12eeeae0c
rstack-skill-evaluator is a skill published in the GitHub repository rstackjs/agent-skills (90 stars, last pushed 6d ago), licensed MIT. It adds 48 tokens to every session and 1,090 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
mcporter
List, auth, and call MCP servers/tools from the terminal.
agent-messaging
Send and receive cryptographically signed messages between AI agents using the Agent Messaging Protocol (AMP). Use when the user asks to "send a message to an agent", "check agent inbox", "message another agent", "reply to a message", "notify an agent", or any inter-agent communication task.
📝 任务完成后归档
重要提醒: 每次完成复杂调试或开发任务后,主动执行此流程! 将学到的经验归档为 skill,供以后参考。不要等用户提醒。.
oracle
Best practices for using the oracle CLI (prompt + file bundling, engines, sessions, and file attachment patterns).
agent-mode
Unified tool for managing agent LLM modes (add, remove, update, list, switch).
agento11y-prod-setup
Sets up production evaluation and guardrails for a DEPLOYED AI agent in Grafana Agent Observability, grounded in the agent's own code and its real ingested traffic. The judgment layer on top of the agento11y skill: it reads the agent's source (system prompt, tools, entrypoint) AND samples its live traffic via gcx…