Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add jeffreytse/grimoire-core --skill run-performance-calibrationgit clone --depth 1 https://github.com/jeffreytse/grimoire-coreWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/jeffreytse/grimoire-core/run-performance-calibration)<a href="https://agentmods.dev/skills/jeffreytse/grimoire-core/run-performance-calibration"><img src="https://agentmods.dev/badge/skills/jeffreytse/grimoire-core/run-performance-calibration/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/jeffreytse/grimoire-core/run-performance-calibration"><img src="https://agentmods.dev/badge/skills/jeffreytse/grimoire-core/run-performance-calibration.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00040 | $0.01916 |
| Opus 5 | $0.00020 | $0.00958 |
| Sonnet 5 | $0.00008 | $0.00383 |
| Haiku 4.5 | $0.00004 | $0.00192 |
Grade A, and why
run-performance-calibration scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 150 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Run Performance Calibration
Facilitate a cross-manager calibration session where each manager's ratings are reviewed against a shared standard — to correct for individual rater bias, surface promotion candidates, and produce fair performance outcomes across the team.
Why This Is Best Practice
Adopted by: Google, Amazon, Microsoft, Meta, and Salesforce all run formal calibration sessions as part of their annual performance cycle; Google's re:Work program documents calibration as a required step before performance ratings are finalized; SHRM's performance management best practice guide identifies calibration as the primary mechanism for reducing rater bias at scale; the US federal government mandates performance calibration under OPM guidelines for Senior Executive Service ratings Impact: A 2019 Deloitte human capital research study found that organizations using structured calibration sessions reduced rating distribution skew by 35% compared to organizations relying on individual manager ratings alone; Google's internal research (published in re:Work) found that calibrated ratings produced 28% less variance in performance scores for comparable performers across different managers; uncalibrated ratings systematically disadvantage employees whose managers are lenient raters, those working in less visible roles, and those on teams with harsher managers Why best: Without calibration, performance ratings measure both the employee's performance and the manager's rating tendencies — a harsh rater's strong performers receive the same ratings as a lenient rater's average performers; this inequity flows directly into compensation and promotion decisions, creating retention risk among the best performers who are rated fairly only in relative terms, and advancement inequity among employees in harsher rating environments
Sources: Google re:Work "Manager Training: Performance Calibration"; Deloitte "Performance Management: Playing a Winning Hand" (2019); SHRM "Calibration Meetings: Best Practices" (shrm.org); Buckingham & Goodall "Reinventing Performance Management" (Harvard Business Review, 2015)
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 150 lines · 40 tokens per session scan A c41a0d06b127
run-performance-calibration is a skill published in the GitHub repository jeffreytse/grimoire-core (4 stars, last pushed 22d ago), licensed MIT. It adds 40 tokens to every session and 1,916 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
org-report
Generate an executive organization brief (leadership chain, direct reports, extended teams, and vendor partners) for a target person. Produces a Word + PDF report from a directory MCP source (WorkIQ is the reference implementation) and displays it in a Copilot canvas. Triggers on "generate org report for X", "build…
coaching-techniques
GROW model, active listening, developmental feedback, and team growth approaches.
meeting-efficiency
Agenda design, time boxing, decision capture, async alternatives, and productive facilitation.
stakeholder-management
Influence mapping, communication strategies, and expectation management for complex organizations.
scaffolding-generator
Pattern-aware code scaffolding that detects existing conventions and generates new components matching the codebase style. Use when asked to "scaffold a component", "generate boilerplate", "create a new module", "bootstrap a service", "add a new endpoint", "create a new controller", or "add a new feature module".…
status-reporting
Create stakeholder-friendly project status updates and progress reports.