Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/mthines/agent-skills/confidencenpx skills add mthines/agent-skills --skill confidencegit clone --depth 1 https://github.com/mthines/agent-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/mthines/agent-skills/confidence)<a href="https://agentmods.dev/skills/mthines/agent-skills/confidence"><img src="https://agentmods.dev/badge/skills/mthines/agent-skills/confidence.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00103 | $0.03871 |
| Opus 5 | $0.00051 | $0.01936 |
| Sonnet 5 | $0.00021 | $0.00774 |
| Haiku 4.5 | $0.00010 | $0.00387 |
Grade A, and why
confidence scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 238 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Confidence Assessment
Rate your confidence that the current work fully solves the stated requirement.
Multi-signal evaluation. A single LLM-confidence number is unreliable as a stand-alone gate (token probability ≠ correctness). This skill combines the LLM's dimensional scoring with deterministic rule checks the agent must run alongside. The final score is gated on BOTH passing.
Contents
- Mode Detection
- Assessment Dimensions
- For
planmode — multi-signal: LLM scoring + rule checks (89% cap on failure) - For
codemode - For
analysismode
- For
- Output Format
- Score Thresholds
- Iteration Protocol (plan mode)
- Auto-Fix (Fix Mode Only)
Mode Detection
Check the arguments: $ARGUMENTS
| Argument | Default | Validates | When to use |
|---|---|---|---|
plan |
Implementation plan completeness | After Phase 1 planning, before autonomous execution | |
code |
yes | Code implementation correctness | After writing code, before PR |
analysis |
Analysis accuracy (root cause, refactor rationale, or skill gap) | During investigation, before proposing a fix, refactor, or skill-source diff | |
bug-analysis |
Deprecated alias for analysis — behaves identically |
Backwards-compatible; emit a one-line deprecation note in the report header |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 238 lines · 103 tokens per session scan A 1b3afa022211
confidence is a skill published in the GitHub repository mthines/agent-skills (12 stars, last pushed today), licensed MIT. It adds 103 tokens to every session and 3,871 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
ego-browser
Skill "ego-browser" from citrolabs/ego-lite, covering ego-browser, quick start, common helpers, task spaces and control handoff.
temporal-developer
Develop, debug, and manage Temporal applications across Python, TypeScript, Go, Java, .NET, Ruby, and Rust. Use when the user is building workflows, activities, or workers with a Temporal SDK, debugging issues like non-determinism errors, stuck workflows, or activity retries, using Temporal CLI, Temporal Server, or…
x-algorithm
Write X (Twitter) posts that the For You algorithm actually rewards. Grounded in the open-sourced X recommendation system — the Grok-based transformer ranker, Phoenix retrieval, Thunder in-network store, and Grox content-understanding pipeline. Use when the user wants to write a post, thread, reply, or quote; plan a…
opencode-memory
Browse local OpenCode history: sessions, messages, plans, prompt history, and prior decisions. Use when the user says history, previous session, last time, remember, recall, plans, prior work, or when resuming/debugging repeated work where earlier OpenCode context may help. Do not use for fresh tasks or when current…
creating-explainers
Use when creating an interactive explainer - a single self-contained HTML page with hand-built Canvas figures. Handles source-file explainers, topic-driven research explainers, and mixed intake where files provide the spine and research adds support. Trigger phrases include "make an explainer", "turn this paper into…
custom-icons
Create or refine custom icon assets as native or traced SVGs and transparent PNG/WebP files. Use when the user asks for a bespoke icon or cohesive icon set, wants an image traced into a clean vector, or needs a detailed or 3D icon with transparency.