Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add jaktestowac/awesome-copilot-for-testers --skill tracking-quality-trendsgit clone --depth 1 https://github.com/jaktestowac/awesome-copilot-for-testersWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/jaktestowac/awesome-copilot-for-testers/tracking-quality-trends)<a href="https://agentmods.dev/skills/jaktestowac/awesome-copilot-for-testers/tracking-quality-trends"><img src="https://agentmods.dev/badge/skills/jaktestowac/awesome-copilot-for-testers/tracking-quality-trends/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/jaktestowac/awesome-copilot-for-testers/tracking-quality-trends"><img src="https://agentmods.dev/badge/skills/jaktestowac/awesome-copilot-for-testers/tracking-quality-trends.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00105 | $0.02171 |
| Opus 5 | $0.00053 | $0.01086 |
| Sonnet 5 | $0.00021 | $0.00434 |
| Haiku 4.5 | $0.00011 | $0.00217 |
Grade A, and why
tracking-quality-trends scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 148 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Tracking Quality Trends
Use this skill when quality gets reported as a number, and nobody can say whether it is better or worse than last time.
A snapshot is almost useless on its own. "Coverage is 71%" prompts an argument about whether 71 is good. "Coverage on changed lines has moved 62 → 71 over three releases, limit 75, goal 85" prompts a decision. Direction is the finding; the absolute value is context.
The second thing this catches is regression in the quality system - a threshold lowered, a job made non-blocking, a waiver renewed for the fourth time. Those never show up in a snapshot, because a snapshot reports what is measured, not what stopped being measured.
When to Use
- quality reporting is a series of unconnected numbers
- a metric is quoted with no baseline and no target
- a team needs to show a quarter of improvement, or explain a quarter without it
- a contract has been re-derived and the question is what moved
- gates are being quietly weakened and nobody has noticed
- a release decision needs direction, not just current state
Operating Principles
- Limit / current / goal, always three numbers. Limit is the threshold that triggers action; current is measured; goal is the target. A metric with only a current value cannot be acted on.
- Direction and magnitude, not just the delta. "−3pp" needs "inside the noise band" or "third consecutive fall" beside it to mean anything.
- Archive the run, do not recompute history. Store each reading with its date, commit, and how it was measured. Recomputed history changes when the method changes, and then the trend is fiction.
- A method change breaks the series. Say so, and start a new one rather than pretending the numbers are comparable.
- Track the quality system, not only the code. Thresholds, gate blocking-ness, waiver counts and ages, suppression counts. A repo whose coverage rose while its threshold fell has got worse.
- Fewer metrics, honestly measured. Six metrics a team acts on beat twenty nobody reads. Every metric needs a named owner and an action if it crosses its limit.
- Every metric carries its caveat. Coverage does not prove correctness; a rising pass rate can mean weaker tests. Report the caveat inline, not in a footnote.
- Never trend a metric that can be gamed without saying so. Coverage, test count, and defect count all move under pressure without quality moving.
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 148 lines · 105 tokens per session scan A 8fdda5422baa
tracking-quality-trends is a skill published in the GitHub repository jaktestowac/awesome-copilot-for-testers (113 stars, last pushed 15d ago), licensed MIT. It adds 105 tokens to every session and 2,171 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
fastapi_stripe
A Stripe Checkout implementation takes four steps.
standalone-python-scripts
Skill "standalone-python-scripts" from iloveitaly/llm-ide-rules, covering standalone python scripts, /// script, requires-python = ">=3.13", dependencies = [] and ///.
secrets
Here's how environment variables are managed in this application.
fix-tests
Focus on all unit + command tests (pytest --exclude tests/integration). Make sure they pass and fix errors. If you run into anything very odd: stop, and let me know. Mutate test code first and let me know if you think you should update application code.
plan-only
As this point, I only want to talk about the plan. How would you do this? What would you refactor to make this design clean? You are an expert software engineer and I want you to think hard about how to plan this project out.
dev-in-browser
Use your browser to view https://verso.localhost which is tied to livereload dev server which is already running. You can inspect that page (including taking screenshots!) to validate that your changes fixed the issue.