evaluate-pr-tests

evaluate-pr-tests is a skill for Claude Code, Codex from dotnet/maui. It costs 94 tokens per session (2,890 once invoked), scanned A, original, MIT.

A review of tests added in a pull request, checking whether they cover the code change and handle important edge cases. It also checks whether the chosen test type is appropriate, such as a unit test versus a device or user-interface test.

In plain words
What is it for?
Use it before merging to assess tests for a bug fix, check their quality and coverage, and identify opportunities to replace slower device or UI tests with unit tests.
Why use it?
It finds missing coverage, weak test cases, convention problems, and situations where a lighter test would be enough. If no tests were added, it reports that the fix lacks test coverage.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/dotnet/maui/evaluate-pr-tests
Any agent
npx skills add dotnet/maui --skill evaluate-pr-tests
Clone the repo
git clone --depth 1 https://github.com/dotnet/maui

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for evaluate-pr-tests

README.md
[![agentmods](https://agentmods.dev/badge/skills/dotnet/maui/evaluate-pr-tests.svg)](https://agentmods.dev/skills/dotnet/maui/evaluate-pr-tests)
Your own site
<a href="https://agentmods.dev/skills/dotnet/maui/evaluate-pr-tests"><img src="https://agentmods.dev/badge/skills/dotnet/maui/evaluate-pr-tests.svg" alt="Measured on agentmods" height="20"></a>
Per session 94 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,890 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00094 $0.02890
Opus 5 $0.00047 $0.01445
Sonnet 5 $0.00019 $0.00578
Haiku 4.5 $0.00009 $0.00289

Measured 3d ago against content hash a30612c77a79, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

evaluate-pr-tests scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

The scan reads SKILL.md. This mod also ships 1 executable file (scripts/Gather-TestContext.ps1), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.github/skills/evaluate-pr-tests/SKILL.md · 330 lines

How it starts

The opening of the file, as written. The whole thing — 330 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Evaluate PR Tests

Evaluates the quality, coverage, and appropriateness of tests added in a PR. Produces a structured report with actionable findings.

When to Use

  • ✅ PR has tests and you want to evaluate their quality
  • ⚠️ PR has no test files -- output a ❌ Fix Coverage verdict noting no tests were added; skip remaining criteria
  • ✅ Reviewing whether tests adequately cover the fix
  • ✅ Checking if a lighter test type could be used instead
  • ✅ Before merging a PR, as part of review

Quick Start

# Auto-detect PR and base branch
pwsh .github/skills/evaluate-pr-tests/scripts/Gather-TestContext.ps1

# With explicit base branch
pwsh .github/skills/evaluate-pr-tests/scripts/Gather-TestContext.ps1 -BaseBranch "origin/main"

Workflow

Step 1: Gather Automated Context

Run the script to get file categorization, convention checks, and anti-pattern detection:

pwsh .github/skills/evaluate-pr-tests/scripts/Gather-TestContext.ps1

This produces a report at CustomAgentLogsTmp/TestEvaluation/context.md with:

  • File categorization (fix files vs test files by type)
  • Convention compliance checks (naming, attributes, anti-patterns)
  • AutomationId consistency (HostApp ↔ test)
  • Existing similar tests
  • Platform scope analysis

Step 2: Understand the Fix

Read the fix files to understand:

  • What changed — which code paths were modified
  • Why it changed — the bug being fixed (from PR description or linked issue)
  • Edge cases — what boundary conditions exist in the changed code

Step 3: Evaluate the Tests

Read each test file and evaluate against all criteria below. For each criterion, provide a verdict (✅ Pass, ⚠️ Concern, ❌ Fail) with explanation.

Step 4: Produce the Report

Output a structured evaluation report (see Output Format below).


Evaluation Criteria

1. Fix Coverage

Question: Does the test exercise the actual code paths changed by the fix?

How to check:

  • Trace the test's actions through the code to the fix location
  • Would the test fail if the fix were reverted?
  • Does the test assert on the specific behavior that was broken?

Read the full file on GitHub · 330 lines

Files

What ships with it

2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 330 lines · 94 tokens per session scan A a30612c77a79

Subscribe to this mod's changes

evaluate-pr-tests is a skill published in the GitHub repository dotnet/maui (23,317 stars, last pushed 3d ago), licensed MIT. It adds 94 tokens to every session and 2,890 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

dpg-migration

Migration logic for Azure SDK for .NET data-plane libraries migrating from AutoRest/Swagger to TypeSpec-based generation. Uses MCP tools from the generator-agent server for automated deterministic fixes.

Azure/azure-sdk-for-net · 41 tokens

csharp-azure-spector-coverage-gaps

Discovers and implements gaps in Spector test coverage for the Azure C# HTTP client emitter. Use when asked to find missing Spector scenarios, add Spector test coverage, or implement a specific Spector spec for the Azure C# emitter. Can also compare coverage between the Azure dashboard and the Standard (TypeSpec core)…

Azure/azure-sdk-for-net · 78 tokens

auto-build-repair

Headless, bounded repair of custom-code build failures in an already-generated Azure SDK PR. Thin wrapper over the shared azure-sdk-mcp:azsdkcustomizedcodeupdate engine in custom-code-only scope (editScope: CustomCode); the skill owns the iterate-until-green loop, capped by a per-language maxIterations read from…

Azure/azure-sdk-for-net · 149 tokens

azure-sdk-mgmt-pr-review

Review Azure SDK management-plane pull requests, check naming conventions, API compatibility, and code quality.

Azure/azure-sdk-for-net · 26 tokens

mgmt-review-comment-resolution

Resolve review comments on Azure management-plane .NET SDK PRs. Handles renaming types/properties, changing property types, and other API surface adjustments by updating TypeSpec client.tsp and regenerating.

Azure/azure-sdk-for-net · 46 tokens

mitigate-breaking-changes

Patterns and techniques for mitigating breaking changes during Azure management-plane SDK migration from Swagger/AutoRest to TypeSpec. Covers SDK-side customizations (partial classes, CodeGenType, CodeGenSuppress) and TypeSpec decorator customizations (clientName, access, markAsPageable, alternateType…

Azure/azure-sdk-for-net · 68 tokens