documenter
25Agent
:::blue.agents.documenter.
76 tagged llms, measured the same way as everything else here.
Browse within: claude-code-skills 20obsidian 20second-brain 20agenticai 14enterprise-solutions 14large-language-models 7human-in-the-loop 6humanlayer 6api 5evaluation-framework 5knowledge 5knowledge-base 5knowledge-management 5llmops 5
Agent
:::blue.agents.documenter.
Agent
:::blue.agents.registry.AgentRegistry.
Agent Claude Code
Part of soundcheck
Audits a codebase for missing security controls — the gaps that pattern-matching auditors won't catch, like no timeout, no cost cap, no rate limit, prose-only guards. Invoked in parallel with vulnerability-audit calls.
Agent Claude Code
Part of soundcheck
Second-pass refutation filter for security-review findings. Reads each candidate finding's cited code and drops the ones with concrete refutation evidence (a guard, middleware, sanitizer, or correct API call at the cited location). Bias is toward keeping; uncertain findings pass through.
Agent Claude Code
Part of soundcheck
Finds security-sensitive code locations in a repository — the files and functions a reviewer should look at. Reads the threat model for context, then enumerates and ranks hotspots. Invoked after threat-modeling, before per-hotspot review.
Agent
Analyze AgentV evaluation results to identify weak assertions, suggest deterministic upgrades for LLM-grader graders, flag cost/quality improvements, and surface cross-run benchmark patterns. Use when reviewing eval quality, improving evaluation configs, or triaging flaky/expensive evaluations.
Agent
Perform bias-free blind comparison of evaluation outputs from multiple providers or configurations. Randomizes labeling, generates task-specific rubrics, scores N-way comparisons, then unblinds results and attributes improvements. Dispatch this agent when comparing outputs across targets or iterations.
Agent
Grade a candidate response for an AgentV evaluation test case. Evaluates all assertion types natively — deterministic checks via string operations, LLM grading via Claude's own reasoning, script-grader via Bash script execution. Zero CLI dependency. Dispatch this agent after a candidate completes a test case.
Agent Claude Code
Use this agent when you need to create a new GitHub release for a project, including analyzing commits, determining version numbers, writing release notes, and publishing the release. This agent should be triggered after significant development work is completed and you're ready to package changes into a formal…
Agent Claude Code
Use this agent when you need to run, create, or review tests for the project, particularly after new features have been introduced or changes have been made to the /src folder. This agent should be invoked to ensure code quality through automated testing, set up new test suites, or review existing test coverage and…
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: