Agent Claude Code
Use before shipping any LLM feature that touches users — reviews for prompt injection, hallucination risk, output misuse, agentic scope creep, and abuse vectors.
Your AI assistant will skip the eval. This pack won't let it. Drop-in enforcement skills + agents for LLM product teams.
Agent Claude Code
Use before shipping any LLM feature that touches users — reviews for prompt injection, hallucination risk, output misuse, agentic scope creep, and abuse vectors.
Agent Claude Code
Use when designing an evaluation suite for a new LLM feature or prompt — selects metrics, builds test sets, and writes eval harness code.
Agent Claude Code
Use when choosing a model for a new feature or evaluating whether to switch models — structured benchmarking and cost-quality analysis.
Agent Claude Code
Use when writing, iterating, or debugging prompts. Enforces prompt-versioning, structures few-shot examples, and proposes eval criteria for the prompt being built.
Agent Claude Code
Use when designing, debugging, or upgrading a retrieval-augmented generation pipeline — chunking strategy, embedding choice, retrieval, reranking, and generation.