Agent Claude Code
Use when designing an evaluation suite for a new LLM feature or prompt — selects metrics, builds test sets, and writes eval harness code.
2 1mo ago A 32 tokens
original MIT
5 tagged eval, measured the same way as everything else here.
Browse within: AI Safety 5prompt-engineering 5
Agent Claude Code
Use when designing an evaluation suite for a new LLM feature or prompt — selects metrics, builds test sets, and writes eval harness code.
Agent Claude Code
Use when choosing a model for a new feature or evaluating whether to switch models — structured benchmarking and cost-quality analysis.
Agent Claude Code
Use when designing, debugging, or upgrading a retrieval-augmented generation pipeline — chunking strategy, embedding choice, retrieval, reranking, and generation.