Agent Claude Code
Generates comprehensive eval reports with quality, semantic (LLM-as-Judge), accuracy, regression, and performance analysis for think-mcp tools. Use after think-mcp-tester completes.
0 7mo ago A 44 tokens
Agent Claude Code
Generates comprehensive eval reports with quality, semantic (LLM-as-Judge), accuracy, regression, and performance analysis for think-mcp tools. Use after think-mcp-tester completes.
Agent Claude Code
Executes quality, regression, semantic, and performance evals for think-mcp tools. Use for ongoing validation and new mental model testing.