iFixAi tools for Claude Code. Claude guides you through an independent, open-source audit of your own agent: is it doing the job it's supposed to do, given your business rules and org structure? Test any model (Anthropic, OpenAI, Gemini, Azure, Bedrock, …) graded by the judge(s) of your choice.
12k 3d agoA
tokens not measured
originalApache-2.0
Independent auditing for AI agents, open source. Claude investigates your agent, authors an iFixAi fixture from its setup, and audits it against the one question other tools skip: is it doing the job it's supposed to do, given your business rules and org structure? Adversarial depth, assurance discipline. Test any…
12k 3d agoA
tokens not measured
originalApache-2.0
Open-source testing and regression detection framework for AI agents. Golden baseline diffing, CI/CD integration, works with LangGraph, CrewAI, OpenAI, Anthropic Claude, HuggingFace, Ollama, and MCP.
132 9d agoA
tokens not measured
originalApache-2.0
Test whether your Claude Code skills actually trigger, and benchmark any coding agent — author, run, and analyze the eval suite that proves it, locally or as a CI gate.
119 3d agoA
tokens not measured
originalApache-2.0
Agent eval for any MCP workflow — score output quality, catch PII and prompt injection, verify citations, and enforce cost budgets. Adds the Iris MCP server (9 tools) plus an agent-eval skill that guides eval-driven development.