langwatch
01Plugin Claude Code
Plugin marketplace listing 1 plugin: langwatch.
31 tagged evaluation, measured the same way as everything else here.
Browse within: browser-automation 11
Plugin Claude Code
Plugin marketplace listing 1 plugin: langwatch.
Plugin Claude Code
Records which repository and branch each coding-agent session worked in, and teaches the agent to read its own traces back from LangWatch.
Plugin Claude Code
Open-source testing and regression detection framework for AI agents. Golden baseline diffing, CI/CD integration, works with LangGraph, CrewAI, OpenAI, Anthropic Claude, HuggingFace, Ollama, and MCP.
Plugin Claude Code
Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs.
Plugin Claude Code
Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs.
ContextJet-ai/awesome-llm-observability
Plugin Claude Code
Plugin marketplace listing 1 plugin: llm-observability.
ContextJet-ai/awesome-llm-observability
Plugin Claude Code
26 Agent Skills (several with runnable, unit-tested scripts) for building, evaluating, securing, and monitoring reliable LLM & AI-agent apps.
redhat-community-ai-tools/harness-eval
Plugin Claude Code
Evaluate AI agent setups for best practices, redundancy, security, and cross-component issues.
redhat-community-ai-tools/harness-eval
Plugin Claude Code
Evaluate AI agent setups for best practices, redundancy, security, and cross-component issues.
Plugin Claude Code
Plugin marketplace listing 19 plugins: full-eval, hackathon-jury, vc-evaluator, oss-readiness, architect.
Plugin Claude Code
App Store and Google Play screenshot creation with exact platform specs, gallery ordering, and preview videos.
Plugin Claude Code
Multidisciplinary idea and project evaluator — dispatches parallel specialist agents, synthesizes unified assessment with scores.
Plugin Claude Code
Plugin marketplace listing 1 plugin: proofloop.
Plugin Claude Code
The first universal judge for Claude Code & Cowork. Auto-evaluates skill and agent execution quality with 7-dimension scoring, configurable rubrics, and dual-mode operation (auto hooks + manual /judge command).
golemfoundation/octant-council-builder
Plugin Claude Code
Council builder — generates multi-agent evaluation councils, not a council itself.
Plugin Claude Code
Plugin marketplace listing 1 plugin: proofrag.
Plugin Claude Code
Evaluate a RAG/LLM app: generate a golden set from your docs, run LLM-as-judge + retrieval metrics, and produce a shareable HTML scorecard with a CI gate.
Plugin Claude Code
AI coding skill evaluator — 11-dimension session scoring, decision tracking, PR reports.
Plugin Claude Code
Author verifiable eval records through a draft → review → revise → submit loop with server-enforced graders; compile to JSONL/CSV/Inspect/lm-eval via MCP. STDIO or Streamable HTTP.
leonardoeverling-spec/governed-context
Plugin Claude Code
Plugin marketplace listing 1 plugin: governed-context.
leonardoeverling-spec/governed-context
Plugin Claude Code
Governed memory (explicit states, fail-closed sensitivity gates, quarantine-only writes) as a dependency-free MCP server, plus a blind-evaluation protocol skill for honest A/B comparisons.
Plugin Claude Code
Plugin marketplace listing 1 plugin: cold-run.
Plugin Claude Code
Cold-run your agent skills to find which ones change nothing, plus 24 skills that survived the same test.