hidai25

11 mods across 1 repository, 132 stars between them.

eval-view

01

hidai25/eval-view

Plugin Claude Code

Open-source testing and regression detection framework for AI agents. Golden baseline diffing, CI/CD integration, works with LangGraph, CrewAI, OpenAI, Anthropic Claude, HuggingFace, Ollama, and MCP.

132 8d ago A tokens not measured original Apache-2.0

evalview

02

hidai25/eval-view

MCP server Claude CodeCodexCursor +2

MCP server "evalview" as configured in hidai25/eval-view. Launched with evalview mcp serve. Needs 2 environment variables to run.

132 8d ago A tokens not measured original Apache-2.0

eval-view AGENTS.md

03

hidai25/eval-view

Instructions file CodexOpenCode

Instructions for hidai25/eval-view, covering evalview agent instructions, what evalview is, core concepts, testcase and evaluationresult.

132 8d ago A 2,634 tokens original Apache-2.0

hidai25/eval-view

Skill Claude CodeCodex

Beat procrastination with task breakdown, 2-minute starts, and accountability tracking.

132 8d ago A 21 tokens original Apache-2.0

code-reviewer

05

hidai25/eval-view

Skill Claude CodeCodex

Performs comprehensive code reviews with security, quality, and best practice checks.

132 8d ago A 18 tokens original Apache-2.0

hello-world

06

hidai25/eval-view

Skill Claude CodeCodex

A simple skill that creates a greeting file.

132 8d ago A 11 tokens original Apache-2.0

code-reviewer

07

hidai25/eval-view

Skill Claude CodeCodex

A skill that helps review code for best practices, bugs, and security issues.

132 8d ago A 19 tokens original Apache-2.0

generate-tests

08

hidai25/eval-view

Skill Claude CodeCodex

Generate EvalView test cases — either from a SKILL.md file using LLM-powered generation, or by capturing real agent interactions through a proxy.

132 8d ago A 32 tokens original Apache-2.0

run-eval

09

hidai25/eval-view

Skill Claude CodeCodex

Run EvalView regression checks against golden baselines to detect regressions in AI agent behavior after code, prompt, or model changes.

132 8d ago A 30 tokens original Apache-2.0

watch

10

hidai25/eval-view

Skill Claude CodeCodex

Start EvalView watch mode to automatically re-run regression checks whenever project files change.

132 8d ago A 18 tokens original Apache-2.0

evalview-mcp

11

hidai25/eval-view

MCP server Claude CodeCodexCursor +2

Open-source testing and regression detection framework for AI agents. Golden baseline diffing, CI/CD integration, works with LangGraph, CrewAI, OpenAI, Anthropic Claude, HuggingFace, Ollama, and MCP. Runs locally from the evalview Python package. Needs 1 environment variable to run.

132 8d ago A tokens not measured original Apache-2.0