proofrag
01Plugin Claude Code
Plugin marketplace listing 1 plugin: proofrag.
Plugin Claude Code
Plugin marketplace listing 1 plugin: proofrag.
Plugin Claude Code
Evaluate a RAG/LLM app: generate a golden set from your docs, run LLM-as-judge + retrieval metrics, and produce a shareable HTML scorecard with a CI gate.
Instructions file CodexOpenCode
Instructions for unshDee/proofrag, covering agents, use it as a skill and install the engine.
Command
Evaluate a RAG/LLM app — generate a golden set, judge it, and produce a scorecard.
Skill Claude CodeCodex
Evaluate a RAG or LLM app. Use when the user wants to test, score, benchmark, or catch regressions in a retrieval/RAG/LLM system, generate an evaluation/golden dataset from their docs, measure hallucination/groundedness/correctness, or gate CI on answer quality. Generates a golden set from the user's own corpus, runs…