iblai-api-agent-eval

A tool for testing and measuring the quality of an ibl.ai agent. It uses evaluation datasets—collections of questions and expected answers—to run experiments and score the agent's responses.

In plain words
What is it for?
Use it to create evaluation cases, run them against an agent, score responses with automated or human review, configure scoring, and export results as CSV.
Why use it?
It gives you repeatable tests instead of judging an agent from a few conversations. Scores can come from an AI judge or human reviewers, and results can be exported for inspection.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/iblai/api/iblai-api-agent-eval
Any agent
npx skills add iblai/api --skill iblai-api-agent-eval
Clone the repo
git clone --depth 1 https://github.com/iblai/api

Made for: Claude Code, Codex.

Per session 67 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,753 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00067 $0.02753
Opus 5 $0.00034 $0.01376
Sonnet 5 $0.00013 $0.00551
Haiku 4.5 $0.00007 $0.00275

Measured 3d ago against content hash 895ac93e73ac, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

iblai-api-agent-eval scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

curl -X POST \
skills/iblai-api-agent-eval/SKILL.md · 185 lines

How it starts

The opening of the file, as written. The whole thing — 185 lines — stays where its author put it; the contents beside it link to each section on GitHub.

iblai-api-agent-eval

Measure and improve an agent's quality from the API: build evaluation datasets, run experiments that send each question to the agent, then grade the results with LLM-as-Judge and/or human scores and export to CSV. Use to test an agent against a dataset and grade the results.

Auth & conventions

  • Base URL: https://api.iblai.app
  • Header: Authorization: Api-Token $IBLAI_API_KEY on every request. (The dev docs phrase this as Authorization: Token <key> — it is the same platform key; use Api-Token.)
  • Path vars: {org} = $IBLAI_ORG, {username} = $IBLAI_USERNAME.
  • Host root: …/dm/api/ai-mentor/orgs/{org}/users/{username}/evaluations/. Below, …/evals = that root. (ai-mentor is the canonical mount; the ai-agent spelling is an accept-only alias for the same routes.)
  • Not connected yet? Run /iblai-api-login first to populate IBLAI_ORG, IBLAI_USERNAME, and IBLAI_API_KEY.

Concepts

  • These eval datasets are not the agent's RAG datasets. evaluations/datasets/ hold graded test cases (input + expected output) for measuring agent quality. They are unrelated to an agent's knowledge/training datasets in /iblai-api-agent-dataset (RAG documents) — do not cross-wire the two.
  • Eval data is org-scoped and isolated. Datasets, items, runs, scores, and score configs belong to $IBLAI_ORG alone — no other org can read them or grade against them.
  • Runs and judges are async task records. Starting a run (POST …/runs/) or an LLM-as-Judge (POST …/evaluate/) dispatches a background task and returns 202 immediately with a task record that moves pending → in_progress → completed (or failed). Poll for status; a run must reach completed before you judge or export it.
  • Three grading paths. LLM-as-Judge (…/evaluate/) scores every item in a run automatically from a free-text criteria rubric. Scores (…/scores/) are individual human/numeric/boolean/categorical annotations on a trace. Score configs (…/score-configs/) are reusable rubrics a score references via config_id.
  • Pagination. List endpoints take ?page= (1-indexed) and ?limit= (default 50, max 200) and return {count, next, previous, results}.
  • On the wire the path segment is users/{user_id}; its value is $IBLAI_USERNAME.

Read the full file on GitHub · 185 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 185 lines · 67 tokens per session scan A 895ac93e73ac

Subscribe to this mod's changes

iblai-api-agent-eval is a skill published in the GitHub repository iblai/api (15 stars, last pushed 4d ago), licensed MIT. It adds 67 tokens to every session and 2,753 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

vindicate

Use when the user wants to write, add, fix, stabilize (flaky), refactor, run, or audit Playwright browser tests, draft requirements/stories from a recording (no tests), find test-coverage gaps, scaffold a Playwright project, or set up Playwright CI. Vindicate's guided workflow for grounded, conformant Playwright test…

OpenEvident/vindicate · 77 tokens

audit

Audits recent work against its Definition of Done and project patterns. Runs the test suite, compares code against the spec, and reports PASS / PARTIAL / FAIL. Also runs the Critical Gate — a safety scan of the diff for destructive or dangerous operations. Generates an incremental prompt pack for any gaps found. With…

pe-menezes/vibeflow · 93 tokens

implement

Implements a feature from its spec following all guardrails: budget, DoD, anti-scope, and pattern compliance. Runs an 8-phase pipeline (find spec → extract guardrails → load patterns → plan → implement → test → refine → self-verify DoD). Use after gen-spec when you're ready to code. The agent has filesystem access and…

pe-menezes/vibeflow · 79 tokens

assistant

Assistant — on any repo, scan README→docs→AGENTS→CONTRIBUTING→PR templates→task runners→devcontainer→CI→configs before code; cite sources; prefer AGENTS.md for agent behavior; portable across Cursor/Copilot/Claude; use agent-toolkit CLI when needed.

ulises-jeremias/agent-toolkit · 64 tokens

design-improvement

WHAT - Browser-grounded iterative design improvement. Consumes design-assessment findings, defines direction, prioritizes safe vs ambiguous changes, implements within existing design system, runs app, captures rendered evidence via browser, reviews and iterates. Reuses evidence model — no new scoring framework.

ulises-jeremias/agent-toolkit · 60 tokens

codeql

CodeQL operational workflow — discover/config, run/inspect, triage SARIF findings (rule/query ID, source→sink, evidence), remediate, re-validate. Distinguishes broad MegaLinter linting from semantic security analysis.

ulises-jeremias/agent-toolkit · 52 tokens