Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/iblai/api/iblai-api-agent-evalnpx skills add iblai/api --skill iblai-api-agent-evalgit clone --depth 1 https://github.com/iblai/apiWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00067 | $0.02753 |
| Opus 5 | $0.00034 | $0.01376 |
| Sonnet 5 | $0.00013 | $0.00551 |
| Haiku 4.5 | $0.00007 | $0.00275 |
Grade A, and why
iblai-api-agent-eval scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
curl -X POST \ How it starts
The opening of the file, as written. The whole thing — 185 lines — stays where its author put it; the contents beside it link to each section on GitHub.
iblai-api-agent-eval
Measure and improve an agent's quality from the API: build evaluation datasets, run experiments that send each question to the agent, then grade the results with LLM-as-Judge and/or human scores and export to CSV. Use to test an agent against a dataset and grade the results.
Auth & conventions
- Base URL:
https://api.iblai.app - Header:
Authorization: Api-Token $IBLAI_API_KEYon every request. (The dev docs phrase this asAuthorization: Token <key>— it is the same platform key; use Api-Token.) - Path vars:
{org}=$IBLAI_ORG,{username}=$IBLAI_USERNAME. - Host root:
…/dm/api/ai-mentor/orgs/{org}/users/{username}/evaluations/. Below,…/evals= that root. (ai-mentoris the canonical mount; theai-agentspelling is an accept-only alias for the same routes.) - Not connected yet? Run
/iblai-api-loginfirst to populateIBLAI_ORG,IBLAI_USERNAME, andIBLAI_API_KEY.
Concepts
- These eval datasets are not the agent's RAG datasets.
evaluations/datasets/hold graded test cases (input + expected output) for measuring agent quality. They are unrelated to an agent's knowledge/training datasets in/iblai-api-agent-dataset(RAG documents) — do not cross-wire the two. - Eval data is org-scoped and isolated. Datasets, items, runs, scores, and
score configs belong to
$IBLAI_ORGalone — no other org can read them or grade against them. - Runs and judges are async task records. Starting a run (
POST …/runs/) or an LLM-as-Judge (POST …/evaluate/) dispatches a background task and returns 202 immediately with a task record that movespending → in_progress → completed(orfailed). Poll for status; a run must reachcompletedbefore you judge or export it. - Three grading paths. LLM-as-Judge (
…/evaluate/) scores every item in a run automatically from a free-textcriteriarubric. Scores (…/scores/) are individual human/numeric/boolean/categorical annotations on a trace. Score configs (…/score-configs/) are reusable rubrics a score references viaconfig_id. - Pagination. List endpoints take
?page=(1-indexed) and?limit=(default 50, max 200) and return{count, next, previous, results}. - On the wire the path segment is
users/{user_id}; its value is$IBLAI_USERNAME.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 185 lines · 67 tokens per session scan A 895ac93e73ac
iblai-api-agent-eval is a skill published in the GitHub repository iblai/api (15 stars, last pushed 4d ago), licensed MIT. It adds 67 tokens to every session and 2,753 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
vindicate
Use when the user wants to write, add, fix, stabilize (flaky), refactor, run, or audit Playwright browser tests, draft requirements/stories from a recording (no tests), find test-coverage gaps, scaffold a Playwright project, or set up Playwright CI. Vindicate's guided workflow for grounded, conformant Playwright test…
audit
Audits recent work against its Definition of Done and project patterns. Runs the test suite, compares code against the spec, and reports PASS / PARTIAL / FAIL. Also runs the Critical Gate — a safety scan of the diff for destructive or dangerous operations. Generates an incremental prompt pack for any gaps found. With…
implement
Implements a feature from its spec following all guardrails: budget, DoD, anti-scope, and pattern compliance. Runs an 8-phase pipeline (find spec → extract guardrails → load patterns → plan → implement → test → refine → self-verify DoD). Use after gen-spec when you're ready to code. The agent has filesystem access and…
assistant
Assistant — on any repo, scan README→docs→AGENTS→CONTRIBUTING→PR templates→task runners→devcontainer→CI→configs before code; cite sources; prefer AGENTS.md for agent behavior; portable across Cursor/Copilot/Claude; use agent-toolkit CLI when needed.
design-improvement
WHAT - Browser-grounded iterative design improvement. Consumes design-assessment findings, defines direction, prioritizes safe vs ambiguous changes, implements within existing design system, runs app, captures rendered evidence via browser, reviews and iterates. Reuses evidence model — no new scoring framework.
codeql
CodeQL operational workflow — discover/config, run/inspect, triage SARIF findings (rule/query ID, source→sink, evidence), remediate, re-validate. Distinguishes broad MegaLinter linting from semantic security analysis.