Getting it into your agent
There is no command for this one: it runs only inside a plugin, and the catalogue could not identify which plugin ships it. The source is linked below.
Wrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/qaml-ai/camelai/eval-reports)<a href="https://agentmods.dev/skills/qaml-ai/camelai/eval-reports"><img src="https://agentmods.dev/badge/skills/qaml-ai/camelai/eval-reports.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00000 | $0.01616 |
| Opus 5 | $0.00000 | $0.00808 |
| Sonnet 5 | $0.00000 | $0.00323 |
| Haiku 4.5 | $0.00000 | $0.00162 |
Grade A, and why
eval-reports scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
`curl -H "CF-Access-Client-Id: ..." -H "CF-Access-Client-Secret: ..."` both use these. How it starts
The opening of the file, as written. The whole thing — 114 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Running agent evals
camelAI agent evals run locally in a qaml-ai/camelAI checkout — there is no remote
runner. This service (https://evals.camelai.dev, workers/eval-reports) is only the shared
results history: a read-only dashboard + JSON API plus an upload endpoint the local reporter uses.
Everything is behind Cloudflare Access.
This document is served by the service itself at
GET /skill, so it always matches the running API. (You're reading it because you fetched that endpoint.)
Run an eval (locally, in chiridion-app)
Requirements: Docker, Bun, and a .dev.vars in the repo root (the eval secret bundle).
bun run test:eval <eval-id> # ids from workers/main/tests/evals/manifest.json
bun run test:eval:dashboard # or :deploy / :sandbox shortcuts
# Custom prompt (not committed to the tree) via the generic harness:
CUSTOM_EVAL_PROMPT="Build a dashboard from fake data." \
bun scripts/run-agent-eval.mjs custom-prompt-live
Common knobs: --model <id>, --timeout-ms <ms>, EVAL_REAL_DEPLOY=0/1 (publish apps to the
testing-grounds namespace for real), CUSTOM_EVAL_PROJECT,
CUSTOM_EVAL_REQUIRED_TRANSCRIPT_SUBSTRINGS. See bun scripts/run-agent-eval.mjs --help.
Suite and matrix invocations automatically mint a shared EVAL_BATCH_ID plus a default
EVAL_BATCH_LABEL, so their member runs group together on the dashboard. Pre-set either env var
when an orchestrator needs to join runs into an existing batch. Solo bun run test:eval <id> runs
stay batchless and render as singleton batches.
Report a run here
Set EVAL_REPORT=1 on the run and the reporter (scripts/report-eval-run.mjs) uploads the output
log and run metadata when the eval finishes — pass or fail. It also uploads the transcript
artifact, including its scorecard, when the eval emitted it:
EVAL_REPORT=1 bun run test:eval dashboard-fake-data-live
Reporting is best-effort and never fails the eval. Artifactless harness failures are still
reported; ingest synthesizes an evaluation_contract failure so they remain visible. Finalization
is retried up to three times with the same run id before the reporter gives up. Re-report an
artifact by hand:
What ships with it
42 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- app/app.css 1.2 KB
- app/components/app-shell.tsx 2.6 KB
- app/components/back-button.tsx 519 B
- app/components/batch-matrix.tsx 3.1 KB
- app/components/batch-result-badge.tsx 1.1 KB
- app/components/batches-table.tsx 10 KB
- app/components/criteria-card.tsx 2.8 KB
- app/components/eval-name-cell.tsx 3.0 KB
- app/components/evals-rollup.tsx 3.6 KB
- app/components/prompt-section.tsx 2.6 KB
- app/components/runs-table.tsx 6.7 KB
- app/components/score.tsx 2.0 KB
- app/components/scorecard-card.tsx 2.8 KB
- app/components/text-viewer.tsx 1.8 KB
- app/components/transcript/markdown.tsx 447 B
- app/components/transcript/tool-call-block.tsx 2.7 KB
- app/components/transcript/transcript-view.tsx 6.2 KB
- app/components/verdict-badge.tsx 733 B
- app/lib/api.ts 1.6 KB runs code
- app/lib/batches.ts 2.8 KB runs code
- app/lib/format.ts 4.1 KB runs code
- app/lib/transcript.ts 4.6 KB runs code
- app/main.tsx 1.5 KB
- app/routes/batch-detail.tsx 8.1 KB
- app/routes/run-detail.tsx 11 KB
- app/routes/runs-list.tsx 11 KB
- index.html 1004 B
- package.json 435 B
- public/camelAI-fullname-logo-darkmode.svg 8.5 KB
- public/camelAI-fullname-logo-lightmode.svg 8.5 KB
- public/favicon.svg 3.7 KB
- README.md 5.0 KB
- src/access.ts 4.9 KB runs code
- src/batches.ts 3.5 KB runs code
- src/global.d.ts 446 B runs code
- src/index.ts 27 KB runs code
- src/ingest.ts 10 KB runs code
- src/types.ts 4.5 KB runs code
- tsconfig.app.json 370 B
- tsconfig.json 336 B
- vite.config.ts 499 B runs code
- wrangler.jsonc 1.0 KB
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago First seen · 114 lines · 0 tokens per session scan A 48bcce7852c4
eval-reports is a skill published in the GitHub repository qaml-ai/camelAI (365 stars, last pushed yesterday), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 1,616 tokens. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
local-ai-agents
Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. Covers Small Language Models (SLMs), the OpenAI-compatible local endpoint, sandboxed local tools, local RAG with Chroma, local MCP servers, hybrid cloud/local routing, and the…
chronicle
Analyze Copilot session history for standup reports, usage tips, session search, and session reindexing. Use when the user asks for a standup, daily summary, usage tips, workflow recommendations, wants to search or find past sessions by keyword/file/PR, wants to reindex their session store, or asks about deleting…
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…