Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/princeofscale/bloxforge/evalsnpx skills add princeofscale/bloxforge --skill evalsgit clone --depth 1 https://github.com/princeofscale/bloxforgeWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00022 | $0.00945 |
| Opus 5 | $0.00011 | $0.00473 |
| Sonnet 5 | $0.00004 | $0.00189 |
| Haiku 4.5 | $0.00002 | $0.00094 |
Grade A, and why
evals scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 82 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Evals
22 symbols | 4 files | Cohesion: 90%
When to Use
- Working with code in
evals/ - Understanding how runSuite, mean, bootstrapTax work
- Modifying evals-related functionality
Key Files
| File | Symbols |
|---|---|
evals/metrics.ts |
bootstrapTax, effectivePaidForEvent, effectivePaidInput, warmBootstrapTax, firstValidActionTokens (+3) |
evals/harness.ts |
startServer, runTask, stopServer, runSuite, mean (+2) |
evals/run.ts |
resolveModelConfig, loadCases, printBucketBreakdown, median, aggregateRepeats (+1) |
evals/adapters/claude-mcp-adapter.ts |
ClaudeMcpAdapter |
Entry Points
Start here when exploring this area:
runSuite(Function) —evals/harness.ts:73mean(Function) —evals/harness.ts:116bootstrapTax(Function) —evals/metrics.ts:55effectivePaidInput(Function) —evals/metrics.ts:84warmBootstrapTax(Function) —evals/metrics.ts:93
Key Symbols
| Symbol | Type | File | Line |
|---|---|---|---|
ClaudeMcpAdapter |
Class | evals/adapters/claude-mcp-adapter.ts |
42 |
runSuite |
Function | evals/harness.ts |
73 |
mean |
Function | evals/harness.ts |
116 |
bootstrapTax |
Function | evals/metrics.ts |
55 |
effectivePaidInput |
Function | evals/metrics.ts |
84 |
warmBootstrapTax |
Function | evals/metrics.ts |
93 |
firstValidActionTokens |
Function | evals/metrics.ts |
105 |
recoveryCostAfterFirstError |
Function | evals/metrics.ts |
116 |
successPer1kInputTokens |
Function | evals/metrics.ts |
127 |
scoreTrajectory |
Function | evals/metrics.ts |
139 |
evaluateGates |
Function | evals/harness.ts |
132 |
McpHarnessAdapter |
Interface | evals/harness.ts |
40 |
startServer |
Method | evals/harness.ts |
41 |
runTask |
Method | evals/harness.ts |
42 |
stopServer |
Method | evals/harness.ts |
43 |
effectivePaidForEvent |
Function | evals/metrics.ts |
74 |
resolveModelConfig |
Function | evals/run.ts |
38 |
loadCases |
Function | evals/run.ts |
60 |
printBucketBreakdown |
Function | evals/run.ts |
76 |
median |
Function | evals/run.ts |
94 |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 82 lines · 22 tokens per session scan A be3751e33331
evals is a skill published in the GitHub repository princeofscale/bloxforge (5 stars, last pushed 6d ago), licensed MIT. It adds 22 tokens to every session and 945 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
etsy-category-listing
Etsy category page scraper: given an Etsy category URL (e.g. https://www.etsy.com/c/jewelry) and optional page number, returns paginated product listings with listingId, shopId, title, url, image, salePrice, originalPrice, currency, rating, reviewCount, shopName, isAd, freeShipping, badge from category and subcategory…
human-approval
Request human approval before performing a SAFETY-CRITICAL, IRREVERSIBLE, or SCOPE-EXPANDING action — submit a structured context (action, scope, risk, consequence) plus options, then STOP the current turn. The platform redispatches the agent after the human decides. NEVER use for routine deliverables (writing docs /…
firebase-analytics
Use when logging analytics events, setting user properties, configuring default event parameters, building funnels, or adding screen-view tracking.
firebase-remote-config
Use when implementing feature flags, running A/B tests, setting parameter defaults, fetching/activating config, or enabling real-time config updates.
git-master
MUST USE whenever a task needs a commit or git-history investigation. Covers atomic commits, staging, commit-message style, rebase, squash, fixup/autosquash, blame, bisect, reflog, git log -S/-G, and questions like who wrote this or when was this added. Do not use for ordinary code edits unless the user asks for git…
prismer-evolve-record
Record the outcome of applying an evolution strategy. Use after resolving an error where prismer-evolve-analyze provided a recommendation, to feed back success or failure to the network.