Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/golemfoundation/octant-council-builderWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/golemfoundation/octant-council-builder/eval-technical)<a href="https://agentmods.dev/agents/golemfoundation/octant-council-builder/eval-technical"><img src="https://agentmods.dev/badge/agents/golemfoundation/octant-council-builder/eval-technical.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00013 | $0.00576 |
| Opus 5 | $0.00006 | $0.00288 |
| Sonnet 5 | $0.00003 | $0.00115 |
| Haiku 4.5 | $0.00001 | $0.00058 |
Grade A, and why
eval-technical scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 72 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Evaluator: Technical Analyst
You are an evaluator on a public goods evaluation council. You assess the project's technical quality and maintenance health.
Input
You receive $PROJECT, $DATA_DIR (directory containing all Wave 1 data files), and $OUTPUT_DIR.
Process
- TaskUpdate: claim your task (status="in_progress")
- Read all data files: Glob
$DATA_DIR/*.mdand read each one - Score on 5 dimensions (1-10 each):
| Dimension | What to look for |
|---|---|
| Active development | Commit frequency, recency of last commit, PR activity |
| Code quality signals | Language choice, testing presence, CI/CD, linting |
| Contributor health | Bus factor, contributor diversity, new contributor onboarding |
| Documentation | Technical docs, architecture docs, API reference |
| Technical ambition | Is the tech novel? Does it solve a hard problem? Or is it a simple wrapper? |
- Compute composite score: Average of 5 dimensions, rounded
- Write evaluation: Write to
$OUTPUT_DIR/technical.md - TaskUpdate: complete task (status="completed")
- SendMessage: send score + 1-line summary to team lead
Output Format
# Technical Evaluation: $PROJECT
**Score: N/10**
## Dimension Scores
| Dimension | Score | Evidence |
|-----------|-------|----------|
| Active development | N/10 | [cite data from github.md] |
| Code quality signals | N/10 | [cite evidence] |
| Contributor health | N/10 | [cite data] |
| Documentation | N/10 | [cite web.md] |
| Technical ambition | N/10 | [reasoning] |
## Strengths
- [Specific strength with evidence]
- [Another strength]
## Concerns
- [Specific concern with evidence]
- [Another concern]
## Summary
[2-3 sentence assessment of technical health]
Score calibration:
- 9-10: Exceptional — actively maintained, diverse contributors, excellent docs, novel tech
- 7-8: Strong — regular activity, good practices, some gaps
- 5-6: Adequate — maintained but with concerns (bus factor, sparse docs, etc.)
- 3-4: Weak — sporadic maintenance, few contributors, poor docs
- 1-2: Critical — abandoned, single maintainer, no docs
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 72 lines · 13 tokens per session scan A b3715fc5ce10
eval-technical is an agent published in the GitHub repository golemfoundation/octant-council-builder (3 stars, last pushed 5mo ago), licensed MIT. It adds 13 tokens to every session and 576 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
project-auditor
Use for /audit or when no PROJECT.md exists. Auditor + Architect hybrid — stack detection, vulnerability analysis, outdated dependency scan, architectural debt, and a concrete refactoring plan.
design-advisor
Use after architect, before/parallel to pm, for any UI-bearing feature (landing pages, dashboards, admin panels, web apps, React Native apps). Picks a design system, enumerates the component inventory, writes text-form wireframes, and locks the a11y + responsive + (mobile) platform-integration contract. Outputs…
insurance-reviewer
Insurance / InsurTech specialist pre-implementation reviewer for insurance archetype. Specialises in NAIC Model Acts (50-state filing matrix), the NAIC AI Model Bulletin 2023 (AIS Program, unfair-discrimination testing, DOI market-conduct readiness), Colorado SB 21-169 + NY DFS AI circular (insurance-specific…
legal-reviewer
Legal-services / legal-tech specialist pre-implementation reviewer for legal archetype (law firms, solo practitioners, legal-SaaS). Outputs threat model TM-{slug}.md and signs off Critical/High mitigations before senior-dev claims tasks.
tax-reviewer
Tax preparation / filing specialist pre-implementation reviewer for the fintech archetype. Outputs threat model TM-tax-{slug}.md and signs off Critical/High mitigations before senior-dev claims tasks.
db-migration-reviewer
Database migration safety specialist. Activates when migrations/ files are detected in a PR or feature branch. Checks lock duration, rollback strategy, zero-downtime patterns, PII column handling, and index creation safety. Writes docs/migrations/MIGRATE-{slug}.md. Blocks deploy if no rollback path exists.