Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add ondrej-svec/heart-of-gold-toolkit --skill architecture-reviewgit clone --depth 1 https://github.com/ondrej-svec/heart-of-gold-toolkitWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/ondrej-svec/heart-of-gold-toolkit/architecture-review)<a href="https://agentmods.dev/skills/ondrej-svec/heart-of-gold-toolkit/architecture-review"><img src="https://agentmods.dev/badge/skills/ondrej-svec/heart-of-gold-toolkit/architecture-review/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/ondrej-svec/heart-of-gold-toolkit/architecture-review"><img src="https://agentmods.dev/badge/skills/ondrej-svec/heart-of-gold-toolkit/architecture-review.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 2 findings, up to medium
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- medium Excessive Agency · line 38 Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.Fix: Add human-in-the-loop confirmation for destructive, irreversible, or high-impact operations. Never auto-execute commands that modify files, send data, or alter system state.
- medium Excessive Agency · line 257 Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.Fix: Add human-in-the-loop confirmation for destructive, irreversible, or high-impact operations. Never auto-execute commands that modify files, send data, or alter system state.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00103 | $0.02493 |
| Opus 5 | $0.00051 | $0.01247 |
| Sonnet 5 | $0.00021 | $0.00499 |
| Haiku 4.5 | $0.00010 | $0.00249 |
Grade A, and why
architecture-review scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 261 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Architecture Review
Deep, evidence-based architectural review that produces a decision-grade document. Not a checklist pass — a systematic investigation that cross-references every claim against actual code, maps failure modes through the full lifecycle, and evaluates whether the architecture can support its intended scale.
Boundaries
This skill MAY: read code, analyze architecture, research platforms, write review documents, ask architectural decision questions via AskUserQuestion. This skill MAY NOT: modify application code, deploy changes, or make architectural changes. It produces a review document — the team decides what to act on.
Common Rationalizations
| Shortcut | Why It Fails | The Cost |
|---|---|---|
| "Just summarize the issue/checklist" | Restating what the team already knows adds no value | CTO handoff that tells the team nothing new |
| "The README says X, so X is true" | Code drifts from docs constantly | False claims in the review become false decisions |
| "Recommend best practices without checking constraints" | A 3-person vibe-coding team can't run Kubernetes | Recommendations that are technically correct but impossible to execute |
| "Skip the edge cases — focus on happy path" | Production failures happen on the unhappy path | Architecture that works in demos but breaks under real load |
| "Trust the security audit's line numbers" | Code changes daily; audits are snapshots | Wrong file:line references destroy reviewer credibility |
Phase 1: Scope and Context
Entry: User invokes /architecture-review with a target (repo path, issue URL, or description).
1.1 Identify the target
| Input | Action |
|---|---|
| GitHub issue URL | Fetch issue body via gh issue view |
| Repo path | Explore codebase structure |
| Description | Clarify scope via AskUserQuestion |
1.2 Gather context (parallel subagents)
Launch these concurrently:
- Codebase explorer (subagent, sonnet): Map the tech stack, directory structure, API routes, database schema, external dependencies, deployment config. Read
package.json/requirements.txt,Dockerfile, CI configs,README.md,CLAUDE.md.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 261 lines · 103 tokens per session scan A 44931832ec01
architecture-review is a skill published in the GitHub repository ondrej-svec/heart-of-gold-toolkit (19 stars, last pushed 23d ago), licensed MIT. It adds 103 tokens to every session and 2,493 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
autoreview
Pre-commit/ship code review: Codex default; optional Claude or Pi.
omh-code-review
This is a Hermes-native code-review workflow skill.
revdiff-plan
Review the last Codex assistant message (plan, analysis, or proposal) with inline annotations in a TUI overlay. Extracts the most recent response from Codex rollout files and opens it in revdiff for review and annotation. Activates on "revdiff-plan", "review plan with revdiff", "annotate plan", "review last response"…
code-reviewer
Code review specialist focused on patterns, bugs, security, and performance.
full-repo-review
Comprehensive four-wave review of all repo source files, producing a prioritized issue backlog.
agent-teams-simplify-and-harden
Implementation + audit loop using parallel agent teams with structured simplify, harden, and document passes. Spawns implementation agents to do the work, then audit agents to find complexity, security gaps, and spec deviations, then loops until code compiles cleanly, all tests pass, and auditors find zero issues or…