Getting it into your agent
This one installs as part of its plugin. Adding the marketplace and installing the plugin brings it with everything else the plugin ships.
/plugin marketplace add vibeic/vibe-ic/plugin install vibe-icWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/vibeic/vibe-ic/benchmark-verify)<a href="https://agentmods.dev/skills/vibeic/vibe-ic/benchmark-verify"><img src="https://agentmods.dev/badge/skills/vibeic/vibe-ic/benchmark-verify/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/vibeic/vibe-ic/benchmark-verify"><img src="https://agentmods.dev/badge/skills/vibeic/vibe-ic/benchmark-verify.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00211 | $0.04045 |
| Opus 5 | $0.00105 | $0.02022 |
| Sonnet 5 | $0.00042 | $0.00809 |
| Haiku 4.5 | $0.00021 | $0.00404 |
Grade A, and why
benchmark-verify scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 212 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Benchmark Verify — normalized doc→production-ready verification
We run benchmark ICs to validate the whole Vibe-IC flow: that starting from a complete set of Design Documents we can reach a production-ready result. A benchmark run is not finished when the flow returns a verdict — it is finished when this skill's five pillars all pass and the single report is written.
This skill is chip-AGNOSTIC and MANDATORY for every benchmark IC. Apply it
to each IC under benchmark_clean/<ic>/ (and any future benchmark set).
⛔ Prerequisite — honesty + the corrected protocol
Read benchmark_clean/METHODOLOGY.md first. Non-negotiable:
- Input = Design Documents ONLY (incl. explicit PDK); nothing produced by a later phase is fed back as input.
- Phase-2 RTL must be GENERATED; IP reuse is allowed but every reused source MUST be
tagged
REUSED-IPinSOURCE_MANIFEST.md(vsGENERATED). Production-readiness credit applies only to GENERATED content; reused content is reported separately. - No vacuous result counts as PASS (e.g. a Magic
gds readthat dropped geometry and reports "0 DRC" is INCONCLUSIVE, not clean). A missing input is PENDING, never a silent PASS. Enforced byprograms/drc_vacuous_pass_check.py— flips a 0-violation DRC verdict to INCONCLUSIVE unless the log proves geometry was loaded; SKIPs (never PASSes) when no DRC log exists. - The upstream open-source IP is the golden oracle for cross-check only — never a Phase-1/2 input.
The five pillars (== report sections == hard gates)
Pillar 1 — Functional Verification Coverage ▸ gate: == 100%
LLM judgment (irreducible): walk L1–L27 (functional spec, interface, register map, command/protocol, timing, behavioral sequences, test cases) and enumerate every requirement — a program cannot reliably know that a prose timing sentence in L8 is a distinct testable requirement (vs a restatement) nor author the directed test that covers it. For each requirement, bind a verification item (directed test, assertion, golden-vector check, or formal property) and confirm it PASSES against the generated RTL.
Functional Coverage = verified_requirements / total_requirements→ must be 100% (the == 100% gate is enforced byprograms/benchmark_verify_report.py).- Emit
reports/functional_coverage.json:{"requirements":[{"id","source","desc","status":"PASS|FAIL|PENDING"}]}. - Closed-loop: if < 100%, write the missing tests and/or fix the RTL (use
rtl-repair) and re-verify until 100%. Do not waive a requirement; if a doc requirement is genuinely untestable, that is a spec defect to record, not a pass.
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 212 lines · 211 tokens per session scan A 08e24f45c369
benchmark-verify is a skill published in the GitHub repository vibeic/vibe-ic (23 stars, last pushed yesterday), licensed Apache-2.0. It adds 211 tokens to every session and 4,045 once invoked, about $0.0011 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
analog-verify
Pre-simulation review and Spectre simulation verification for analog circuits. Reviews circuit netlist and testbench, runs simulation, produces margin report. Use after analog-design completes a netlist.
boardrepo
Read and review real PCB projects on BoardRepo. Search published hardware designs, open a board's schematic connectivity, bill of materials and files, run KiCad's DRC and ERC, and check a design against a fabrication house's limits. Use whenever the user names a BoardRepo board or URL, asks to find a published board…
analog-netlist-crawl
Crawl and analyze post-layout parasitic netlists without running SPICE. Answers "what's the effective resistance from node A to node B across this massive R mesh?", "inside the VREFN mesh, which device pins are electrically farthest apart?", "which nets have the worst coupling?", "where does settling bottleneck?" — by…
analog-design
Transistor-level circuit design for one analog sub-block. Produces Spectre netlist with hand-calculation rationale. Use when designing a specific circuit block after architecture is defined.
analog-pipeline
MANDATORY — MUST load this skill when the user mentions: OTA, ADC, PLL, comparator, bandgap, LDO, amplifier, opamp, or any analog/mixed-signal IC design task. Full analog design pipeline: spec -> architecture -> design -> verify -> deliver. Orchestrates analog-decompose, analog-behavioral, analog-design…
analog-audit
Audit analog circuit netlists for correctness, quality, and risks. Supports both pre-layout (schematic) and post-layout (extracted) netlists. Post-layout mode filters massive parasitic netlists before auditing. Works without EDA. TRIGGER on: "audit", "review netlist", "check this circuit", "design review"…