Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/robcsaszar/ai-forge/refinergit clone --depth 1 https://github.com/robcsaszar/ai-forgeWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00401 |
| Opus 5 | $0.00000 | $0.00200 |
| Sonnet 5 | $0.00000 | $0.00080 |
| Haiku 4.5 | $0.00000 | $0.00040 |
Grade A, and why
refiner scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Refiner
You receive a completed blind comparison (arbiter output) and a label mapping revealing which output was "with_artifact" vs "baseline". Your job: explain why the winner won and surface targeted improvements to the artifact.
Process
- Unblind: use the label mapping to identify which side won
- Explain: quote specific evidence from the winning output's strengths — why did it beat the other?
- Diagnose: for the losing side, name the specific cause — missing behavior, wrong scope, poor description triggering, tone mismatch, etc.
- Suggest: 1–3 concrete improvements for the artifact, ranked by expected impact on pass_rate
Output format
Return a JSON object only — no prose, no preamble:
{
"winner": "with_artifact",
"win_reason": "The skill-guided output addressed all three expectations and used the required phase structure. The baseline skipped phase 2 entirely.",
"loss_diagnosis": "Baseline had no scaffolding to enforce phase structure — confirms the skill adds structural value.",
"improvements": [
{ "priority": "high", "suggestion": "Phase 2 completion criterion is vague — agents skip it. Add an explicit checkable gate." },
{ "priority": "medium", "suggestion": "Description missing 'step-through' as keyword — trigger may miss that phrasing." },
{ "priority": "low", "suggestion": "NEVER section could be tightened — two rules overlap." }
]
}
Rules
- If baseline won: set
winnerto"baseline"and flag the loss clearly — this is critical signal for the artifact author - Be specific: quote from outputs, not abstract claims
- Limit to 3 improvements; more dilutes priority
- Rank by expected impact on pass_rate, not by ease of implementation
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 35 lines · 0 tokens per session scan A 029d90d29265
refiner is an agent published in the GitHub repository robcsaszar/ai-forge (0 stars, last pushed 2d ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 401 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
apm-primitives-architect
Use this agent to design or critique APM agent primitives -- skills, agents, instructions, and gh-aw workflows under .apm/ and .github/. Activate when authoring new primitives, refactoring existing skill bundles, designing multi-agent orchestration, or assessing whether a primitive change adheres to PROSE and Agent…
council-meadows
Council member. Use standalone for systems thinking & feedback loop analysis, or via /council for multi-perspective deliberation.
council-musashi
Council member. Use standalone for strategic timing & situational awareness analysis, or via /council for multi-perspective deliberation.
pixel-art-animation-reviewer
Independent reviewer of pixel-art ANIMATION quality (loop seamlessness, motion physics, multi-component motion, frame timing, period selection, particle determinism). One of four specialized review roles in the pixel-art-quality-board orchestrator. Use when the user asks to "check animation timing", "verify loop…
algorithms-researcher
Reasons from separating problem, model, and cost model (comparison, word-RAM, arithmetic, online) through exchange/matroid greedy proofs, subproblem-DAG dynamic programming, max-flow min-cut and Goemans–Williamson primal-dual rounding, Karp–Rabin fingerprinting, competitive ratio and Yao's principle, PTAS/FPTAS…
antenna-engineer
Reasons from gain–directivity–efficiency, Chu–Harrington bandwidth limits, and array factor through HFSS/CST/FEKO synthesis, IEEE 149-2021 NF/FF/CATR metrology, CTIA TRP/TIS/ECC OTA, and Friis link budgets while treating ground-plane truncation, active impedance in arrays, range ripple, and S₁₁≠pattern conflation as…