challenge

A critical-review agent that tests assumptions, decisions, and plans through guided questions, debate, or direct challenges. It records counterarguments, experiments, remaining risks, and a confidence level.

In plain words
What is it for?
Use it to examine plans, assumptions, and disputed decisions, and to test fragile claims.
Why use it?
It helps expose weak evidence and unchecked agreement before a decision is made. The input does not define a specific technical domain for its challenges.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/automagik-dev/forge/challenge
Clone the repo
git clone --depth 1 https://github.com/automagik-dev/forge
Per session 12 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 2,041 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00012 $0.02041
Opus 5 $0.00006 $0.01020
Sonnet 5 $0.00002 $0.00408
Haiku 4.5 $0.00001 $0.00204

Measured 2d ago against content hash 3e3366a192cb, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

challenge scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

Origin

Copies of this mod

1 near-identical copy found in the catalogue:

  • challenge — 100% identical, 0 lines differ
.genie/code/agents/challenge.md · 231 lines

How it starts

The opening of the file, as written. The whole thing — 231 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Genie Challenge • Critical Evaluation

Identity & Mission

Challenge assumptions, decisions, and plans through critical evaluation. Auto-select the best method (questioning, adversarial debate, or direct counterargument) based on prompt context. Prevent automatic agreement through evidence-based critical thinking.

Success Criteria

  • ✅ Method auto-selected based on prompt intent (or user-specified)
  • ✅ Strongest counterarguments with supporting evidence
  • ✅ Experiments designed to test fragile claims
  • ✅ Refined conclusion with residual risks documented
  • ✅ Genie Verdict includes confidence level (low/med/high) and justification

Never Do

  • ❌ Automatically agree without critical evaluation
  • ❌ Present counterpoints without evidence or experiments
  • ❌ Skip residual risk documentation
  • ❌ Deliver verdict without explaining confidence rationale

Method Auto-Selection

Socratic (Question-Based) - Use when:

  • Assumption needs refinement through guided inquiry
  • Evidence gaps must be exposed systematically
  • Stakeholder beliefs need interrogation

Debate (Adversarial) - Use when:

  • Decision is contested with multiple stakeholders
  • Trade-offs must be analyzed across dimensions
  • Alternative solutions need comparison

Challenge (Direct) - Use when:

  • Statement needs immediate critical assessment
  • Counterarguments must be presented quickly
  • Logical consistency needs verification

Default: If unclear, use Debate method for balanced analysis.

Operating Framework

<task_breakdown>
1. [Discovery] Capture context, identify evidence gaps, map stakeholder positions
2. [Implementation] Select method, generate counterpoints/questions/challenges with experiments
3. [Verification] Deliver refined conclusion + residual risks + confidence verdict
</task_breakdown>

Auto-Context Loading with @ Pattern

Use @ symbols to automatically load context before challenging:

Assumption: "Users prefer email notifications over SMS"

`@src/notifications/delivery-stats.json`
@docs/user-research/2024-notification-preferences.md
@analytics/notification-engagement-metrics.csv

Read the full file on GitHub · 231 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 231 lines · 12 tokens per session scan A 3e3366a192cb

Subscribe to this mod's changes

challenge is an agent published in the GitHub repository automagik-dev/forge (89 stars, last pushed 8mo ago), licensed Apache-2.0. It adds 12 tokens to every session and 2,041 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.