evaluate-tradeoffs

evaluate-tradeoffs is a skill for Claude Code, Codex from echoo19/decision-simulator. It costs 97 tokens per session (489 once invoked), scanned A, original, MIT.

A guide for comparing different technical ways to reach the same engineering goal. It focuses on practical effects such as delivery risk, code ownership, testing, rollback, and future changes.

In plain words
What is it for?
Use it to compare a rewrite with an incremental refactor, a service split with extending a monolith, or a direct integration with a wrapper library.
Why use it?
It shows where complexity and migration work will land instead of judging an approach only by how neat it looks on paper. It also identifies when a recommendation would change.

Skill for Claude CodeCodex

Part of the decision-simulator plugin — 4 skills, 1 agent shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/echoo19/decision-simulator/evaluate-tradeoffs
Any agent
npx skills add echoo19/decision-simulator --skill evaluate-tradeoffs
Clone the repo
git clone --depth 1 https://github.com/echoo19/decision-simulator

Made for: Claude Code, Codex.

Or install decision-simulator, the plugin that ships this one along with the rest of its 4 skills, 1 agent.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for evaluate-tradeoffs

README.md
[![agentmods](https://agentmods.dev/badge/skills/echoo19/decision-simulator/evaluate-tradeoffs.svg)](https://agentmods.dev/skills/echoo19/decision-simulator/evaluate-tradeoffs)
Your own site
<a href="https://agentmods.dev/skills/echoo19/decision-simulator/evaluate-tradeoffs"><img src="https://agentmods.dev/badge/skills/echoo19/decision-simulator/evaluate-tradeoffs.svg" alt="Measured on agentmods" height="20"></a>
Per session 97 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 489 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00097 $0.00489
Opus 5 $0.00048 $0.00244
Sonnet 5 $0.00019 $0.00098
Haiku 4.5 $0.00010 $0.00049

Measured 4d ago against content hash d1aeb711829f, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

evaluate-tradeoffs scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/decision-simulator/skills/evaluate-tradeoffs/SKILL.md · 74 lines

What it actually says

Evaluate Tradeoffs

Compare options across the dimensions that are actually decision-driving in this case. Do not blindly enumerate every possible factor.

Method

  1. Name the options clearly.
  2. Choose the smallest set of relevant dimensions.
  3. For each dimension, explain the mechanism behind the difference.
  4. Identify which dimensions are decisive versus merely nice-to-have.
  5. Say when one option dominates and why.
  6. State the conditions that would flip the recommendation.

Common Dimensions

Use only the dimensions that matter here:

  • UX
  • DX
  • system boundaries
  • maintainability
  • delivery speed
  • scalability
  • performance
  • reliability
  • security and privacy
  • infra cost
  • engineering time
  • team coordination overhead
  • debugging difficulty
  • flexibility
  • reversibility
  • operational burden

Quality Bar

  • Avoid empty labels like "faster" or "cleaner" without explaining what becomes faster or cleaner.
  • Make asymmetry visible. Some options will have one major win and three serious liabilities.
  • Highlight non-obvious tradeoffs first.
  • Prefer saying "Option A dominates because..." over pretending the choice is balanced.

Suggested Output

Options

List the options.

Relevant dimensions

Name the dimensions that matter and why they matter here.

Tradeoff analysis

For each decisive dimension, compare the options with concrete implications.

Dominant option

Name the option that wins on the dimensions that matter most.

If no option clearly dominates, do not hedge — name the deciding unknown: the single missing input that, once resolved, would break the tie. Do not leave this section as "it depends."

Second-order effects

What does this choice change three weeks or three months later? Consider: team behavior, debugging surface, architecture shape, ownership burden, or rule drift.

Recommendation

Make a call and explain what would change the answer.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 74 lines · 97 tokens per session scan A d1aeb711829f

Subscribe to this mod's changes

evaluate-tradeoffs is a skill published in the GitHub repository echoo19/decision-simulator (4 stars, last pushed 3mo ago), licensed MIT. It adds 97 tokens to every session and 489 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

brainstorming

You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.

obra/superpowers · 37 tokens

auto-perf-optimize

Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.

microsoft/vscode · 62 tokens

chat-perf

Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.

microsoft/vscode · 51 tokens

chat-pet-sprite-creation

Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.

microsoft/vscode · 53 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens