running-with-without-evals

running-with-without-evals is a skill for Claude Code, Codex from marcusrbrown/systematic. It costs 57 tokens per session (1,313 once invoked), scanned A, original, MIT.

A recorded comparison of the same coding task run with and without a plugin or skill loaded in OpenCode, a coding-agent environment.

In plain words
What is it for?
Use it to validate a plugin change, create a with-versus-without demonstration, or produce evidence for a public page without presenting a small sample as a benchmark.
Why use it?
It shows what the add-on actually changes while preserving a clean baseline and acknowledging what the agent already handles without it.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/marcusrbrown/systematic/running-with-without-evals
Any agent
npx skills add marcusrbrown/systematic --skill running-with-without-evals
Clone the repo
git clone --depth 1 https://github.com/marcusrbrown/systematic

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for running-with-without-evals

README.md
[![agentmods](https://agentmods.dev/badge/skills/marcusrbrown/systematic/running-with-without-evals.svg)](https://agentmods.dev/skills/marcusrbrown/systematic/running-with-without-evals)
Your own site
<a href="https://agentmods.dev/skills/marcusrbrown/systematic/running-with-without-evals"><img src="https://agentmods.dev/badge/skills/marcusrbrown/systematic/running-with-without-evals.svg" alt="Measured on agentmods" height="20"></a>
Per session 57 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,313 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00057 $0.01313
Opus 5 $0.00028 $0.00656
Sonnet 5 $0.00011 $0.00263
Haiku 4.5 $0.00006 $0.00131

Measured 5d ago against content hash 374675e72129, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

running-with-without-evals scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.agents/skills/running-with-without-evals/SKILL.md · 70 lines

How it starts

The opening of the file, as written. The whole thing — 70 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Running With/Without Evals

Overview

A with/without eval captures real opencode run transcripts on the same task, once with a plugin loaded and once without, to show what the plugin actually changes. The output is recorded sessions, not a benchmark.

Core principle: the result is only worth publishing if it would survive a hostile reader who has the transcripts. That means pre-registering the bar before you run, keeping a clean baseline, including the control that could disprove your thesis, and conceding what the baseline already does well.

The reusable harness lives at tests/manual/with-without-eval/run-arm.sh. Reuse it; don't reinvent the isolation.

When to use

  • Building or refreshing the with-without-systematic demo page
  • Validating that a plugin/skill change actually moves model behavior
  • Producing a credibility artifact for public promotion

When NOT to use

  • You need statistical claims — this is n=1–3 recorded sessions, not a benchmark. Don't dress it up as one.
  • The task is trivial (the baseline will look fine and the eval proves nothing)

The method

  1. Pre-register first. Write the task, the models, and the publishable-delta bar to a PRE-REGISTRATION*.md BEFORE any run. State the no-ship condition (if the bar isn't cleared, ship an honest decision guide, not a staged win). Committing this first is what lets you claim the result wasn't retrofitted.
  2. Run each arm via run-arm.sh <baseline|treatment> <model> <prompt-file> <out-dir> [seed-dir]. Baseline = --pure + empty plugin array; treatment = plugin loaded. Use a paid model — free models rate-limit and the failure masquerades as empty output.
  3. Verify the baseline is clean and the treatment loaded — the script checks systematic_skill is absent (baseline) / present (treatment). A contaminated baseline is a dishonest demo.
  4. Add the controls that matter (see below) before drawing conclusions.
  5. Have Oracle review the methodology and the drafted page before publishing. Council for the design if the matrix is non-trivial.

Read the full file on GitHub · 70 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 70 lines · 57 tokens per session scan A 374675e72129

Subscribe to this mod's changes

running-with-without-evals is a skill published in the GitHub repository marcusrbrown/systematic (24 stars, last pushed today), licensed MIT. It adds 57 tokens to every session and 1,313 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.