model-race

model-race is a skill for Claude Code, Codex from machbuilds/atom. It costs 24 tokens per session (483 once invoked), scanned A, original, MIT.

A workflow for sending the same feature specification to several AI models in separate Git worktrees, which are independent working copies of a repository. It compares their results and lets you merge the selected version.

In plain words
What is it for?
Use it to start a model comparison, check its status, launch a model, run scorecards or judging, merge a winner, or abort the run.
Why use it?
It makes competing implementation approaches easier to test side by side. Separate worktrees keep each model's changes isolated while you compare them.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/machbuilds/atom/model-race
Any agent
npx skills add machbuilds/atom --skill model-race
Clone the repo
git clone --depth 1 https://github.com/machbuilds/atom

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for model-race

README.md
[![agentmods](https://agentmods.dev/badge/skills/machbuilds/atom/model-race.svg)](https://agentmods.dev/skills/machbuilds/atom/model-race)
Your own site
<a href="https://agentmods.dev/skills/machbuilds/atom/model-race"><img src="https://agentmods.dev/badge/skills/machbuilds/atom/model-race.svg" alt="Measured on agentmods" height="20"></a>
Per session 24 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 483 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00024 $0.00483
Opus 5 $0.00012 $0.00242
Sonnet 5 $0.00005 $0.00097
Haiku 4.5 $0.00002 $0.00048

Measured 3d ago against content hash e5fe4efebab4, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

model-race scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

scaffold/.claude/skills/model-race/SKILL.md · 59 lines

How it starts

The opening of the file, as written. The whole thing — 59 lines — stays where its author put it; the contents beside it link to each section on GitHub.

model-race skill (Claude wrapper)

Source of truth: see the Tooling > model-race section of AGENTS.md in this project. That file holds the canonical instructions every AI tool reads.

This file is the Claude-specific wrapper. It exists so Claude Code's Skill tool can register model-race as an invocable skill. Behavior comes from AGENTS.md.

Quick reference

model-race start <feature> --spec <file>     # create worktrees, write spec
model-race status                             # show race state
model-race launch <model>                     # open AI CLI in the model's worktree
model-race score                              # run automated scorecard
model-race judge                              # opt-in LLM evaluation
model-race merge <winner>                     # cherry-pick winner, clean losers
model-race abort                              # tear down (destructive)

model-race --help and model-race <command> --help show all flags.

When this skill activates (Claude-specific)

You should suggest model-race to the user only when:

  • The user is about to make a non-obvious decision with multiple reasonable approaches (algorithm choice, API shape, refactor pattern).
  • The work is high-stakes enough that the comparison cost is worth paying — typically anything that will be hard to change later.
  • The user has not already chosen a path.

Do NOT suggest it for:

  • CRUD endpoints, boilerplate, glue code.
  • Anything where the answer is obvious.
  • Bug fixes (these have one correct answer).

When the user races, the spec is critical

A race is only as good as its spec. Before model-race start, help the user write a spec that:

  • States the feature in one paragraph.
  • Lists testable acceptance criteria (what makes it correct).
  • Names constraints (perf budgets, API contracts, file boundaries).
  • Excludes solution details (don't pre-decide the approach).

A weak spec produces three confused implementations and no clear winner. A strong spec produces three differentiated approaches you can actually compare.

Read the full file on GitHub · 59 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 59 lines · 24 tokens per session scan A e5fe4efebab4

Subscribe to this mod's changes

model-race is a skill published in the GitHub repository machbuilds/atom (2 stars, last pushed 2mo ago), licensed MIT. It adds 24 tokens to every session and 483 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.