benchmark

benchmark is a skill for Claude Code, Codex from tmargolis/career-navigator. It costs 41 tokens per session (1,227 once invoked), scanned A, original, Apache-2.0.

A career-data benchmarking tool that compares your job-search pipeline with typical results for similar roles, experience levels, company sizes, and locations.

In plain words
What is it for?
Use it to review application-to-response, interview, and offer rates; response timelines; applicant-tracking-system scores; and compensation positioning.
Why use it?
It shows whether your application progress, response times, interview conversions, and pay expectations are within a normal range. It also requires enough application history to make the comparison meaningful.

Skill for Claude CodeCodex

Part of the career-navigator plugin — 46 skills, 3 MCP servers shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/tmargolis/career-navigator/benchmark
Any agent
npx skills add tmargolis/career-navigator --skill benchmark
Clone the repo
git clone --depth 1 https://github.com/tmargolis/career-navigator

Made for: Claude Code, Codex.

Or install career-navigator, the plugin that ships this one along with the rest of its 46 skills, 3 MCP servers.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for benchmark

README.md
[![agentmods](https://agentmods.dev/badge/skills/tmargolis/career-navigator/benchmark.svg)](https://agentmods.dev/skills/tmargolis/career-navigator/benchmark)
Your own site
<a href="https://agentmods.dev/skills/tmargolis/career-navigator/benchmark"><img src="https://agentmods.dev/badge/skills/tmargolis/career-navigator/benchmark.svg" alt="Measured on agentmods" height="20"></a>
Per session 41 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,227 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00041 $0.01227
Opus 5 $0.00020 $0.00613
Sonnet 5 $0.00008 $0.00245
Haiku 4.5 $0.00004 $0.00123

Measured 4d ago against content hash 1ef342316d6d, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

benchmark scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/benchmark/SKILL.md · 97 lines

How it starts

The opening of the file, as written. The whole thing — 97 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Benchmark the user's pipeline performance against industry norms for their role, level, company size, and geography.

Workflow

1. Check data threshold

Application data uses the split layout defined in references/tracker-schema.md — read it before any read or write.

Read {user_dir}/CareerNavigator/tracker.json. Count the total number of applications (any status) from the summary rows — no detail files needed for this check. If fewer than 5:

"You need at least 5 applications to run a meaningful benchmark — you have {n} so far. Keep logging applications via /career-navigator:track-application and run this again once you have more history."

Stop here if below threshold.

If ≥5 but fewer than 10 resolved outcomes, proceed with a note that results are preliminary.

2. Load the stage history the conversion math needs

Every PIPELINE CONVERSION and TIMELINES figure below is computed from stage_history[], and tracker.json alone contains no stage history — it moved to the per-application detail files. Computing app → response, screen → interview, interview → offer, or days-to-response from tracker.json by itself silently yields zeros and reports the user as far below norm when they are not.

  1. Read {user_dir}/CareerNavigator/tracker.json and take applications[].
  2. Iterate every row and load its detail_file (relative to CareerNavigator/) for that application's stage_history[]. Skip a row only when its stage_count is 0 — that row genuinely has no stages to count.
  3. Use the summary fields where they answer the question directly: latest_stage and latest_stage_date give the current funnel position and last movement date, and stage_count / notes_count / contact_count give volumes without a read.

3. Invoke analyst — Operation 4

Hand off to the analyst agent with:

  • CareerNavigator/tracker.json (summary rows) plus the loaded applications/<application_id>.json detail files — pass both; conversion and timeline math is impossible from summary rows alone
  • The full CareerNavigator/artifacts-index.json
  • The full CareerNavigator/profile.md
  • Instruction to run Operation 4: Market Benchmark

Read the full file on GitHub · 97 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 97 lines · 41 tokens per session scan A 1ef342316d6d

Subscribe to this mod's changes

benchmark is a skill published in the GitHub repository tmargolis/career-navigator (14 stars, last pushed 6d ago), licensed Apache-2.0. It adds 41 tokens to every session and 1,227 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

brainstorming

You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.

obra/superpowers · 37 tokens

auto-perf-optimize

Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.

microsoft/vscode · 62 tokens

chat-perf

Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.

microsoft/vscode · 51 tokens

chat-pet-sprite-creation

Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.

microsoft/vscode · 53 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens