local-platform-e2e

A procedure for running benchmark tests against a local copy of the benchmarks platform. The setup uses local database and object-storage services, so it does not require cloud access or provider credentials.

In plain words
What is it for?
Use it to start the platform locally and run a real `@benchsdk/runner` benchmark. It covers creating a benchmark, running workers, recording results, uploading artifacts, and completing the run.
Why use it?
It lets developers test the full benchmark workflow without depending on external services. It helps investigate problems in worker planning, task reporting, artifact uploads, and run completion.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/computesdk/benchmarks/local-platform-e2e
Any agent
npx skills add computesdk/benchmarks --skill local-platform-e2e
Clone the repo
git clone --depth 1 https://github.com/computesdk/benchmarks

Made for: Claude Code, Codex.

Per session 82 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 3,273 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00082 $0.03273
Opus 5 $0.00041 $0.01636
Sonnet 5 $0.00016 $0.00655
Haiku 4.5 $0.00008 $0.00327

Measured 2d ago against content hash 78b4c0a322c7, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

local-platform-e2e scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

curl -s "http://127.0.0.1:8123/?user=default&password=chpass" \
.agents/skills/local-platform-e2e/SKILL.md · 227 lines

How it starts

The opening of the file, as written. The whole thing — 227 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Local end-to-end: @benchsdk/runner ↔ benchmarks-platform

Goal: exercise upsert benchmark → create run → planWorkers → claim → heartbeat → task_results → artifact upload → complete → dashboard, with zero external credentials.

1. Infra (docker)

docker run -d --name pg -p 5432:5432 -e POSTGRES_PASSWORD=postgres -e POSTGRES_DB=bench postgres:16
docker run -d --name minio -p 9000:9000 -e MINIO_ROOT_USER=minioadmin -e MINIO_ROOT_PASSWORD=minioadmin \
  quay.io/minio/minio server /data
docker run -d --name ch -p 8123:8123 -e CLICKHOUSE_PASSWORD=chpass clickhouse/clickhouse-server:25.6
docker run --rm --network host --entrypoint sh quay.io/minio/mc -c \
  "mc alias set local http://127.0.0.1:9000 minioadmin minioadmin && mc mb -p local/bench"
  • Use ClickHouse >= 25.x: 24.8 fails ch:migrate with TTL expression result column should have DateTime or Date type, but has DateTime64(3,'UTC').
  • The importer passes clickhouse_settings: { date_time_input_format: 'best_effort' } per insert, so a default-configured server works. If you hit Cannot parse input: expected '"' before: 'Z"...' (older code), work around it server-side, and remember to REMOVE the override before verifying an importer fix — otherwise the server setting masks it:
    docker exec ch bash -c 'mkdir -p /etc/clickhouse-server/users.d && printf "<clickhouse><profiles><default><date_time_input_format>best_effort</date_time_input_format></default></profiles></clickhouse>" > /etc/clickhouse-server/users.d/besteffort.xml'
    docker restart ch
    # verify which mode is actually active:
    curl -s "http://127.0.0.1:8123/?user=default&password=chpass" \
      --data-binary "SELECT value FROM system.settings WHERE name='date_time_input_format'"
    

2. benchmarks-platform .env.local

Point TIGRIS_* at MinIO — the events and artifacts routes require S3 or they return 502 on every batch. region: "auto" + presigned PUT works with MinIO.

DATABASE_URL=postgresql://postgres:[email protected]:5432/bench
DATABASE_URL_UNPOOLED=postgresql://postgres:[email protected]:5432/bench
CLICKHOUSE_URL=http://127.0.0.1:8123
CLICKHOUSE_DATABASE=default
CLICKHOUSE_USER=default
CLICKHOUSE_PASSWORD=chpass
ADMIN_API_KEY=local-admin-key
BETTER_AUTH_SECRET=<openssl rand -hex 32>
BETTER_AUTH_URL=http://localhost:3000
TIGRIS_ACCESS_KEY_ID=minioadmin
TIGRIS_SECRET_ACCESS_KEY=minioadmin
TIGRIS_STORAGE_ENDPOINT=http://127.0.0.1:9000
TIGRIS_BUCKET=bench

Read the full file on GitHub · 227 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 227 lines · 82 tokens per session scan A 78b4c0a322c7

Subscribe to this mod's changes

local-platform-e2e is a skill published in the GitHub repository computesdk/benchmarks (110 stars, last pushed 4d ago), licensed MIT. It adds 82 tokens to every session and 3,273 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

next-cache-components-adoption

Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…

vercel/next.js · 95 tokens

babysit-pr

Babysit a GitHub pull request after creation by continuously polling review comments, CI checks/workflow runs, and mergeability state until the PR is merged/closed or user help is required. Diagnose failures, retry likely flaky failures up to 3 times, auto-fix/push branch-related issues when appropriate, and keep…

openai/codex · 114 tokens

imagegen

Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output…

openai/codex · 113 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens

next-cache-components-optimizer

Drive a Next.js route to instant navigation by setting up an agentic loop, under Cache Components / PPR, on initial load (hard navigation) and client-side navigation (soft navigation). Encode the goal as a failing @next/playwright instant() e2e and work it to green, one verified route at a time; the shipped test then…

vercel/next.js · 170 tokens