surface-benchmarker

surface-benchmarker is a skill for Claude Code, Codex from mitchdenny/hex1b. It costs 38 tokens per session (2,970 once invoked), scanned A, original, MIT.

A set of instructions for running and reading performance benchmarks for the Surface API, a part of the Hex1b library that represents terminal display surfaces.

In plain words
What is it for?
Use it after modifying files under src/Hex1b/Surfaces/ to run all, targeted, or quick dry-run benchmarks.
Why use it?
It helps check whether changes to surface-related code make the library slower. It identifies which benchmark commands to run for affected files and areas.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/mitchdenny/hex1b/surface-benchmarker
Any agent
npx skills add mitchdenny/hex1b --skill surface-benchmarker
Clone the repo
git clone --depth 1 https://github.com/mitchdenny/hex1b

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for surface-benchmarker

README.md
[![agentmods](https://agentmods.dev/badge/skills/mitchdenny/hex1b/surface-benchmarker.svg)](https://agentmods.dev/skills/mitchdenny/hex1b/surface-benchmarker)
Your own site
<a href="https://agentmods.dev/skills/mitchdenny/hex1b/surface-benchmarker"><img src="https://agentmods.dev/badge/skills/mitchdenny/hex1b/surface-benchmarker.svg" alt="Measured on agentmods" height="20"></a>
Per session 38 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,970 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00038 $0.02970
Opus 5 $0.00019 $0.01485
Sonnet 5 $0.00008 $0.00594
Haiku 4.5 $0.00004 $0.00297

Measured 5d ago against content hash 989f935cc8e6, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

surface-benchmarker scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.github/skills/surface-benchmarker/SKILL.md · 382 lines

How it starts

The opening of the file, as written. The whole thing — 382 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Surface Benchmarker Skill

This skill provides guidelines for AI agents to run and interpret performance benchmarks for the Hex1b Surface API. Run benchmarks whenever you modify code in src/Hex1b/Surfaces/ to ensure performance is not regressed.

Quick Reference

Action Command
Run all Surface benchmarks dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "Surface*"
Run specific benchmark dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "*WriteText*"
Quick dry-run (sanity check) dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "Surface*" --job dry

When to Run Benchmarks

ALWAYS run benchmarks after modifying:

File/Area Critical Benchmarks
SurfaceCell.cs All (cells are used everywhere)
Surface.cs CreateSurface_*, WriteText_*, Fill_*, Clone_*
CompositeSurface.cs CompositeSurface_*
SurfaceComparer.cs Compare_*, ToTokens_*, ToAnsiString_*
SurfaceDiff.cs Compare_*
ComputeContext.cs CompositeSurface_* (computed cells use this)

Benchmark Harness Architecture

Project Structure

benchmarks/Hex1b.Benchmarks/
├── Hex1b.Benchmarks.csproj    # Console app with BenchmarkDotNet
├── Program.cs                  # Entry point
└── SurfaceBenchmarks.cs        # Surface API benchmarks

How BenchmarkDotNet Works

BenchmarkDotNet is a .NET performance benchmarking library that:

  1. Warms up the JIT by running the benchmark multiple times before measuring
  2. Runs multiple iterations to get statistically significant results
  3. Reports mean, median, standard deviation for each benchmark
  4. Tracks memory allocations (when [MemoryDiagnoser] is enabled)

Benchmark Lifecycle

┌─────────────────────────────────────────────────────────────────┐
│ 1. GlobalSetup                                                  │
│    - Runs once before all benchmarks                            │
│    - Creates pre-allocated surfaces, diffs, etc.                │
│    - Sets up shared test data                                   │
└─────────────────────────────────────────────────────────────────┘
                              ▼
┌─────────────────────────────────────────────────────────────────┐
│ 2. For each [Benchmark] method:                                 │
│    a. Warmup phase (JIT compilation, cache warming)             │
│    b. Pilot phase (determine optimal iteration count)           │
│    c. Actual measurements (multiple iterations)                 │
│    d. Results aggregation                                       │
└─────────────────────────────────────────────────────────────────┘
                              ▼
┌─────────────────────────────────────────────────────────────────┐
│ 3. Report Generation                                            │
│    - Console output with statistics                             │
│    - Optional: HTML, CSV, JSON exports                          │
└─────────────────────────────────────────────────────────────────┘

Read the full file on GitHub · 382 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 382 lines · 38 tokens per session scan A 989f935cc8e6

Subscribe to this mod's changes

surface-benchmarker is a skill published in the GitHub repository mitchdenny/hex1b (173 stars, last pushed 6d ago), licensed MIT. It adds 38 tokens to every session and 2,970 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

write-documentation

Write and format Rust documentation correctly. Apply proactively when writing code with rustdoc comments (//! or ///). Covers voice & tone, prose style (opening lines, explicit subjects, verb tense), structure (inverted pyramid), intra-doc links (crate:: paths, reference-style), constant conventions (binary/byte…

r3bl-org/r3bl-open-core · 0 tokens

organize-modules

Apply private modules with public re-exports (barrel export) pattern for clean API design. Includes conditional visibility for docs and tests. Use when creating modules, organizing mod.rs files, or before creating commits.

r3bl-org/r3bl-open-core · 47 tokens

check-bounds-safety

Apply type-safe bounds checking patterns using VPIndex/VPLength types instead of usize. Use when working with arrays, buffers, cursors, viewports, or any code that handles indices and lengths.

r3bl-org/r3bl-open-core · 46 tokens

check-code-quality

Run comprehensive Rust code quality checks including compilation, linting, documentation, and tests. Use after completing code changes and before creating commits.

r3bl-org/r3bl-open-core · 31 tokens

analyze-performance

Establish performance baselines and detect regressions using flamegraph analysis. Use when optimizing performance-critical code, investigating performance issues, or before creating commits with performance-sensitive changes.

r3bl-org/r3bl-open-core · 38 tokens

run-clippy

Run clippy linting, enforce comment punctuation rules, format code with cargo fmt, and verify module organization patterns. Use after code changes and before creating commits.

r3bl-org/r3bl-open-core · 36 tokens