Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/mitchdenny/hex1b/surface-benchmarkernpx skills add mitchdenny/hex1b --skill surface-benchmarkergit clone --depth 1 https://github.com/mitchdenny/hex1bWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/mitchdenny/hex1b/surface-benchmarker)<a href="https://agentmods.dev/skills/mitchdenny/hex1b/surface-benchmarker"><img src="https://agentmods.dev/badge/skills/mitchdenny/hex1b/surface-benchmarker.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00038 | $0.02970 |
| Opus 5 | $0.00019 | $0.01485 |
| Sonnet 5 | $0.00008 | $0.00594 |
| Haiku 4.5 | $0.00004 | $0.00297 |
Grade A, and why
surface-benchmarker scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 382 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Surface Benchmarker Skill
This skill provides guidelines for AI agents to run and interpret performance benchmarks for the Hex1b Surface API. Run benchmarks whenever you modify code in src/Hex1b/Surfaces/ to ensure performance is not regressed.
Quick Reference
| Action | Command |
|---|---|
| Run all Surface benchmarks | dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "Surface*" |
| Run specific benchmark | dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "*WriteText*" |
| Quick dry-run (sanity check) | dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "Surface*" --job dry |
When to Run Benchmarks
ALWAYS run benchmarks after modifying:
| File/Area | Critical Benchmarks |
|---|---|
SurfaceCell.cs |
All (cells are used everywhere) |
Surface.cs |
CreateSurface_*, WriteText_*, Fill_*, Clone_* |
CompositeSurface.cs |
CompositeSurface_* |
SurfaceComparer.cs |
Compare_*, ToTokens_*, ToAnsiString_* |
SurfaceDiff.cs |
Compare_* |
ComputeContext.cs |
CompositeSurface_* (computed cells use this) |
Benchmark Harness Architecture
Project Structure
benchmarks/Hex1b.Benchmarks/
├── Hex1b.Benchmarks.csproj # Console app with BenchmarkDotNet
├── Program.cs # Entry point
└── SurfaceBenchmarks.cs # Surface API benchmarks
How BenchmarkDotNet Works
BenchmarkDotNet is a .NET performance benchmarking library that:
- Warms up the JIT by running the benchmark multiple times before measuring
- Runs multiple iterations to get statistically significant results
- Reports mean, median, standard deviation for each benchmark
- Tracks memory allocations (when
[MemoryDiagnoser]is enabled)
Benchmark Lifecycle
┌─────────────────────────────────────────────────────────────────┐
│ 1. GlobalSetup │
│ - Runs once before all benchmarks │
│ - Creates pre-allocated surfaces, diffs, etc. │
│ - Sets up shared test data │
└─────────────────────────────────────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────────────┐
│ 2. For each [Benchmark] method: │
│ a. Warmup phase (JIT compilation, cache warming) │
│ b. Pilot phase (determine optimal iteration count) │
│ c. Actual measurements (multiple iterations) │
│ d. Results aggregation │
└─────────────────────────────────────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────────────┐
│ 3. Report Generation │
│ - Console output with statistics │
│ - Optional: HTML, CSV, JSON exports │
└─────────────────────────────────────────────────────────────────┘
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 382 lines · 38 tokens per session scan A 989f935cc8e6
surface-benchmarker is a skill published in the GitHub repository mitchdenny/hex1b (173 stars, last pushed 6d ago), licensed MIT. It adds 38 tokens to every session and 2,970 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
write-documentation
Write and format Rust documentation correctly. Apply proactively when writing code with rustdoc comments (//! or ///). Covers voice & tone, prose style (opening lines, explicit subjects, verb tense), structure (inverted pyramid), intra-doc links (crate:: paths, reference-style), constant conventions (binary/byte…
organize-modules
Apply private modules with public re-exports (barrel export) pattern for clean API design. Includes conditional visibility for docs and tests. Use when creating modules, organizing mod.rs files, or before creating commits.
check-bounds-safety
Apply type-safe bounds checking patterns using VPIndex/VPLength types instead of usize. Use when working with arrays, buffers, cursors, viewports, or any code that handles indices and lengths.
check-code-quality
Run comprehensive Rust code quality checks including compilation, linting, documentation, and tests. Use after completing code changes and before creating commits.
analyze-performance
Establish performance baselines and detect regressions using flamegraph analysis. Use when optimizing performance-critical code, investigating performance issues, or before creating commits with performance-sensitive changes.
run-clippy
Run clippy linting, enforce comment punctuation rules, format code with cargo fmt, and verify module organization patterns. Use after code changes and before creating commits.