factor-benchmark

A set of benchmark tests for comparing a factor-mining system with baseline methods and reproducing reported research results. An ablation test removes or changes one part of a system to measure its contribution.

In plain words
What is it for?
Use it for Top-K comparisons, memory and strategy ablations, rising transaction-cost tests, engine-efficiency profiling, or the complete benchmark suite.
Why use it?
It turns one factor-mining run into evidence about performance, memory, strategy choices, costs, runtime, and reproducibility.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/minihellboy/factorminer/factor-benchmark
Any agent
npx skills add minihellboy/factorminer --skill factor-benchmark
Clone the repo
git clone --depth 1 https://github.com/minihellboy/factorminer

Made for: Claude Code, Codex.

Per session 79 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 517 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00079 $0.00517
Opus 5 $0.00039 $0.00259
Sonnet 5 $0.00016 $0.00103
Haiku 4.5 $0.00008 $0.00052

Measured 2d ago against content hash c5b78a8da140, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

factor-benchmark scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

integrations/factor-researcher/plugin/skills/factor-benchmark/SKILL.md · 57 lines

What it actually says

Factor Benchmark

This skill runs FactorMiner's canonical benchmark surface — the rigorous comparison layer that turns a single run into evidence.

Modes

Mode What it answers
table1 Top-K freeze benchmark across configured universes vs. baselines — the headline reproduction.
ablation-memory How much does experience memory contribute?
ablation-strategy Effect of memory policy × dependence metric × backend.
cost-pressure How does the library hold up under rising transaction costs?
efficiency Operator- and factor-level runtime/compute cost.
suite The full benchmark suite in one run.

Workflow

Run a benchmark

factorminer -o output/bench benchmark table1 --data path/to/market_data.csv
factorminer -o output/bench benchmark suite --data path/to/market_data.csv

Pass a pre-mined library to benchmark a specific run rather than mining fresh:

factorminer -o output/bench benchmark table1 \
  --data market_data.csv \
  --factor-miner-library output/run1/factor_library.json

efficiency takes no data — it profiles the engine itself.

Read the result

The CLI prints a per-universe summary (library IC, ICIR, avg |ρ|) and writes JSON payloads into the output directory. Fold those JSON files into the research note with factor-report --benchmark.

Interpreting ablations

  • An ablation that removes a feature and barely moves the metric means that feature is not earning its compute on this dataset — report that plainly.
  • cost-pressure is the honesty check: a library that only wins at zero cost is not a result.

Guardrails

  • Benchmark numbers are comparative research evidence, not a performance guarantee.
  • Use the same dataset and splits across compared runs, or the comparison is meaningless.
  • Reproduction claims must cite the exact config and run directory.
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 57 lines · 79 tokens per session scan A c5b78a8da140

Subscribe to this mod's changes

factor-benchmark is a skill published in the GitHub repository minihellboy/factorminer (105 stars, last pushed 16d ago), licensed MIT. It adds 79 tokens to every session and 517 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

brainstorming

You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.

obra/superpowers · 37 tokens

chat-pet-sprite-creation

Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.

microsoft/vscode · 53 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens

agent-host-chat-contributions

Build and review cross-cutting agent-host chat behavior through lifecycle contributions. Use when adding turn lifecycle side effects, prompt or context injection, restored-history transformation, protocol-action observation, or when reviewing changes that add code to AgentSideEffects or AgentService.

microsoft/vscode · 56 tokens

auto-perf-optimize

Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.

microsoft/vscode · 62 tokens