Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add hajibabaie/combinatorial-optimization-skills --skill pandas-experiment-managementgit clone --depth 1 https://github.com/hajibabaie/combinatorial-optimization-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/hajibabaie/combinatorial-optimization-skills/pandas-experiment-management)<a href="https://agentmods.dev/skills/hajibabaie/combinatorial-optimization-skills/pandas-experiment-management"><img src="https://agentmods.dev/badge/skills/hajibabaie/combinatorial-optimization-skills/pandas-experiment-management/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/hajibabaie/combinatorial-optimization-skills/pandas-experiment-management"><img src="https://agentmods.dev/badge/skills/hajibabaie/combinatorial-optimization-skills/pandas-experiment-management.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00140 | $0.09174 |
| Opus 5 | $0.00070 | $0.04587 |
| Sonnet 5 | $0.00028 | $0.01835 |
| Haiku 4.5 | $0.00014 | $0.00917 |
Grade A, and why
pandas-experiment-management scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Runs shell commandslowCapability
Expected in a hook, worth knowing in a rule or an instructions file.
`subprocess.run(["git", "rev-parse", "--short", "HEAD"], capture_output=True, text=True)`, the How it starts
The opening of the file, as written. The whole thing — 800 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Pandas Experiment Management
You are an expert in managing computational-experiment data for combinatorial optimization research. This skill covers tidy result tables (one row per run), run-metadata capture, atomic CSV/parquet writing, aggregation across instances and seeds, and pivot tables ready for papers. Use the pattern catalog below to build a results pipeline that survives crashes, parallel workers, and reviewer questions — and that turns thousands of raw runs into one table you can defend.
Initial Assessment
Establish these facts before recommending a results pipeline:
- Campaign size. Count expected rows: instances × algorithms × configurations × seeds. A 3-algorithm, 30-instance, 10-seed study is 900 rows (CSV is fine); a tuning campaign with 500 configurations is 150,000 rows (parquet, partitioning).
- Run cost. Seconds per run or hours per run? Expensive runs make crash-safe writing and resume logic mandatory, not optional.
- Parallelism. Single process,
multiprocessingpool, or cluster array jobs writing to a shared filesystem? This decides the write strategy (one file per run vs. one shared file). - What is recorded per run. Final objective only, or also the incumbent trace over time? Traces need their own table (long format), never list-valued cells in the runs table.
- Optimization sense. Minimization or maximization? Mixed across problems? Store the raw
objective plus a
sensecolumn; convert only at aggregation time. - Reference values. Are best-known solutions (BKS) available for gap computation, or is the reference the best value found inside the campaign itself?
- Failure modes. Can runs time out, crash, or end infeasible? The schema needs a
statuscolumn from day one; retrofitting it later contaminates every aggregate already computed. - Downstream consumers. Statistical tests, convergence plots, LaTeX tables for a paper, or all three? The tidy runs table must serve all of them without re-running experiments.
- Storage stack. Is
pyarrowavailable for parquet? Is the filesystem local or networked (atomicity ofos.replaceholds within one filesystem only)? - Existing data. Is there a legacy spreadsheet or ad-hoc CSV to migrate? Migrate once, into the schema below, and freeze the old files as read-only.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 800 lines · 140 tokens per session scan A 93b9fa89a058
pandas-experiment-management is a skill published in the GitHub repository hajibabaie/combinatorial-optimization-skills (7 stars, last pushed 3mo ago), licensed MIT. It adds 140 tokens to every session and 9,174 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 1 finding (runs shell commands). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
phx-deps-audit
Audit Hex deps for supply-chain security risk — bidi chars, compile-time exec, maintainer changes, typosquats, CVEs. Use after mix deps.update, when checking if a package upgrade is safe, or reviewing mix.lock PR diffs.
release
CONTRIBUTOR TOOL - Cut a plugin release: bump plugin.json version, finalize CHANGELOG, update README if needed, gate on make ci, commit, tag vX.Y.Z, and create the GitHub release. Use when shipping a new plugin version. NOT distributed.
session-deep-dive
Deep qualitative analysis of high-signal sessions. Spawns subagents with v2 template, synthesizes patterns, compares against known findings. Use after /session-scan.
catchup
Summarize and review what changed while you were away. Use after a weekend, vacation, or flight to check missed PRs, git commits, Linear tickets, and meetings — one prioritized brief, not a firehose.
brainstorm
Brainstorm Elixir/Phoenix features — explore ideas, compare approaches, gather requirements. Use when vague idea, not sure how to approach, or want to discuss before plan.
learn-from-fix
Capture Elixir/Ecto/LiveView lessons and Hex API rules. Use after corrections or when asked to document learning, record a lesson, prevent a fixed mistake, or remember package guidance with --library.