Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add 001TMF/harness-forge --skill meta-harness-proteusgit clone --depth 1 https://github.com/001TMF/harness-forgeWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/001tmf/harness-forge/meta-harness-proteus)<a href="https://agentmods.dev/skills/001tmf/harness-forge/meta-harness-proteus"><img src="https://agentmods.dev/badge/skills/001tmf/harness-forge/meta-harness-proteus/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/001tmf/harness-forge/meta-harness-proteus"><img src="https://agentmods.dev/badge/skills/001tmf/harness-forge/meta-harness-proteus.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00024 | $0.00822 |
| Opus 5 | $0.00012 | $0.00411 |
| Sonnet 5 | $0.00005 | $0.00164 |
| Haiku 4.5 | $0.00002 | $0.00082 |
Grade A, and why
meta-harness-proteus scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
Copies of this mod
1 near-identical copy found in the catalogue:
- meta-harness-proteus — 100% identical, 0 lines differ
How it starts
The opening of the file, as written. The whole thing — 78 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Meta-Harness — proteus memory-summary evolution
Run ONE iteration. Do all work in the main session — do NOT delegate to subagents.
You do NOT run benchmarks. You analyze prior results, prototype a mechanism,
and write new candidate summary compressors. The outer loop (meta_harness.py)
scores them on (fidelity, chars) separately, with no model and no network.
What a candidate is
A summary compressor: it turns one campaign-memory record (a dict — see
corpus.py) into the short string injected into the policy's context on
retrieval. The proteus analog of a memory system. The grading is in
corpus.py::score_fidelity: the fraction of load-bearing facts (target,
surface, strategy, outcome, quality, difficulty, transfer hint) that survive in
your summary. Context cost = len(summary).
The objective
Preserve fidelity (>= the floor in config.yaml, currently 0.70 worst-record)
while using FEWER characters than agents/baseline_incumbent.py. The frontier
is Pareto: fidelity up, chars down. You cannot win by dropping facts — a summary
that loses a required fact loses fidelity and falls off the frontier.
CRITICAL CONSTRAINTS
- Implement exactly 3 new compressors this iteration.
- Each must change a mechanism, not a constant. Bad: "same template, drop the organism." Good ideas: abbreviation/symbol encoding of fixed vocab (surface types, outcomes); a key:value micro-syntax instead of prose; dropping only provably-redundant words; reordering so the highest-value facts survive truncation; field-name elision where the value is self-identifying.
- No record-specific hints. Never hardcode a target name, campaign_id, or
any value from
corpus.pyinto a compressor. It must generalize to unseen records. (This is the anti-leakage rule — load-bearing for proteus.) - Do not abort early or write "the frontier is optimal".
Workflow
- Analyze. Read
logs/evolution_summary.jsonl(what's been tried),logs/frontier.json(current best),corpus.py(records + rubric),agents/baseline_incumbent.py(the system to beat). - Prototype (mandatory). Write a throwaway script in
/tmp/that runs your compression idea over a couple ofcorpus.pyrecords and checks fidelity by eye before committing. Delete it after. - Implement. For each of 3 candidates: copy
agents/baseline_incumbent.pytoagents/<snake_name>.py, subclassSummaryCompressor, implementsummarize(self, record) -> str. Import fromcandidate_base. Self-critique: is this a new mechanism or just a tweaked constant? If the latter, rewrite. - Validate.
python -c "import agents.<name>; print('OK')"from the repo root. - Write
logs/pending_eval.json:
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 78 lines · 24 tokens per session scan A 87f42a3d44d0
meta-harness-proteus is a skill published in the GitHub repository 001TMF/harness-forge (78 stars, last pushed 2mo ago), licensed MIT. It adds 24 tokens to every session and 822 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
knowledge_store_skill
Skill for working with local .knowledge.yaml files via KnowledgeStore. Use this when you need to recall, search, or manage directory-local memories and knowledge links stored in plain YAML alongside the user's project files. KnowledgeStore is directory-scoped. Each directory that contains a .knowledge.yaml file…
cocoscout
Relevance-ranked context loading — Tier 2 async subagent (Haiku, <5s) that fires after Tier 1 deterministic checks in UserPromptSubmit. Injects ranked context from CocoGrove, CocoContext, Environment Inspector, Prompt Studio, and CocoDream.
cocohealth
Context utilization monitor — background monitor that samples context window utilization via PostToolUse hook, surfaces advisory at 60% and critical warning with recovery decision matrix at 70%.
pull-search
CocoPull session archive search — full-text search across past sessions. Handles $pull search " " with --since and --feature filters, and $pull index rebuild.
pod-kb
Display the project knowledge base (lifecycle/kb.md) — project-specific patterns, decisions, and gotchas accumulated by CocoCupper across sessions.
pod-resume
Reconstruct context for a returning developer. Reads .cocoplus/ state and generates a narrative summary of project status, last achievements, pending tasks, recent decisions, available patterns, CocoCupper insights, and recommended next action.