Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/putervision/spc/vision-memory-mcpnpx skills add putervision/spc --skill vision-memory-mcpgit clone --depth 1 https://github.com/putervision/spcWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/putervision/spc/vision-memory-mcp)<a href="https://agentmods.dev/skills/putervision/spc/vision-memory-mcp"><img src="https://agentmods.dev/badge/skills/putervision/spc/vision-memory-mcp.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00034 | $0.01401 |
| Opus 5 | $0.00017 | $0.00700 |
| Sonnet 5 | $0.00007 | $0.00280 |
| Haiku 4.5 | $0.00003 | $0.00140 |
Grade B, and why
vision-memory-mcp scanned grade B with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Reads agent configuration directoriesmediumAgent snooping
.claude/, .codex/, .gemini/ hold keys, settings and other credentials a mod has no legitimate need for.
* **Google Antigravity (`~/.gemini/config/config.json`)**: Add these rules to your `"globalPermissionGrants"` -> `"allow"` list: This is a copy
100% identical to vision-memory-mcp — 0 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.
How it starts
The opening of the file, as written. The whole thing — 58 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Visual Memory (vision-memory-mcp)
This project utilizes vision-memory-mcp to cache visual states, record layout transitions, and avoid repetitive LLM vision calls.
1. Priority Order & Checklist
Whenever you capture a screenshot, examine a webpage, or need to verify a visual state, you MUST run this sequence:
- Orient: Call
get_session_contextto load active transitions and recent visual states. - Search (Optional): Call
recall_memoryto search past visual states by text query or image query. - Ingest/Verify: Call
analyze_screenshotwith the base64 screenshot before calling any vision LLMs. - Cache Hit: If
is_knownistrue, read the returneddescriptionand do NOT call your vision LLM. - Cache Miss: If
is_knownisfalse, inspect the image with your vision model, summarize the layout, and register it back by callinganalyze_screenshotwith both thescreenshotanddescriptionparameters. - Log Transitions: Right after taking any UI action (click, type, navigate, scroll), call
record_outcometo build the navigation graph. - Snapshotting: Call
manage_snapshot(action: "save") when reaching milestones, andmanage_snapshot(action: "diff") to check for visual regressions.
2. Complete Tool Reference
| Tool Name | Key Inputs | Description |
|---|---|---|
analyze_screenshot |
screenshot? (base64), file_path?, description?, items? |
Main ingestion (single or batch) and visual state retrieval tool. |
recall_memory |
query?, screenshot?, file_path?, strategy?, limit? |
Search visual memory by text query or image query (read-only). |
record_outcome |
from_state_id, to_state_id?, action, action_type? ('blocker' | 'click' | etc.) |
Record UI action transitions or log visual blockers for state-memory. |
get_navigation_paths |
from_state_id?, to_state_id?, to_description?, max_hops? |
Find historical path or instructions between states. |
predict_next_action |
current_state_id, goal_description?, goal_state_id? |
Predict best next UI action and grounded element handles (target_selector, target_coords). |
compare_states |
state_a_id & state_b_id OR video_a_id & video_b_id |
Compare two states visually (has_layout_change) or compare video runs. |
get_session_context |
include_recent?, include_frequent? |
Get recent/frequent states, transition graphs, disk stats, cache metrics, and version info. |
manage_snapshot |
action ('save' | 'diff' | 'export' | 'restore'), name?, archive_json? |
Unified snapshot management for visual checkpoints and regression detection. |
manage_visual_spec |
action ('set' | 'verify' | 'list'), name?, screenshot?, tolerance? |
Register and verify visual design contract baselines (Visual SDD). |
manage_video |
action ('ingest' | 'search' | 'timeline'), file_path?, query?, video_id? |
Ingest WebM/MP4 recordings, search video keyframes, or retrieve timelines. |
create_evidence_pack |
keyframe_state_ids, source_video_id?, linked_state_memory_nodes? |
Package immutable evidence packs linking video keyframes to state-memory DAGs. |
export_trajectories |
format? ('json' | 'llava' | 'qwen2_vl' | 'joint'), trace_id? |
Export multimodal trajectories for model fine-tuning or joint workflow exports. |
undo_visual_mutation |
type? ('state' | 'transition' | 'any') |
Revert the last visual state ingestion or transition edge addition. |
forget_state |
state_id |
Purge a specific state and vector embedding for privacy. |
wait_for_visual_state |
target_state_id, timeout_ms? |
Poll for target visual state until present or timeout occurs. |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 58 lines · 34 tokens per session scan B 2babda58358c
vision-memory-mcp is a skill published in the GitHub repository putervision/spc (39 stars, last pushed 13d ago), licensed MIT. It adds 34 tokens to every session and 1,401 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it B with 1 finding (reads agent configuration directories). It is 100% identical to vision-memory-mcp, differing in 0 lines, and is treated as a copy.
Other skills, from other repositories
claw-roam
Sync OpenClaw workspace between multiple machines (local Mac and remote VPS) via Git. Enables seamless migration of OpenClaw personality, memory, and skills. Use when user wants to (1) push workspace changes to remote before shutdown, (2) pull latest workspace on a new machine, (3) check sync status between machines…
session-watchdog
Monitor session context levels and proactively save checkpoints before compaction. Use when: (1) session context exceeds 80% capacity, (2) user asks about session status or memory, (3) at the start of each new session to check context, or (4) before long tasks that might push context over threshold.
profiling
Profile Rubydex indexer performance — CPU flamegraphs, memory usage, phase-level timing. Use this skill whenever the user mentions profiling, performance, flamegraphs, benchmarking, "why is X slow", bottlenecks, hot paths, memory usage, or wants to understand where time is spent during indexing/resolution. Also…
extract-knowledge
Extract knowledge (decisions, facts, session metadata) from the current Claude Code session into the Grafema Knowledge Base. Run after completing a task or at any point when substantive knowledge was produced. Follows runbook ai/runbooks/02-claude-sessions.md.
mem0-oss-to-platform
Plan and then execute a migration of a project from the mem0 open-source / self-hosted SDK (the local Memory class) to the mem0 Platform / hosted / managed SDK (the MemoryClient class). Use this whenever a developer wants to move, switch, or migrate their mem0 usage off OSS/self-hosted to the hosted API — e.g.…
Cortex
Operate Cortex, the LifeOS memory system — the typed Knowledge Archive (People, Companies, Ideas, Research with typed related: links) plus recall of prior work sessions, ISAs, and conversations. Search, add, harvest, develop, ingest, distill, graph-navigate, recall. USE WHEN cortex, knowledge, knowledge base, search…