Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/thinkfleetai/memmesh/memmesh-test-integrationnpx skills add ThinkfleetAI/memmesh --skill memmesh-test-integrationgit clone --depth 1 https://github.com/ThinkfleetAI/memmeshWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/thinkfleetai/memmesh/memmesh-test-integration)<a href="https://agentmods.dev/skills/thinkfleetai/memmesh/memmesh-test-integration"><img src="https://agentmods.dev/badge/skills/thinkfleetai/memmesh/memmesh-test-integration.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00136 | $0.00682 |
| Opus 5 | $0.00068 | $0.00341 |
| Sonnet 5 | $0.00027 | $0.00136 |
| Haiku 4.5 | $0.00014 | $0.00068 |
Grade A, and why
memmesh-test-integration scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
memmesh-test-integration
Prove the integration actually works — not just that it typechecks.
Preconditions
- A
.memmesh-integration/directory exists (else tell the user to runmemmesh-integratefirst). - Credentials: Hosted ⇒
MEMMESH_API_KEYset; Local ⇒memmesh doctorpasses.
Steps
-
Static gate. Typecheck / lint / build. Any failure ⇒ stop, report.
-
Native suite. Run the repo's own tests twice — once with
MEMMESH_ENABLEDunset (must be byte-for-byte the original behavior) and once set. Both must pass. -
Real E2E smoke against the live engine (not a mock):
observea known fact for a throwawayuserId(e.g.smoke-<runid>).searchfor it; assert the fact comes back.buildContextfor the subject; assert the fact appears in the bundle.- If prediction was wired: call
predict/predictTargetand assert you get either a calibrated probability or an honest abstention — both are a pass; a crash or an uncalibrated 1.0/0.0 with no evidence is a fail. - Clean up:
memory_delete(orcompliance.hardDeleteSubject) the throwaway subject so the smoke run leaves no residue.
-
Scorecard. Write
.memmesh-integration/scorecard.md:Check Result Typecheck / build ✅ / ❌ Native tests (flag off) ✅ / ❌ Native tests (flag on) ✅ / ❌ E2E observe→search ✅ / ❌ E2E buildContext ✅ / ❌ E2E predict (calibrated OR abstained) ✅ / ❌ / n/a Smoke cleanup ✅ / ❌
Definition of done
Every applicable row green, throwaway data removed, and a one-paragraph verdict: ship / needs-work, with the failing checks called out.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 61 lines · 136 tokens per session scan A d7ca3ddf39e5
memmesh-test-integration is a skill published in the GitHub repository ThinkfleetAI/memmesh (441 stars, last pushed 10d ago), licensed Apache-2.0. It adds 136 tokens to every session and 682 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
use-agent-browser-for-airi
Test AIRI display-model imports with agent-browser across stage-tamagotchi Electron, stage-web, and stage-pocket mobile web layouts. Use when uploading and verifying contributor-supplied Live2D ZIP, VRM, or MMD ZIP/PMX/PMD files through AIRI's model selector, including onboarding bypass, format-specific import…
run-integration-tests
Build, pack, and run .NET MAUI integration tests locally. Validates templates, samples, and end-to-end scenarios using the local workload.
cli-e2e-testcase-writer
Use when adding or updating Go CLI E2E coverage for one tests/clie2e/{domain} domain of the compiled lark-cli, especially when the work requires live --help or schema exploration, scenario-based clie2e.RunCmd workflows, and per-domain coverage.md maintenance.
webapp-testing
Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
__SKILL_ID__
This fixture verifies that skill content can be written into the target sandbox and queried back immediately.
harness-test-writer
Add regression test cases to the Bifrost provider harness (the Postman collection run via make run-provider-harness-test) based on a merged PR or a GitHub issue. Fetches the PR/issue, traces the affected wire path in the codebase, checks existing harness coverage, designs cases following harness conventions, inserts…