Borrowing it
Nothing to install: this file belongs to on1659/memradar. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/on1659/memradar/master/.claude/skills/memtest/SKILL.mdgit clone --depth 1 https://github.com/on1659/memradarWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/on1659/memradar/memtest)<a href="https://agentmods.dev/skills/on1659/memradar/memtest"><img src="https://agentmods.dev/badge/skills/on1659/memradar/memtest.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00097 | $0.00725 |
| Opus 5 | $0.00048 | $0.00362 |
| Sonnet 5 | $0.00019 | $0.00145 |
| Haiku 4.5 | $0.00010 | $0.00072 |
Grade A, and why
memtest scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
memtest — AI 역할 평가 실행 스킬
즉시 실행 (Claude에게 지시)
이 스킬이 호출되면 반드시 아래 단계를 순서대로 실행해라:
Step 1: 테스트 스크립트 실행
Bash 도구로 다음 명령을 실행:
cd /Users/radar/Work/memradar && npx tsx scripts/test-eval-and-report.mts
Step 2: 결과 확인
스크립트 출력에서 정확도 수치를 확인하고 사용자에게 요약 보고:
- 전체 정확도
- 카테고리별 정확도 (pure/mixed/ambiguous/consistency)
- 난이도별 정확도 (easy/normal/hard)
Step 3: HTML 리포트 열기
Bash 도구로 리포트 파일을 브라우저에서 열기:
open /Users/radar/Work/memradar/docs/eval-report.html
Step 4: 사용자에게 최종 보고
다음 형식으로 짧게 보고:
✅ 테스트 완료
📊 정확도: X.X% (N/총합)
📄 리포트: docs/eval-report.html (브라우저에서 열림)
주요 발견:
- [카테고리/난이도별 간단 요약]
- [개선 포인트 1-2개]
관련 파일
- 실행 스크립트:
scripts/test-eval-and-report.mts - 분류 로직:
src/lib/usageProfile.ts - 샘플 디렉토리:
tests/fixtures/role-eval-samples/(218개) - 출력 HTML:
docs/eval-report.html - 출력 JSON:
docs/eval-results.json - 출력 마크다운:
docs/AI-ROLE-EVAL-RESULTS.md
에러 처리
- 샘플 파일 없음:
tests/fixtures/role-eval-samples/디렉토리 확인 지시 - tsx 에러:
scripts/test-eval-and-report.mts파일 존재 확인 - HTML 생성 실패: 에러 메시지 그대로 보고
중요 제약
- 이 스킬은 스크립트를 실행하는 역할만 한다
- 코드를 수정하지 마라 (읽기 전용)
analyzeUsageTopCategories()로직을 변경하지 마라- 샘플 JSON 파일을 수정하지 마라
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 86 lines · 97 tokens per session scan A 7e9cecd583c7
memtest is a skill published in the GitHub repository on1659/memradar (11 stars, last pushed 6d ago), licensed MIT. It adds 97 tokens to every session and 725 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-02.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
local-ai-agents
Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. Covers Small Language Models (SLMs), the OpenAI-compatible local endpoint, sandboxed local tools, local RAG with Chroma, local MCP servers, hybrid cloud/local routing, and the…
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…
next-partial-prefetching-adoption
Turn on Partial Prefetching in a Next.js app and work through the insights it surfaces. Use when the user wants to enable or adopt Partial Prefetching, flip the partialPrefetching flag, opt routes in with export const prefetch = 'partial', audit Link prefetch={true} behavior, preserve existing prefetched UI with…
chronicle
Analyze Copilot session history for standup reports, usage tips, session search, and session reindexing. Use when the user asks for a standup, daily summary, usage tips, workflow recommendations, wants to search or find past sessions by keyword/file/PR, wants to reindex their session store, or asks about deleting…
babysit-pr
Babysit a GitHub pull request after creation by continuously polling review comments, CI checks/workflow runs, and mergeability state until the PR is merged/closed or user help is required. Diagnose failures, retry likely flaky failures up to 3 times, auto-fix/push branch-related issues when appropriate, and keep…