eval-memory
01Skill Claude CodeCodex
Autonomously check the health of the personal-knowledge memory system AND validate whether the autolearn improvements are actually working (not just active). Runs a blind LLM-as-judge efficacy eval against a frozen probe set, scores re-rank vs raw retrieval, tracks the trend in a ledger, and judges how well the…