Skill Claude CodeCodex
Autonomous LongMemEval benchmark iteration toward ≥95% strict J-Score on full N=500. Use when the user asks to run, iterate, improve, or continue the LongMemEval campaign on branch longmemeval-iter. Loops baseline → cluster-analyze → propose fix → re-test → net-positive decision → commit → repeat. Terminates only when…