longmemeval-iterate

longmemeval-iterate is a skill for Claude Code, Codex from OpenUploading/CogniFold. It costs 122 tokens per session (1,262 once invoked), scanned A, original, Apache-2.0.

An automated improvement loop for LongMemEval, a benchmark that tests how well an AI system uses information from long conversations. It repeatedly measures, diagnoses, changes, and retests a system on 500 examples.

In plain words
What is it for?
It runs baselines, groups and analyzes failures, proposes fixes, retests them, decides whether they help, and commits accepted changes on the dedicated iteration branch.
Why use it?
It organizes benchmark improvement into comparable rounds and prevents changes from being accepted without checking whether they improve the score.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/openuploading/cognifold/longmemeval-iterate
Any agent
npx skills add OpenUploading/CogniFold --skill longmemeval-iterate
Clone the repo
git clone --depth 1 https://github.com/OpenUploading/CogniFold

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for longmemeval-iterate

README.md
[![agentmods](https://agentmods.dev/badge/skills/openuploading/cognifold/longmemeval-iterate.svg)](https://agentmods.dev/skills/openuploading/cognifold/longmemeval-iterate)
Your own site
<a href="https://agentmods.dev/skills/openuploading/cognifold/longmemeval-iterate"><img src="https://agentmods.dev/badge/skills/openuploading/cognifold/longmemeval-iterate.svg" alt="Measured on agentmods" height="20"></a>
Per session 122 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,262 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00122 $0.01262
Opus 5 $0.00061 $0.00631
Sonnet 5 $0.00024 $0.00252
Haiku 4.5 $0.00012 $0.00126

Measured 4d ago against content hash 465c2ba438cf, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

longmemeval-iterate scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

The scan reads SKILL.md. This mod also ships 4 executable files (scripts/check_setup.sh, scripts/compute_net.py, scripts/drop_qids.py, …), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/skills/longmemeval-iterate/SKILL.md · 104 lines

How it starts

The opening of the file, as written. The whole thing — 104 lines — stays where its author put it; the contents beside it link to each section on GitHub.

LongMemEval Autonomous Iteration

When to use

  • User says "iterate LongMemEval" / "run the longmemeval loop" / "continue R10"
  • After a fresh clone, before any iteration: walk §0 setup
  • Any time the autonomous loop is mid-cycle and needs to resume

Hard rules (never violate)

  1. Branch lock: only commit/push on longmemeval-iter. Verify git branch --show-current returns longmemeval-iter before any git commit. Never touch main / iter / public-release / etc.
  2. Judge lock: --judge-model openai:gpt-4o always. Substituting breaks comparability with Mastra / Hindsight numbers.
  3. Symbolic stack on: --symbolic-resolver --symbolic-temporal --symbolic-bypass must all stay enabled (~5 pp on the score).
  4. Full N=500 each round: no stratified < 133, no sampled subsets. Resume makes incremental cost ≈ wall-clock of one batch anyway.
  5. Cluster-then-diagnose-then-propose: every fix must follow the protocol in references/iteration-rules.md. Skipping this step is the #1 historical cause of regressions.

Setup (one-time per fresh machine)

Run scripts/check_setup.sh — it verifies branch, push credentials, remote, model config in scripts/parallel_longmemeval.sh, and that history_max_effort.md + .max_effort_round exist (creates them if not). Halt and surface any failures.

The loop

loop forever:
    ROUND = read+bump .max_effort_round

    # (1) Baseline: full N=500 run
    bash scripts/parallel_longmemeval.sh <N_PARALLEL> 133 500
    # N_PARALLEL from references/model-config.md Tier table

    # (2) Measure
    metrics = json.load("benchmarks/longmemeval/output/metrics.json")
    correct = metrics["correct"]

    # (3) Terminate if ≥475 AND confirmation rerun also ≥475
    if correct >= 475:
        run confirmation rerun (rm hypothesis.jsonl, re-run full N)
        if confirmed correct2 >= 475:
            commit FINAL + push + EXIT
        # else fall through with corrected (lower) baseline

    # (4) Snapshot pre-fix state
    cp -r output/ output_v${ROUND}/
    append baseline metric to history_max_effort.md
    git add + commit + git push origin longmemeval-iter

    # (5) Analyze failures per references/iteration-rules.md §A-B-C
    # (6) Propose fix (estimate trigger isolation; not a gate)
    # (7) Drop test_set qids, re-run with SAME N_PARALLEL (see scripts/drop_qids.py)
    # (8) Compute net = fixes - regressions vs output_v${ROUND}/
    # (9) Apply references/iteration-rules.md decision table:
    #     net ≥ +1                   → keep
    #     net ∈ {0, -1} + reusable   → keep (infra)
    #     net ≤ -2                   → revert (restore verdicts + git revert)
    # (10) Commit + push the post-fix state
    # Loop back to (1)

Read the full file on GitHub · 104 lines

Files

What ships with it

7 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 104 lines · 122 tokens per session scan A 465c2ba438cf

Subscribe to this mod's changes

longmemeval-iterate is a skill published in the GitHub repository OpenUploading/CogniFold (59 stars, last pushed 10d ago), licensed Apache-2.0. It adds 122 tokens to every session and 1,262 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

mem0-oss-to-platform

Plan and then execute a migration of a project from the mem0 open-source / self-hosted SDK (the local Memory class) to the mem0 Platform / hosted / managed SDK (the MemoryClient class). Use this whenever a developer wants to move, switch, or migrate their mem0 usage off OSS/self-hosted to the hosted API — e.g.…

mem0ai/mem0 · 273 tokens

Cortex

Operate Cortex, the LifeOS memory system — the typed Knowledge Archive (People, Companies, Ideas, Research with typed related: links) plus recall of prior work sessions, ISAs, and conversations. Search, add, harvest, develop, ingest, distill, graph-navigate, recall. USE WHEN cortex, knowledge, knowledge base, search…

danielmiessler/LifeOS · 196 tokens

auditing-subgroup-fairness

Audit an OpenMed NER or de-identification model for performance disparities across demographic subgroups (sex, age band, race/ethnicity when available) using openmed.eval.fairnessreport. Use when the user wants per-subgroup recall and leakage, wants to check whether de-identification under-protects a group, wants to…

maziyarpanahi/openmed · 148 tokens

agent-memory

../../../engineering/agent-memory/skills/agent-memory/SKILL.md.

alirezarezvani/claude-skills · 0 tokens

memory

Use when the user asks to remember, recall, forget, update, search, or inspect durable OpenSquilla memory, including profile facts in USER.md and long-term notes in MEMORY.md or memory//.md.

opensquilla/opensquilla · 44 tokens

ha-data-stores

Map of Hope Agent's local data stores and safe read-only query workflow. Use when the user asks where Hope Agent stores data, wants to inspect sessions/messages/memory/logs/background jobs/knowledge indexes/settings, asks the model to query local app data, or debugging requires checking persisted state. Trigger…

shiwenwen/hope-agent · 115 tokens