Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/openuploading/cognifold/longmemeval-iteratenpx skills add OpenUploading/CogniFold --skill longmemeval-iterategit clone --depth 1 https://github.com/OpenUploading/CogniFoldWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/openuploading/cognifold/longmemeval-iterate)<a href="https://agentmods.dev/skills/openuploading/cognifold/longmemeval-iterate"><img src="https://agentmods.dev/badge/skills/openuploading/cognifold/longmemeval-iterate.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00122 | $0.01262 |
| Opus 5 | $0.00061 | $0.00631 |
| Sonnet 5 | $0.00024 | $0.00252 |
| Haiku 4.5 | $0.00012 | $0.00126 |
Grade A, and why
longmemeval-iterate scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 104 lines — stays where its author put it; the contents beside it link to each section on GitHub.
LongMemEval Autonomous Iteration
When to use
- User says "iterate LongMemEval" / "run the longmemeval loop" / "continue R10"
- After a fresh clone, before any iteration: walk §0 setup
- Any time the autonomous loop is mid-cycle and needs to resume
Hard rules (never violate)
- Branch lock: only commit/push on
longmemeval-iter. Verifygit branch --show-currentreturnslongmemeval-iterbefore anygit commit. Never touchmain/iter/public-release/ etc. - Judge lock:
--judge-model openai:gpt-4oalways. Substituting breaks comparability with Mastra / Hindsight numbers. - Symbolic stack on:
--symbolic-resolver --symbolic-temporal --symbolic-bypassmust all stay enabled (~5 pp on the score). - Full N=500 each round: no stratified < 133, no sampled subsets. Resume makes incremental cost ≈ wall-clock of one batch anyway.
- Cluster-then-diagnose-then-propose: every fix must follow the
protocol in
references/iteration-rules.md. Skipping this step is the #1 historical cause of regressions.
Setup (one-time per fresh machine)
Run scripts/check_setup.sh — it verifies branch, push credentials,
remote, model config in scripts/parallel_longmemeval.sh, and that
history_max_effort.md + .max_effort_round exist (creates them if
not). Halt and surface any failures.
The loop
loop forever:
ROUND = read+bump .max_effort_round
# (1) Baseline: full N=500 run
bash scripts/parallel_longmemeval.sh <N_PARALLEL> 133 500
# N_PARALLEL from references/model-config.md Tier table
# (2) Measure
metrics = json.load("benchmarks/longmemeval/output/metrics.json")
correct = metrics["correct"]
# (3) Terminate if ≥475 AND confirmation rerun also ≥475
if correct >= 475:
run confirmation rerun (rm hypothesis.jsonl, re-run full N)
if confirmed correct2 >= 475:
commit FINAL + push + EXIT
# else fall through with corrected (lower) baseline
# (4) Snapshot pre-fix state
cp -r output/ output_v${ROUND}/
append baseline metric to history_max_effort.md
git add + commit + git push origin longmemeval-iter
# (5) Analyze failures per references/iteration-rules.md §A-B-C
# (6) Propose fix (estimate trigger isolation; not a gate)
# (7) Drop test_set qids, re-run with SAME N_PARALLEL (see scripts/drop_qids.py)
# (8) Compute net = fixes - regressions vs output_v${ROUND}/
# (9) Apply references/iteration-rules.md decision table:
# net ≥ +1 → keep
# net ∈ {0, -1} + reusable → keep (infra)
# net ≤ -2 → revert (restore verdicts + git revert)
# (10) Commit + push the post-fix state
# Loop back to (1)
What ships with it
7 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 104 lines · 122 tokens per session scan A 465c2ba438cf
longmemeval-iterate is a skill published in the GitHub repository OpenUploading/CogniFold (59 stars, last pushed 10d ago), licensed Apache-2.0. It adds 122 tokens to every session and 1,262 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
mem0-oss-to-platform
Plan and then execute a migration of a project from the mem0 open-source / self-hosted SDK (the local Memory class) to the mem0 Platform / hosted / managed SDK (the MemoryClient class). Use this whenever a developer wants to move, switch, or migrate their mem0 usage off OSS/self-hosted to the hosted API — e.g.…
Cortex
Operate Cortex, the LifeOS memory system — the typed Knowledge Archive (People, Companies, Ideas, Research with typed related: links) plus recall of prior work sessions, ISAs, and conversations. Search, add, harvest, develop, ingest, distill, graph-navigate, recall. USE WHEN cortex, knowledge, knowledge base, search…
auditing-subgroup-fairness
Audit an OpenMed NER or de-identification model for performance disparities across demographic subgroups (sex, age band, race/ethnicity when available) using openmed.eval.fairnessreport. Use when the user wants per-subgroup recall and leakage, wants to check whether de-identification under-protects a group, wants to…
agent-memory
../../../engineering/agent-memory/skills/agent-memory/SKILL.md.
memory
Use when the user asks to remember, recall, forget, update, search, or inspect durable OpenSquilla memory, including profile facts in USER.md and long-term notes in MEMORY.md or memory//.md.
ha-data-stores
Map of Hope Agent's local data stores and safe read-only query workflow. Use when the user asks where Hope Agent stores data, wants to inspect sessions/messages/memory/logs/background jobs/knowledge indexes/settings, asks the model to query local app data, or debugging requires checking persisted state. Trigger…