dream-beta-test

dream-beta-test is a skill for Claude Code from Zenetusken/consolidate-memory. It costs 189 tokens per session (1,595 once invoked), scanned A, original, MIT.

A testing skill for the consolidate-memory skill. It runs automated checks and human-style review lenses to find defects, regressions, or unsupported claims in that memory-consolidation process.

In plain words
What is it for?
Use it to test new or changed versions of consolidate-memory, stress-test its behavior on a repository, and produce a structured report of confirmed and suspected problems.
Why use it?
It helps maintainers discover when memory consolidation stops preserving correct information or reports results too confidently.

Skill for Claude Code

Written for Claude Code: $CLAUDE_PLUGIN_ROOT variable.

Runs only inside its plugin — its command needs a path that Claude Code sets for a plugin’s own hooks and for nothing else. Install the plugin, not this.

Part of the dream-beta-tester plugin — 1 skill shipped together

Good fit Use it to test new or changed versions of consolidate-memory, stress-test its behavior on a repository, and produce a structured report of confirmed and suspected problems.

Compare 6 skills from other repositories ↓
Install

Getting it into your agent

This one installs as part of its plugin. Adding the marketplace and installing the plugin brings it with everything else the plugin ships.

Claude Code
/plugin marketplace add Zenetusken/consolidate-memory
Claude Code
/plugin install dream-beta-tester

Made for: Claude Code.

Or install dream-beta-tester, the plugin that ships this one along with the rest of its 1 skill.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for dream-beta-test

README.md
[![agentmods](https://agentmods.dev/badge/skills/zenetusken/consolidate-memory/dream-beta-test/github.svg)](https://agentmods.dev/skills/zenetusken/consolidate-memory/dream-beta-test)
Your own site
<a href="https://agentmods.dev/skills/zenetusken/consolidate-memory/dream-beta-test"><img src="https://agentmods.dev/badge/skills/zenetusken/consolidate-memory/dream-beta-test/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for dream-beta-test

Your own site · 80×15
<a href="https://agentmods.dev/skills/zenetusken/consolidate-memory/dream-beta-test"><img src="https://agentmods.dev/badge/skills/zenetusken/consolidate-memory/dream-beta-test.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 189 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,595 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00189 $0.01595
Opus 5 $0.00095 $0.00797
Sonnet 5 $0.00038 $0.00319
Haiku 4.5 $0.00019 $0.00160

Measured 9d ago against content hash 784f836556dc, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-08, from the pricing page.

Security

Grade A, and why

dream-beta-test scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/dream-beta-tester/skills/dream-beta-test/SKILL.md · 104 lines

How it starts

The opening of the file, as written. The whole thing — 104 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Dream beta-test

Beta-test the consolidate-memory "dream" skill in a target repo and emit one structured, version-stamped report. You are the dream skill's consumer / beta-tester — you run it faithfully, instrument it, and report defects. You never patch the dream skill (nothing under its plugin dir); this tooling ships as the dream-beta-tester plugin — its scripts live under $CLAUDE_PLUGIN_ROOT/scripts/ and you run them from there.

The harness has two detectors and your job is to run both:

  1. a deterministic invariant oracle (already built: beta_checks.py, driven by run_beta.py) that re-tests known defect families every run — the trustworthy floor;
  2. the judgment lenses (this skill's contribution) that you apply by hand to this run's actual artifacts to catch novel defects the families don't encode — the dynamic half.

Two properties make this worth doing, and everything below serves them:

  • Dynamic, not overfit — adapt to whatever this run surfaces; never just re-run a frozen checklist. The oracle re-tests known classes; the lenses find new ones.
  • Honest — a confident-wrong defect is worse than a missed one. The harness has already retracted two of its own hand-diagnoses, so every finding is re-verified against source before it is called confirmed, and suspected ones are quarantined, never counted.

Flow

1 · Run the deterministic engine

python3 "$CLAUDE_PLUGIN_ROOT/scripts/run_beta.py" --repo <TARGET_REPO> --test --json

--test (default) restores the store afterward — a pure beta-test that leaves no mutation; default repo is cwd. The runner snapshots the store, runs the oracle (which drives the dream skill's own read-only scripts from a clean subprocess pinned to the target repo), renders a report to reports/<slug>__<version>__<ts>.md, diffs the store, and restores.

Note the consolidate-memory version under test (stamped in the report header) and the report path. Read the rendered report.

Read the full file on GitHub · 104 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 9d ago First seen · 104 lines · 189 tokens per session scan A 784f836556dc

Subscribe to this mod's changes

dream-beta-test is a skill published in the GitHub repository Zenetusken/consolidate-memory (6 stars, last pushed yesterday), licensed MIT. It adds 189 tokens to every session and 1,595 once invoked, about $0.0009 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

alive:save

The human wants to checkpoint. Or: the stash has grown heavy — 5+ items, 30+ minutes, a natural pause in the work. The squirrel doesn't decide when to save. It surfaces the need and lets the human pull the trigger. Runs the full save protocol: confirms stash, writes log, updates state, generates projections…

alivecontext/alive · 78 tokens

alive:capture-context

Use when external content arrives in the session — emails, transcripts, screenshots, documents, files, or in-session research worth keeping. Also use when there's nothing obvious to capture — the skill checks 03Inbox/ for unrouted files and enters inbox scan mode. Stores raw content, routes to bundles, extracts tasks…

alivecontext/alive · 74 tokens

alive:load-context

The human mentions a walnut to work on, asks about a specific venture/experiment/project, or wants to check status — not just explicit 'load X'. Load the brief pack (3 files), resolve the people involved, check the active bundle — then surface one observation and ask what to work on. Context loads in tiers: walnut and…

alivecontext/alive · 81 tokens

alive:session-history

Revive sessions (quick or heavy), browse, and search — 'what happened recently?', 'find the session where we discussed X', 'revive yesterday's session'. For single-session recall and multi-session browsing. If the human needs to merge multiple sessions into one working context or detect conflicts between parallel…

alivecontext/alive · 76 tokens

alive:mine-for-context

Deep context extraction from source material. Creates reference bundles, builds extraction plans, tracks what's been extracted, and discovers new targets — people, subjects, patterns, connections. The archaeologist that turns raw sources into structured knowledge. Can be invoked by alive:session-history for targeted…

alivecontext/alive · 63 tokens

alive:my-context-graph

Render an interactive map of your world. Generates the world index from all walnut and bundle frontmatter, then produces a force-directed graph showing connections between walnuts, people, bundles, and tags. Opens in the browser.

alivecontext/alive · 50 tokens