Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/kastalien-research/thoughtbox/thoughtbox-evolutionnpx skills add Kastalien-Research/thoughtbox --skill thoughtbox-evolutiongit clone --depth 1 https://github.com/Kastalien-Research/thoughtboxWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/kastalien-research/thoughtbox/thoughtbox-evolution)<a href="https://agentmods.dev/skills/kastalien-research/thoughtbox/thoughtbox-evolution"><img src="https://agentmods.dev/badge/skills/kastalien-research/thoughtbox/thoughtbox-evolution.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00102 | $0.01081 |
| Opus 5 | $0.00051 | $0.00541 |
| Sonnet 5 | $0.00020 | $0.00216 |
| Haiku 4.5 | $0.00010 | $0.00108 |
Grade A, and why
thoughtbox:evolution scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 132 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Thought Evolution (A-Mem Pattern)
When you add a new insight to a reasoning session, earlier thoughts don't automatically update. Thought 1 might say "consider rate limiting" while thought 15 decides "use sliding window algorithm" — but thought 1 doesn't know about the sliding window decision. This skill checks which prior thoughts should evolve.
Based on the A-Mem paper (arxiv.org/abs/2502.12110): when new memory is added, find related existing memories and update their context.
When to Trigger
Run evolution checks when the new thought:
- Resolves ambiguity from earlier thoughts
- Contradicts an earlier assumption
- Adds implementation detail to a high-level earlier thought
- Synthesizes multiple earlier threads into a conclusion
Don't run for every thought — only on significant ones (synthesis, conclusions, decisions, revisions).
Workflow
Phase 1: Retrieve Session Content
// thoughtbox_execute
async () => {
const session = await tb.session.get("current-session-id");
return session.thoughts.map((t, i) => ({
number: t.thoughtNumber,
content: t.thought.slice(0, 200) // Truncate for efficiency
}));
}
Phase 2: Spawn Evolution Checker
Dispatch a Haiku subagent for cost efficiency (~400 tokens in subagent context, ~50 tokens returned):
Spawn subagent (model: haiku):
"Evaluate which prior thoughts should be updated based on a new insight.
NEW INSIGHT:
[Your new thought content]
PRIOR THOUGHTS:
S1: [thought 1 content]
S2: [thought 2 content]
...
For each thought, respond ONLY with:
S1: [UPDATE|NO_UPDATE] - [brief reason if UPDATE]
S2: [UPDATE|NO_UPDATE] - [brief reason if UPDATE]
...
Be selective. Only suggest UPDATE if the new insight meaningfully enriches
the prior thought's context. Keyword overlap alone is not enough."
Phase 3: Apply Revisions
For each thought marked UPDATE, create a revision:
async () => {
await tb.thought({
thought: "EVOLVED: [original content] — Now contextualized: [how new insight relates]",
thoughtType: "reasoning",
isRevision: true,
revisesThought: 1, // The thought number being updated
thoughtNumber: 20, // Current thought number (advances the chain)
totalThoughts: 25,
nextThoughtNeeded: true
});
}
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 132 lines · 102 tokens per session scan A 3368d88b93f4
thoughtbox:evolution is a skill published in the GitHub repository Kastalien-Research/thoughtbox (64 stars, last pushed 1mo ago), licensed MIT. It adds 102 tokens to every session and 1,081 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
unity-addressables
Manage Addressables groups, entries, profiles and content builds (com.unity.addressables, reflection-based).
impeccable
Use when the user wants to design, redesign, shape, critique, audit, polish, clarify, distill, harden, optimize, adapt, animate, colorize, extract, or otherwise improve a frontend interface. Covers websites, landing pages, dashboards, product UI, app shells, components, forms, settings, onboarding, and empty states.…
render-airdrop-carousel
Assemble a viral iOS "AirDrop" notification-carousel video ad (≈6–8s, 9:16) from a brand line plus 6–16 real product photos — a native AirDrop share-sheet card ("Brand would like to share a · Decline / Accept") springs up and its preview window CYCLES through the products, landing on a range/lineup payoff with an…
render-3d-product-showcase
Assemble a premium 3D product-showcase ad from a config — four beat clips (an orbiting hero rotation, a macro push-in, a physics reveal, a typographic close) normalized to the brand-color canvas, hard-concatenated in order, closed on a deterministic Playwright brand end card, and mixed under one instrumental bed at…
opik-diagnose
Surface the Opik traces worth a developer's attention, ranked by signal — errors, failed tool calls, latency, regressions, and low online-eval scores — plus Diagnostics issues. Reads live/production traces via the SDK (searchtraces and agentinsights) and works with no MCP; uses the MCP issue entity when connected.…
Threat Hunting & IOC Analysis
IOC extraction, threat intelligence correlation, MITRE ATT&CK mapping, hunt hypothesis generation, and detection rule creation.