Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add richfrem/agent-plugins-skills --skill os-experiment-loggit clone --depth 1 https://github.com/richfrem/agent-plugins-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/richfrem/agent-plugins-skills/os-experiment-log)<a href="https://agentmods.dev/skills/richfrem/agent-plugins-skills/os-experiment-log"><img src="https://agentmods.dev/badge/skills/richfrem/agent-plugins-skills/os-experiment-log.svg" alt="Measured on agentmods" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00110 | $0.00891 |
| Opus 5 | $0.00055 | $0.00445 |
| Sonnet 5 | $0.00022 | $0.00178 |
| Haiku 4.5 | $0.00011 | $0.00089 |
Grade A, and why
os-experiment-log scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 74 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Overview
The experiment log is the unified cross-cutting record for all agentic-os experiments.
One file per run, all files in context/experiment-log/, with index.md as a
queryable table of all runs.
context/experiment-log/
index.md ← one row per run (date, source, target, verdict)
2026-04-25-verifier-os-architect-round1.md ← from os-evolution-verifier
2026-04-25-tester-os-architect.md ← from os-architect-tester
2026-04-25-os-improvement-loop-os-eval-runner.md ← from os-improvement-loop
2026-04-25-planner-0024.md ← from os-evolution-planner
2026-04-25-survey-session.md ← from post_run_survey
Source Types and Result Kinds
Agents must check result_type in a log entry's header before parsing it:
--source-type |
Produced by | result_type |
Key fields |
|---|---|---|---|
verifier |
os-evolution-verifier | qualitative |
PASS/PARTIAL/FAIL counts, HANDOFF_BLOCK validity |
tester |
os-architect-tester | qualitative |
AC-1–4 pass/fail per scenario |
orchestrator |
os-improvement-loop | numeric |
best_score, baseline, delta, KEEP/DISCARD counts |
planner |
os-evolution-planner | qualitative |
workstream count, gaps identified |
survey |
post_run_survey | mixed |
friction item count, north_star metric |
Numeric entries (result_type: numeric) carry quantitative metrics suitable for trending and charting.
Qualitative entries (result_type: qualitative) carry pass/fail verdicts and gap analysis prose.
Mixed entries (result_type: mixed) carry both — agents must check which fields are present before parsing.
Phase 1 — Resolve Mode
Read the argument or invocation context to determine mode:
append --source-type TYPE: log a new run from a completed experimentquery <term>: search all files incontext/experiment-log/by keywordsummary: print aggregate stats across all runs, broken down by source type
What ships with it
5 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- today Changed · -113 lines 02a790861adc
- 8d ago First seen · 187 lines · 110 tokens per session scan A f11830458bf1
os-experiment-log is a skill published in the GitHub repository richfrem/agent-plugins-skills (6 stars, last pushed today), licensed MIT. It adds 110 tokens to every session and 891 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
composio
Use Composio from Agent Swarm through the agent-swarm x composio CLI route, the swarmx MCP tool, or a registered ctx.api.composio script connection. Trigger when a task needs connected third-party app tools such as Gmail, Google Calendar, Google Docs, Google Drive, GitHub, Slack, Notion, or HubSpot through Tool Router…
close-issue
Close a GitHub or GitLab issue with a summary comment.
soul
Embody this digital identity. Read SOUL.md first, then STYLE.md, then examples/. Become the person—opinions, voice, worldview.
implement-issue
Implement a GitHub issue or GitLab issue and create a PR/MR.
bbc-skill
A read-only tool for collecting all comments from a Bilibili video, including replies and pinned comments. Bilibili is a Chinese video-sharing platform.
feature-implementation
Implement software features, fix bugs, and write production-quality code. Use this for hands-on development work including writing new code, modifying existing code, and debugging issues.