Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add vasuag09/harness-claude --skill observegit clone --depth 1 https://github.com/vasuag09/harness-claudeWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/vasuag09/harness-claude/observe)<a href="https://agentmods.dev/skills/vasuag09/harness-claude/observe"><img src="https://agentmods.dev/badge/skills/vasuag09/harness-claude/observe/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/vasuag09/harness-claude/observe"><img src="https://agentmods.dev/badge/skills/vasuag09/harness-claude/observe.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00131 | $0.01261 |
| Opus 5 | $0.00066 | $0.00630 |
| Sonnet 5 | $0.00026 | $0.00252 |
| Haiku 4.5 | $0.00013 | $0.00126 |
Grade A, and why
observe scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 67 lines — stays where its author put it; the contents beside it link to each section on GitHub.
/observe — triage the signal, locate the area, route to /fix
Goal: turn a raw production signal into something the harness can act on. The user brings the
signal (the harness does not watch anything); /harness-claude:observe grounds it in real code —
which component, what failure, what would reproduce it — and hands that package to
/harness-claude:fix, which runs the reproduce-first discipline. This is the step that makes the
pipeline a loop rather than a line: …/ship → /deploy → (prod) → /observe → /fix → ….
Opt-in. Invoked explicitly. Bring-a-signal only — no polling, no credentials, no new MCP; you paste/point at the signal. Git boundary:
/harness-claude:observeis read-only — it never commits, pushes, or branches; the fix is committed (if at all) only later through/harness-claude:ship.
When /harness-claude:observe, when /harness-claude:fix
/harness-claude:observeis for a prod signal you don't yet know how to reproduce in code — it triages and locates, then routes. Start here when all you have is a stack trace / log / alert./harness-claude:fixis the lane itself — reproduce-as-a-failing-test-first, root-cause, minimal fix. If you already have a red test or a known repro, skipobserveand go straight tofix.
How it works
parse the brought signal → locate the failing area (mgrep/graph) → shape a repro SEED → hand to /fix
Do this
- Parse the brought signal. From the pasted stack trace / log / error / issue URL, extract a structured failure description: the service or component, the error message or pattern, the conditions it appears under, and any data visible in the signal. No polling, no credentials, no new MCP — you work only from what was handed to you. (AC-O1)
- Locate the failing area in real code. Use
mgrep/ the knowledge graph to trace the signal to its origin — the function, module, or path the error points at. Ground the signal in the codebase rather than guessing; name the suspected location. (AC-O2) - Shape a repro SEED — not the finished test. Describe the smallest input/state likely to trigger
the failure (or concrete manual repro steps when automation genuinely isn't possible — the same
carve-out as
/harness-claude:fixstep 1). Do NOT write the finished failing test here — capturing the bug RED is/harness-claude:fix's non-negotiable first step; duplicating it blurs the boundary. (AC-O3) - Hand off to
/harness-claude:fix— then stop. Pass the package: the structured failure description (1), the located area (2), and the repro seed (3)./harness-claude:fixruns reproduce → root-cause → minimal fix-plan and rides the existing review/verify/ship gates./harness-claude:observedoes not duplicate any of that, and stays read-only throughout. (AC-O4)
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 67 lines · 131 tokens per session scan A a8805dd402e4
observe is a skill published in the GitHub repository vasuag09/harness-claude (2 stars, last pushed 2mo ago), licensed MIT. It adds 131 tokens to every session and 1,261 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
issue-triage
3-phase issue backlog management with audit, deep analysis, and validated triage actions. Use when triaging GitHub issues, sorting bug reports, cleaning up stale tickets, or detecting duplicate issues. Args: 'all' to analyze all, issue numbers to focus (e.g. '42 57'), 'en'/'fr' for language, no arg = audit only.
check-cache-bugs
Audit Claude Code setup for cache bugs (CC#40524): sentinel, --resume/--continue, attribution header + ArkNill B3/B4/B5.
eval-rules
Audit .claude/rules/ files for structural correctness, glob validity, and real-world usefulness. Resolves each paths: pattern against actual project files, then asks the user whether each rule is still relevant and useful. Can update rules in-place based on answers. Use when setting up rules for the first time…
audit-codebase
Codebase health audit scoring 7 categories with progression plan.
pentest-forensics
Digital forensics — evidence acquisition, memory/disk imaging analysis, timeline reconstruction, IOC extraction advisory. Triggers on forensics, DFIR, Volatility, memory analysis, disk image, Autopsy, FTK, timeline, IOC extraction, evidence chain, log analysis.
pentest-malware
Malware analysis — triage, static analysis, dynamic sandbox, IOC extract, YARA signature writing advisory. Triggers on malware analiz, malware triage, sandbox, Cuckoo, IDA, Ghidra, dynamic analysis, IOC, YARA imza, packer, unpacker, reverse malware.