Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Mark393295827/third-brain-v7-skills --skill harness-engineeringgit clone --depth 1 https://github.com/Mark393295827/third-brain-v7-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/mark393295827/third-brain-v7-skills/harness-engineering)<a href="https://agentmods.dev/skills/mark393295827/third-brain-v7-skills/harness-engineering"><img src="https://agentmods.dev/badge/skills/mark393295827/third-brain-v7-skills/harness-engineering/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/mark393295827/third-brain-v7-skills/harness-engineering"><img src="https://agentmods.dev/badge/skills/mark393295827/third-brain-v7-skills/harness-engineering.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00034 | $0.01795 |
| Opus 5 | $0.00017 | $0.00898 |
| Sonnet 5 | $0.00007 | $0.00359 |
| Haiku 4.5 | $0.00003 | $0.00179 |
Grade A, and why
harness-engineering scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 135 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Harness Engineering
<skill_contract> An agent workflow, runtime environment, tools, data sensitivity, effects, cadence, risk, and operator constraints. An auditable runtime kernel with scoped permissions, scheduling, observability, recovery, and eval controls. An end-to-end trace and failure-path tests prove bounded, replayable, recoverable delegated action. <non_goals>Business-task decomposition, prompt-only safety, broad credentials, or unbounded scheduled autonomy.</non_goals>
Treat the harness as the kernel around an LLM OS: context is RAM, durable state is disk, tools are system calls, skills are programs, the scheduler is control, and evals are verifiers. Load references/runtime-control-patterns.md for matrices and schemas. Start guarded automation from references/runtime-envelope-example.json and validate it with scripts/validate_runtime_envelope.py --strict.
Usage Template
Provide: workflow, users, agent roles, environment, tools/connections, data sensitivity, delegated actions, cadence, throughput/SLA, failure history, and risk tolerance.
Workflow
Run the trace gate: the harness must be able to show what the agent saw, proposed, called, changed, and verified. Separate Intent Plan (human-reviewable source), Compiled Contract (validated runtime envelope), Agent (instructions/capabilities), Environment (network/files/credential broker), and Session (mounted context/events/state). Define one auditable control path for high-risk intent and final joins. Model output is never execution authority.
<unknowns_gate>
If state ownership, credential scope, external side effects, retention, or approval authority is unclear, return NEEDS_INPUT. Probe tools with read-only discovery where possible; unknown side effects default to denied.
</unknowns_gate>
- Pass Four-C: Context truth/retrieval, Connections scoped accounts/APIs, Capabilities versioned skills/scripts/evals, Cadence trigger/receipt/anomaly/stop.
- Compile the reviewed intent plan into a versioned runtime envelope. Validate
plan hash,
tool_execution_owner: host, filesystem/network/secret boundaries, output cardinality, legal no-op, budgets, approvals, audit paths, and rollback before execution. - Map runtime: stored program, control unit, hot context, durable disk, event bus, I/O tools, verifier, and garbage collector.
- Choose the lowest-context primitive: deterministic script/hook, skill, static Graph, connector, dynamic workflow, or agent team. Load capabilities lazily. Graph Engineering owns dependency semantics; the harness owns the ready queue, leases, duplicate delivery, concurrency, and executor health.
- Define each tool as a narrow host-owned system call with purpose, explicit inputs, bounds, timeout, idempotency, failure path, evidence, and audit location. Validate model-proposed arguments before dispatch.
- Enforce zero trust and least privilege in the environment, not only prose: bind access to task, resource, operation, and time; use exact network allowlists and opaque secret handles; never expose raw credentials to model context. Stage and vet writes before external commit.
- Normalize each model
termination_reasoninto complete, tool request, checkpoint/truncation, refusal/error, or unknown. The host decides whether to execute, continue, checkpoint, or escalate; success prose cannot override the control signal or verifier. - For delegated action require mandate, scope, limit, preview, receipt, and rollback. Human approval governs irreversible/shared/financial/published/credentialed actions.
- Define allowed output types, maximum external outputs, and a verifiable
NO_OPcondition. Quiet execution is success only when eligibility was checked and no side effect occurred. - Add deterministic feedback (tests, lint, LSP, policy checks) outside context when possible; add independent evaluator/red team for high-risk semantic output.
- Persist an append-only session event log and checkpoint; define alerts, fallback, incident response, cleanup, permission review, and stale-context/rule review.
- For scheduled work define Trigger, Context, Steering, Receipt, budget, stop, recovery, and executor health. A schedule firing is not task success.
- For Graph execution, persist node/edge/join transitions before releasing successors, make delivery idempotent, recover from the last verified checkpoint, and test permission denial, worker loss, duplicate events, and compensation without relying on in-memory scheduler state.
What ships with it
4 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 135 lines · 34 tokens per session scan A 4f7070f08b96
harness-engineering is a skill published in the GitHub repository Mark393295827/third-brain-v7-skills (138 stars, last pushed 20d ago), licensed MIT. It adds 34 tokens to every session and 1,795 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
obsidian-ai-setup
Bootstrap an Obsidian vault with the AI Pro Obsidian Plugin. Creates the complete vault structure, system files, Obsidian configuration, memory system, hooks, and output styles, then runs personalized onboarding to tailor the workspace to the user. Uses a single universal structure with no mode selection. Trigger when…
wiki-retrieve
Build and query a vault-local contextual BM25 retrieval index with optional multilingual Nomic cosine reranking; use for retrieve, hybrid retrieval, BM25, rerank, contextual retrieval, chunk search, vault search, semantic search, find relevant passages, or retrieval diagnostics. Derived caches stay under .vault-meta…
obsidian
Comprehensive guidelines for Obsidian.md plugin development including ESLint rules from eslint-plugin-obsidianmd v0.4.1, TypeScript best practices, memory management, API usage (requestUrl vs fetch), UI/UX standards, popout window compatibility, community.obsidian.md submission process, and Scorecard optimization. Use…
superbrain-distill
Internal SuperBrain skill — run by the detached capture child to distill a session-event delta into routed Obsidian notes. Not for direct user invocation.
pos-verify
Use this immediately after files are created, edited, moved, deleted, or materially rewritten inside PersonalOS. Verifies that new truth was routed to the correct owner, written in the correct file shape, and still follows POS conventions. Do NOT use for whole-vault deep audits; use system-health-check.
skillify
Use this when {{username}} asks to skillify a repeated workflow, determine whether it deserves a reusable PersonalOS skill, or harden an existing workflow into a tested resolver-reachable capability. Do NOT use for one-off notes, ordinary execution, or already-specified skill authoring; use write-skill.