harness-engineering

harness-engineering is a skill for Claude Code, Codex from Mark393295827/graph-engineering-architectures. It costs 34 tokens per session (1,425 once invoked), scanned A, original, MIT.

A control layer for agent workflows that sets boundaries around context, tools, permissions, logging, scheduling, evaluation, recovery, and maintenance. It treats the agent as one part of a larger runtime environment.

In plain words
What is it for?
Use it when designing production-like agent workflows with sensitive data, scheduled work, external tools, delegated actions, audit trails, or recovery requirements.
Why use it?
It makes delegated agent work easier to inspect and constrain, including what the agent saw, decided, called, changed, and verified. It also addresses failure recovery and repeatable testing.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it when designing production-like agent workflows with sensitive data, scheduled work…

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/mark393295827/graph-engineering-architectures/harness-engineering
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add Mark393295827/graph-engineering-architectures --skill harness-engineering
Clone the repo
git clone --depth 1 https://github.com/Mark393295827/graph-engineering-architectures

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for harness-engineering

README.md
[![agentmods](https://agentmods.dev/badge/skills/mark393295827/graph-engineering-architectures/harness-engineering.svg)](https://agentmods.dev/skills/mark393295827/graph-engineering-architectures/harness-engineering)
Your own site
<a href="https://agentmods.dev/skills/mark393295827/graph-engineering-architectures/harness-engineering"><img src="https://agentmods.dev/badge/skills/mark393295827/graph-engineering-architectures/harness-engineering.svg" alt="Measured on agentmods" height="20"></a>
Per session 34 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,425 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00034 $0.01425
Opus 5 $0.00017 $0.00713
Sonnet 5 $0.00007 $0.00285
Haiku 4.5 $0.00003 $0.00143

Measured 6d ago against content hash 1f48e3c566a0, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-06, from the pricing page.

Security

Grade A, and why

harness-engineering scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/harness-engineering/SKILL.md · 114 lines

How it starts

The opening of the file, as written. The whole thing — 114 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Harness Engineering

<skill_contract> An agent workflow, runtime environment, tools, data sensitivity, effects, cadence, risk, and operator constraints. An auditable runtime kernel with scoped permissions, scheduling, observability, recovery, and eval controls. An end-to-end trace and failure-path tests prove bounded, replayable, recoverable delegated action. <non_goals>Business-task decomposition, prompt-only safety, broad credentials, or unbounded scheduled autonomy.</non_goals>

Treat the harness as the kernel around an LLM OS: context is RAM, durable state is disk, tools are system calls, skills are programs, the scheduler is control, and evals are verifiers. Load references/runtime-control-patterns.md for matrices and schemas.

Usage Template

Provide: workflow, users, agent roles, environment, tools/connections, data sensitivity, delegated actions, cadence, throughput/SLA, failure history, and risk tolerance.

Workflow

Run the trace gate: the harness must be able to show what the agent saw, decided, called, changed, and verified. Separate Agent (instructions/capabilities), Environment (network/files/credentials), and Session (mounted context/events/state). Define one auditable control path for high-risk intent and final joins.

<unknowns_gate>

If state ownership, credential scope, external side effects, retention, or approval authority is unclear, return NEEDS_INPUT. Probe tools with read-only discovery where possible; unknown side effects default to denied.

</unknowns_gate>

  1. Pass Four-C: Context truth/retrieval, Connections scoped accounts/APIs, Capabilities versioned skills/scripts/evals, Cadence trigger/receipt/anomaly/stop.
  2. Map runtime: stored program, control unit, hot context, durable disk, event bus, I/O tools, verifier, and garbage collector.
  3. Choose the lowest-context primitive: deterministic script/hook, skill, static Graph, connector, dynamic workflow, or agent team. Load capabilities lazily. Graph Engineering owns dependency semantics; the harness owns the ready queue, leases, duplicate delivery, concurrency, and executor health.
  4. Define each tool as a narrow system call with purpose, explicit inputs, bounds, timeout, idempotency, failure path, evidence, and audit location.
  5. Enforce least privilege in the environment, not only prose. Escalate autonomy through observe -> co-drive -> scoped reversible action -> monitored routine -> audited low-risk autonomy.
  6. For delegated action require mandate, scope, limit, preview, receipt, and rollback. Human approval governs irreversible/shared/financial/published/credentialed actions.
  7. Add deterministic feedback (tests, lint, LSP, policy checks) outside context when possible; add independent evaluator/red team for high-risk semantic output.
  8. Persist an append-only session event log and checkpoint; define alerts, fallback, incident response, cleanup, permission review, and stale-context/rule review.
  9. For scheduled work define Trigger, Context, Steering, Receipt, budget, stop, recovery, and executor health. A schedule firing is not task success.
  10. For Graph execution, persist node/edge/join transitions before releasing successors, make delivery idempotent, recover from the last verified checkpoint, and test permission denial, worker loss, duplicate events, and compensation without relying on in-memory scheduler state.

Read the full file on GitHub · 114 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 6d ago First seen · 114 lines · 34 tokens per session scan A 1f48e3c566a0

Subscribe to this mod's changes

harness-engineering is a skill published in the GitHub repository Mark393295827/graph-engineering-architectures (2 stars, last pushed 16d ago), licensed MIT. It adds 34 tokens to every session and 1,425 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

agent-framework-py-release

Use when cutting a Python release for the microsoft/agent-framework monorepo. Triggers on "bump py versions", "cut a python release", "prepare release PR for python", "release py packages", "bump python to X.Y.Z", or similar requests to bump Python package versions and prepare a release PR. Handles all four lifecycle…

microsoft/agent-framework · 103 tokens

python-package-management

Guide for managing packages in the Agent Framework Python monorepo, including creating new connector packages, versioning, and the lazy-loading pattern. Use this when adding, modifying, or releasing packages.

microsoft/agent-framework · 43 tokens

foundry-hosted-agent-validation

Step-by-step process for validating a Python Foundry hosted agent sample (under python/samples/04-hosting/foundry-hosted-agents/) end to end — running it locally (native runtime and azd ai agent run) and after deploying it to an Azure AI Foundry project with azd. Use this when asked to validate a hosted agent sample.

microsoft/agent-framework · 82 tokens

verify-samples-tool

How to use the verify-samples tool to run, verify, and manage sample definitions in the Agent Framework repository. Use this when adding, updating, or running sample verification.

microsoft/agent-framework · 40 tokens

build-and-test

How to build and test .NET projects in the Agent Framework repository. Use this when verifying or testing changes.

microsoft/agent-framework · 26 tokens

python-feature-lifecycle

Guidance for package and feature lifecycle in the Agent Framework Python codebase, including stage meanings, feature-stage decorators, feature enums, and how to move APIs from one stage to the next.

microsoft/agent-framework · 43 tokens