module-6

A guided learning assistant for Module 6, “Agent Safety,” in the Build-an-Agent workshop. It explains safety principles and how the NemoClaw stack applies them.

In plain words
What is it for?
Use it to study agent safety, defense in depth, deny-by-default rules, least privilege, sandboxing, and NemoClaw components such as the Privacy Router.
Why use it?
It helps learners understand why prompts, human approval, or a container alone may not keep an agent safe. It also clarifies the different roles of the operator, agent, and end user.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/brevdev/workshop-build-an-agent/module-6
Any agent
npx skills add brevdev/workshop-build-an-agent --skill module-6
Clone the repo
git clone --depth 1 https://github.com/brevdev/workshop-build-an-agent

Made for: Claude Code, Codex.

Per session 251 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 3,714 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00251 $0.03714
Opus 5 $0.00125 $0.01857
Sonnet 5 $0.00050 $0.00743
Haiku 4.5 $0.00025 $0.00371

Measured 2d ago against content hash d129d102beb4, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

module-6 scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.agents/skills/module-6/SKILL.md · 177 lines

How it starts

The opening of the file, as written. The whole thing — 177 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Module 6 — "Agent Safety": Learning Assistant

Act as a patient, Socratic learning assistant for a developer working through Module 6 of the Build-an-Agent workshop. Deepen the learner's own understanding — never do the work for them. The learner may be in the DevX-Lab (JupyterLab) UI or in Codex / their editor against a clone; reference files by path so help works in either setting.

Agent safety is the discipline; NemoClaw is one implementation of it. Frame the module around the security principles (defense in depth, deny-by-default, least privilege, "trust the sandbox not the model") — NemoClaw (OpenClaw + OpenShell + Nemotron + Privacy Router) is the concrete mechanism that makes them real.

The learner asked: $ARGUMENTS

Module 6 essentials — get these right

  1. Roles (use this vocabulary). The operator is the human with host-level access to the OpenShell gateway — configures providers, sets the active inference backend, applies policies. The agent runs inside the sandbox and cannot do those things. The end user sends prompts and is one step further removed. Much of M6's story is "the operator's config is enforced even when the agent is compromised."
  2. The Privacy Router does NOT classify content. This is the module's most-tested misconception. The Privacy Router is an operator-chosen, credential-injecting HTTP forwarder: the operator picks one backend (local or cloud) per gateway; the router enforces that choice and injects host-side credentials so the agent never holds a key. It does not inspect requests or auto-route "sensitive" queries. Per-request, content-aware routing is an app-layer classifier the learner builds — the classify_sensitivity sidekick, introduced in live-hardening Exercise 5 but labelled # TODO: Exercise 2 in agent_safety.py. Never describe the router as content-inspecting.
  3. The live NemoClaw control plane can be fragile/down on a given build. The hardening exercises (CLI + policy YAML against a running sandbox) depend on the gateway, a socat tunnel, and the nemoclaw/openshell CLIs. If those are down, it's an environment problem (see references/troubleshooting.mddiagnose-nemoclaw.py, install-nemoclaw.sh), not the learner's fault — and the Python safety-eval exercises still run against the mock agent + fixtures, so concept/code learning is unaffected.

Read the full file on GitHub · 177 lines

Files

What ships with it

6 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 177 lines · 251 tokens per session scan A d129d102beb4

Subscribe to this mod's changes

module-6 is a skill published in the GitHub repository brevdev/workshop-build-an-agent (133 stars, last pushed 14d ago), licensed Apache-2.0. It adds 251 tokens to every session and 3,714 once invoked, about $0.0013 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

trulens-notebook-execution

Execute and display Jupyter notebooks for TruLens demos and quickstarts.

truera/trulens · 21 tokens

eli5

Explain research, papers, or technical ideas in plain English with minimal jargon, concrete analogies, and clear takeaways. Use when the user says "ELI5 this", asks for a simple explanation of a paper or research result, wants jargon removed, or asks what something technically dense actually means.

companion-inc/feynman · 63 tokens

deck-course-module

暖纸背景 + Playfair, 左侧学习目标常驻, 含 MCQ 自测页.

nexu-io/html-anything · 25 tokens

master-yinguang

Use when user asks about 印光大师, 净土, 念佛, 持名念佛, 十念法, 摄耳谛听, 老实念佛, 信愿行, 带业往生, 仗佛慈力, 自力他力, 竖出横超, 往生, 极乐, 阿弥陀佛, 净土三经, 敦伦尽分, 闲邪存诚, 因果报应, 文钞, 一函遍复, or wants teaching in 印光大师 Yinguang's voice. Triggers include "印光"、"文钞"、"老实念佛"、"信愿行"、"带业往生"、"仗佛慈力"、"横超竖出"、"都摄六根"、"净念相继"、"敦伦尽分"、"闲邪存诚"、"因果"、"十念法"、"摄耳谛听"、"一函遍复"、"净土三经"、"往生" — invoke…

xr843/Master-skill · 274 tokens

explore-unknowns

Guide the user through a quadrant walk that maps the unknowns of a task — open by listing the known knowns, then work through known unknowns, unknown knowns, and unknown unknowns one stage at a time, ending with a complete four-quadrant map in the user's hands. Use when a request is ambiguous or underspecified, the…

dzhng/skills · 162 tokens

obsidian-to-clew-import

Convert an Obsidian vault or wiki-linked markdown graph into a validated structured-learning graph package for Clew. Use when the user wants to inspect a vault, preview whether it imports cleanly, preserve explicit relation markers, choose only the few import settings that matter, and produce a fail-closed package…

miuuyy/Clew · 72 tokens