Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/cocorof/geny-executor/tool-buildernpx skills add CocoRoF/geny-executor --skill tool-buildergit clone --depth 1 https://github.com/CocoRoF/geny-executorWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00076 | $0.00843 |
| Opus 5 | $0.00038 | $0.00421 |
| Sonnet 5 | $0.00015 | $0.00169 |
| Haiku 4.5 | $0.00008 | $0.00084 |
Grade A, and why
Sandbox Tool Builder scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 77 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Building your own sandboxed tool
When you lack a tool you need, write one. A sandboxed tool is just a script
in your workspace that reads a JSON request on stdin and writes a JSON response
on stdout. You author it, test it, then forge_tool registers it as a real tool
you can call from the next turn — and it always runs inside your sandbox, not
on the host.
You need a sandbox (an isolated workspace) attached to this session. If you
don't have one, forge_tool will tell you.
The loop
-
Write the script into your workspace with your file-write tool, at a stable path like
tools/<name>/main.py. The contract:- read the whole request from stdin (a JSON object),
- write the result to stdout as a JSON object,
- on failure, print
{"error": "..."}(or exit non-zero) — that surfaces as a tool error.
Minimal Python example (
tools/wordcount/main.py):import sys, json req = json.load(sys.stdin) text = req.get("text", "") print(json.dumps({"words": len(text.split()), "chars": len(text)})) -
Test it with your shell tool, exactly how the tool will be invoked:
echo '{"text":"a b c"}' | python3 tools/wordcount/main.pyIterate until the output JSON is right.
-
Forge it — register it live:
env(action="forge_tool", args={ "name": "wordcount", "description": "Count words and characters in text.", "entrypoint": "tools/wordcount/main.py", "runtime": "python3", "input_schema": {"type":"object","properties":{"text":{"type":"string"}},"required":["text"]} })From the next turn you can call
wordcount({"text": "..."})like any tool. -
Use it. Call it by name. If it misbehaves, fix the script, test again, and re-
forge_tool(disable the old one first if the name clashes).
Persisting it (reuse across sessions)
forge_tool is for this session. To keep it, save a Sandbox Tool Pack:
env(action="save_pack", args={"name": "text-utils", "description": "word + slug tools"})
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 77 lines · 76 tokens per session scan A 762a7bd4d755
Sandbox Tool Builder is a skill published in the GitHub repository CocoRoF/geny-executor (2 stars, last pushed 3d ago), licensed Apache-2.0. It adds 76 tokens to every session and 843 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
agent-framework-py-release
Use when cutting a Python release for the microsoft/agent-framework monorepo. Triggers on "bump py versions", "cut a python release", "prepare release PR for python", "release py packages", "bump python to X.Y.Z", or similar requests to bump Python package versions and prepare a release PR. Handles all four lifecycle…
foundry-hosted-agent-validation
Step-by-step process for validating a Python Foundry hosted agent sample (under python/samples/04-hosting/foundry-hosted-agents/) end to end — running it locally (native runtime and azd ai agent run) and after deploying it to an Azure AI Foundry project with azd. Use this when asked to validate a hosted agent sample.
python-feature-lifecycle
Guidance for package and feature lifecycle in the Agent Framework Python codebase, including stage meanings, feature-stage decorators, feature enums, and how to move APIs from one stage to the next.
build-and-test
How to build and test .NET projects in the Agent Framework repository. Use this when verifying or testing changes.
python-code-quality
Code quality checks, linting, formatting, and type checking commands for the Agent Framework Python codebase. Use this when running checks, fixing lint errors, or troubleshooting CI failures.
python-development
Coding standards, conventions, and patterns for developing Python code in the Agent Framework repository. Use this when writing or modifying Python source files in the python/ directory.