Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/internlm/wildclawbench/03_task3npx skills add InternLM/WildClawBench --skill 03_task3git clone --depth 1 https://github.com/InternLM/WildClawBenchWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00041 | $0.00501 |
| Opus 5 | $0.00020 | $0.00251 |
| Sonnet 5 | $0.00008 | $0.00100 |
| Haiku 4.5 | $0.00004 | $0.00050 |
Grade A, and why
03_task3 scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Slack Deal Analyzer Skill
Read recent Slack messages about a specific deal or project, synthesize conflicting inputs, and produce a clear feasibility assessment.
Tools
All tools are defined in tmp_workspace/utils.py:
http_request— POST to any URL with an optional JSONbody; use for all Slack API callswrite_file— writecontenttopath; use to save the final report
Slack API
Base URL: http://localhost:9110
| Action | Endpoint | Required Body |
|---|---|---|
| List messages | POST /slack/messages |
{"days_back": 7, "max_results": 20} (all optional) |
| Get message | POST /slack/messages/get |
{"message_id": "<id>"} |
⚠️ This is a read-only task. Do not call
slack_send_message(POST /slack/send).
Workflow
- List messages — fetch recent messages with
slack_list_messages - Identify relevant messages — filter for messages related to the deal/project in question
- Read each in full — retrieve complete content via
slack_get_message - Synthesize — across all messages, extract:
- What has been promised or proposed
- Requirements that have shifted or are in conflict
- Risks and blockers flagged by different stakeholders
- What is realistically deliverable vs. what will likely fail
- Write report — save findings to
/tmp_workspace/results/results.md
Report Format
# Deal Feasibility Assessment: [Deal Name]
## What's Been Committed / Proposed
- ...
## Shifting or Conflicting Requirements
- ...
## Risks & Likely Failure Points
- ...
## Realistic Deliverability
- What we can deliver: ...
- What we cannot deliver: ...
## Recommended Stance
...
Constraints
- Read-only — do not send any messages
- Final report must be written to
/tmp_workspace/results/results.md
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 73 lines · 41 tokens per session scan A 959cc38ca506
03_task3 is a skill published in the GitHub repository InternLM/WildClawBench (516 stars, last pushed 16d ago), licensed MIT. It adds 41 tokens to every session and 501 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
agent-browser
Browser automation CLI for AI agents. Use when the user needs to interact with websites, including navigating pages, filling forms, clicking buttons, taking screenshots, extracting data, testing web apps, or automating any browser task. Triggers include requests to "open a website", "fill out a form", "click a…
naga-config
Naga 自身配置管理技能。用于查看和修改 Naga 系统设置、添加 MCP 工具服务、导入自定义技能、搜索可用 MCP 工具。当用户要求修改设置、添加工具或技能时使用此技能。.
naga_control
通过 agentType: "nagacontrol" 调用,直接控制 Naga 自身的运行状态和配置。.
file-manager
文件管理技能。用于创建、移动、复制、删除文件和文件夹,整理目录结构。当用户需要管理文件、整理文件夹或批量处理文件时使用。.
live2d_controller
// 读取:system.characterbundle.loadcharacterskillsections -> system.config.buildtier1variables.characterbuiltinskillsprompt.
code-review
代码审查和质量分析技能。用于审查代码、发现潜在问题、提供改进建议。当用户请求代码审查、代码质量分析或最佳实践建议时使用。.