task

A guide for writing YAML files that define tasks for coder-eval, a tool that tests coding agents against set criteria.

In plain words
What is it for?
It is for creating or extending coding-agent evaluation tasks and checking their scoring criteria with coder-eval.
Why use it?
It helps turn a coding goal into a small, checkable evaluation and verifies that the evaluation tool is available first.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/uipath/coder_eval/task
Any agent
npx skills add UiPath/coder_eval --skill task
Clone the repo
git clone --depth 1 https://github.com/UiPath/coder_eval

Made for: Claude Code, Codex.

Per session 46 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,816 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00046 $0.02816
Opus 5 $0.00023 $0.01408
Sonnet 5 $0.00009 $0.00563
Haiku 4.5 $0.00005 $0.00282

Measured 2d ago against content hash a7d6184cba96, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

task scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/coder-eval/skills/task/SKILL.md · 244 lines

How it starts

The opening of the file, as written. The whole thing — 244 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Author a coder-eval task

You are writing coder-eval task YAML. The user's request is: $ARGUMENTS

If $ARGUMENTS is empty, ask what the task should test. Do not invent a subject.

Good tasks use simple prompts: state the goal and the expected output, then let the agent work out the approach. A single request can produce several task files — "create tasks for all the registry subcommands" means one task per subcommand.

Step 1 — Understand the request, and check the CLI is there

Run coder-eval --version first. Steps 6 and 7 both shell out to it, and finding that out after writing several task files means the user gets a bare command not found with nothing to act on. Installing this plugin did not install the CLI.

If it is missing, follow ${CLAUDE_PLUGIN_ROOT}/reference/cli-setup.md: offer the install, ask before running it, and confirm with coder-eval --version afterwards. Never install unprompted, and do not write any task files if the user declines.

That reference also covers the other half of the version check — whether this project pins a coder-eval version, and what to do when the installed one does not match it.

Then establish:

  • What is being tested — which tool, SDK, CLI, skill, or capability?
  • How many tasks — one operation, or several?
  • Difficulty — smoke, basic, or intermediate?
  • Dependencies — network, packages, starter files, external services?

State any assumptions you make rather than silently picking.

Step 2 — Look at what already exists

Find the repository's task tree by following ${CLAUDE_PLUGIN_ROOT}/reference/repo-layout.md, and say what you resolved. If a task already covers this ground, say so and offer to modify it instead of adding a near-duplicate.

Repo-local convention beats anything bundled with this plugin — where the two disagree, the repo wins. Before writing, read what the repository declares about task authoring: its own contributor or convention documents, a task template if it ships one, and a few neighbouring tasks. Adopt what you find — naming, tags, thresholds, weights, where files go — and say in your report which conventions you adopted, so the choice is visible rather than implied.

Read the full file on GitHub · 244 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 244 lines · 46 tokens per session scan A a7d6184cba96

Subscribe to this mod's changes

task is a skill published in the GitHub repository UiPath/coder_eval (119 stars, last pushed 4d ago), licensed Apache-2.0. It adds 46 tokens to every session and 2,816 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

execution

M-1.4 execution skill — 跑 single task 产 patch + 提交 envelope。.

Towow-ai/Flowness · 22 tokens

fix-self-check

M-1.6 envelope self-check——独立性保证不自欺欺人 (5 blockingcheck)。由 CLI ./tw fix complete --self-check-mode fork(默认即 fork)自动派起,不经 Skill 工具调用;fix 主会话产 FixCompleted 前直读本文,是为理解双层验证关系。.

Towow-ai/Flowness · 74 tokens

review

M-1.5 review skill — 在 patch 跟 contract 之间找 finding,produce Finding 一等对象。.

Towow-ai/Flowness · 26 tokens

dependency-analyze

从 task 的 read/write set + concept statemachine 推导 6 种依赖类型的提案。主 planner 决定边的真实性。派它时只给 read/write set 与疑点、不给预期边集;已有预判逐条标「待复核」交它取证。.

Towow-ai/Flowness · 70 tokens

execution-self-check

Pre-submit 自检——envelope 提交 commit gate 前必跑。独立 OPUS fork 逐项判 blocking checks(清单以 dispatch prompt 注入为准),executor 不能 self-assess(运动员不当裁判)。.

Towow-ai/Flowness · 55 tokens

fix

修复者 — 把一条被发现的问题(finding)按它的闭合合约修干净,修一个不制造下一个。产 FixProposed + 临时的 FixCompleted,不自判问题关闭(那是复查的权)。当 daemon 派一条 finding 来修、或需要闭合一个已发现的问题时用,即使只说"修一下这个 finding""把这个问题闭合"也触发。调用名就是 fix(Skill 工具)或 /fix(命令),没有 harness: 之类的前缀。.

Towow-ai/Flowness · 129 tokens