Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add rules/xclaw-bot/benchmark-task-authoring/60-pipelinegit clone --depth 1 https://github.com/Xclaw-bot/benchmark-task-authoringWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00964 | $0.00964 |
| Opus 5 | $0.00482 | $0.00482 |
| Sonnet 5 | $0.00193 | $0.00193 |
| Haiku 4.5 | $0.00096 | $0.00096 |
Grade A, and why
60-pipeline scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 82 lines — stays where its author put it; the contents beside it link to each section on GitHub.
description: Submission pipeline: fork/PR flow, the automated check stages, pass@2 and pass@5 economics, reading feedback, revision limits. alwaysApply: false
Pipeline and iteration
Flow
Proposal (gated before you build) → fork → build on submission → local
oracle/nop → PR → validity check → pass@2 at your timeout → Automated Review
(all Blocking Issues fixed) → pass@5.
gh pr create --repo <program-org>/<task-repo> --fill
Iterate by pushing to the same branch — checks re-run and sticky comments update
in place. Use the house PR description shape: one-sentence problem, numbered
success criteria mirroring instruction.md, calibration results (oracle 1.0 /
nop <1.0), how to run, notes on anything you had to interpret.
The difficulty gate
Acceptance bar: pass@5 ≤ 2/5.
| pass@5 | Meaning | Outcome |
|---|---|---|
| 0/5 valid failures | Fully stumped, oracle still solves | ✅ strongest |
| 0/5 invalid failures | Timeout / agent / verifier error / unfair prompt | ❌ fix the cause |
| 1–2/5 | Solvable and genuinely hard | ✅ |
| 3–5/5 | Too easy | ❌ |
A valid failure is the model finishing and being wrong on a fair problem.
Timeouts and agent/verifier errors never count. Gate formula:
(good valid fails) + (soft-timeout fails) >= 3 and good valid >= 1 —
soft timeouts count, infra/in-progress timeouts do not.
Never game the score. Lowering the timeout or padding busywork makes the task harder to finish, not harder to get right; reviewers send it back.
Economics — treat runs as scarce
- Pass@2 is 6 runs per day, per fellow, per task. Never spend one before the oracle is green locally.
- A push cancels queued runs for the same PR (per-PR concurrency), so superseded runs auto-cancel — no quota wasted.
- Fork contributors cannot cancel upstream runs (
actions:write→ HTTP 404). - Pass@5 auto-starts only if pass@2 was valid and Automated Review passed.
- Review pipeline: max 2 revisions, then Holding-Rejection. Official guidance is to revise sent-back tasks before claiming new ones — they are closest to the bonus.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 82 lines · 964 tokens per session scan A 8682404c559b
60-pipeline is a cursor rule published in the GitHub repository Xclaw-bot/benchmark-task-authoring (2 stars, last pushed 18d ago), licensed MIT. It adds 964 tokens to every session, about $0.0048 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other cursor rules, from other repositories
angular-20
This rule provides comprehensive best practices and coding standards for Angular development, focusing on modern TypeScript, standalone components, signals, and performance optimizations.
dev-standard
Apache Superset development standards and guidelines for Cursor IDE.
cli-error-handling
CLI command error handling patterns.
prefer-direct-imports-over-module-mocks
Prefer extracting a testable core over vi.mock / vi.resetModules when unit tests need to reach production logic entangled with config, env, or singletons.
control-plane-descriptors
Control plane descriptor and instance implementation patterns.
family-instance-domain-actions
Family instance domain action implementation patterns.