dare-bench

Um verificador que roda casos de teste preparados para medir se correções resolvem problemas conhecidos. Ele usa resultados esperados e não usa inteligência artificial durante a medição.

In plain words
What is it for?
Serve para medir a taxa de correções bem-sucedidas, comparar resultados com uma referência e interromper a verificação quando há regressão acima do limite definido.
Why use it?
Ajuda a detectar regressões, ou seja, quando uma mudança faz o sistema voltar a falhar em algo que já funcionava.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/dewtech-technologies/dare-method/dare-bench
Any agent
npx skills add dewtech-technologies/dare-method --skill dare-bench
Clone the repo
git clone --depth 1 https://github.com/dewtech-technologies/dare-method

Made for: Claude Code, Codex.

Per session 36 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 162 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00036 $0.00162
Opus 5 $0.00018 $0.00081
Sonnet 5 $0.00007 $0.00032
Haiku 4.5 $0.00004 $0.00016

Measured 2d ago against content hash 858c4b9f2e43, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

dare-bench scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

implementations/antigravity/.agents/skills/dare-bench/SKILL.md · 22 lines

What it actually says

DARE Bench — harness de verificação

Roda fixtures versionadas com patches golden/errados; mede Fix·Rate e solve-rate. Determinístico — sem LLM.

Como rodar

dare bench --suite fixtures/bench --json
dare bench --suite fixtures/bench --baseline bench-baseline.json --fail-on-regression 3

Exit codes

  • 0 — ok (sem regressão vs baseline)
  • 1 — regressão de solve-rate > limiar
  • 2 — suite inválida
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 22 lines · 36 tokens per session scan A 858c4b9f2e43

Subscribe to this mod's changes

dare-bench is a skill published in the GitHub repository dewtech-technologies/dare-method (5 stars, last pushed 1mo ago), licensed MIT. It adds 36 tokens to every session and 162 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

write-test-plan

Generate a QA/UAT test plan from product specifications and task definitions, covering acceptance testing, integration flows, and exploratory testing. Unit tests are out of scope (handled by write-unit-tests skill).

andresharpe/dotbot · 43 tokens

implement-telegram-bot

Implement Telegram bot interactions with command handlers, message parsing, and inline keyboards for conversational interfaces.

andresharpe/dotbot · 23 tokens

setup-background-job

Set up scheduled background jobs using Quartz.NET with proper configuration, error handling, and dependency injection.

andresharpe/dotbot · 22 tokens

write-unit-tests

Write comprehensive unit tests with proper setup, assertions, and coverage of happy paths, edge cases, and error scenarios.

andresharpe/dotbot · 26 tokens

blazor-component-design

Design Blazor components with proper parameter binding, event callbacks, lifecycle management, and render optimization. Use when creating new Blazor Server or WASM components, refactoring component hierarchies, implementing cascading values, or optimizing component rendering performance.

andresharpe/dotbot · 53 tokens

hedgehog-planning-intake

Use on any core for first-run planning intake — Phase 0 runs the vendored BMAD-METHOD planning shelf, shared by every core, and Phase 1 (mining 04-prd.md into intent records plus the Add-ons/sync-and-remote-entities decision) is full-stack-app's and pwa-app's shared procedure — identical mechanics, a different…

skyf0xx/hedgehog · 374 tokens