flake-check

A helper for finding why a software test fails only sometimes, such as because of an unreliable selector, timing, test data, or a real product bug.

In plain words
What is it for?
Use it with a test file, test name, failure log, stack trace, or description of the intermittent behavior to get a cause, confidence level, and concrete fix.
Why use it?
It separates test problems from product problems and points to evidence for the diagnosis, so intermittent failures are easier to fix.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/kint4/autoframe/flake-check
Any agent
npx skills add kint4/autoframe --skill flake-check
Clone the repo
git clone --depth 1 https://github.com/kint4/autoframe

Made for: Claude Code, Codex.

Per session 40 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 445 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00040 $0.00445
Opus 5 $0.00020 $0.00222
Sonnet 5 $0.00008 $0.00089
Haiku 4.5 $0.00004 $0.00044

Measured 2d ago against content hash 2398c0532238, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

flake-check scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/skills/flake-check/SKILL.md · 47 lines

What it actually says

/flake-check — Diagnose Flaky Tests

Type: Functional Description: Analyzes failing or unstable tests and diagnoses whether the cause is a flaky selector, a timing/race issue, test data, or a real product bug — then proposes a fix.


Input Format

Any of:

  • A path to a spec file or test name
  • A pasted failure log / stack trace
  • A description of the intermittent behavior

Output Format

A diagnosis report containing:

  • Classification: flaky selector / timing issue / test data / environment / real bug
  • Evidence: the specific lines or signals that point to the cause
  • Recommended fix: concrete code change following Autoframe conventions
  • Confidence: high / medium / low

Step-by-Step Instructions

  1. Read the relevant spec and its Page Object(s).
  2. Inspect the failure signal (log, trace, or description).
  3. Check for the common flake causes, in order:
    • Selectors: brittle CSS / page.$() / nth-based locators → recommend semantic getByRole/getByLabel.
    • Timing: hard waits (waitForTimeout), missing auto-waiting assertions, race conditions → recommend web-first expect assertions and removing fixed sleeps.
    • Test data: shared/mutable state, ordering dependencies → recommend factories and isolation.
    • Environment: network, auth token expiry, base URL.
    • Real bug: behavior is genuinely wrong → recommend /bug-from-failure.
  4. State the classification with evidence and a confidence level.
  5. Propose the fix as a concrete diff that follows POM and spec conventions.
  6. If it's a real bug, route the user to /bug-from-failure.

Rules

  • Prefer fixing the root cause over adding retries.
  • Never recommend waitForTimeout as a fix.
  • Always classify before proposing a change.
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 47 lines · 0 tokens per session scan A 2398c0532238

Subscribe to this mod's changes

flake-check is a skill published in the GitHub repository kint4/autoframe (6 stars, last pushed 2mo ago), licensed MIT. It adds 40 tokens to every session and 445 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

k6-load-testing

Comprehensive k6 load testing skill for API, browser, and scalability testing. Write realistic load scenarios, analyze results, and integrate with CI/CD.

LiHongwei-cn/lihongwei-cn · 35 tokens

qawolf-cli

Manage QA Wolf through the qawolf CLI. Use when asked to create, update, or list coverage requests, bug reports, or maintenance reports; start a run of flows or tags on the QA Wolf platform or read a run's results; list, set, or delete environment variables; manage environments, flows, or tags; request automation of…

qawolf/cli · 118 tokens

tokenless

Use when a task can be delegated through the globally installed Tokenless CLI without directly writing to the workspace; route it to a visible AI provider website to save agent tokens.

jazelly/tokenless · 37 tokens

tokenless-install

Install, upgrade, repair, and verify Tokenless, its agent skills, and local Playwright runtime. Use only when the user explicitly asks for installation, upgrade, repair, browser sign-in handoff, a failed doctor check, or an installation integrity check.

jazelly/tokenless · 56 tokens

agent-qa-authoring

Use when creating, editing, validating, or running agent-qa tests, suites, or hooks. Prefer agent-qa MCP tools, enforce canonical agent-qa IDs, and use the bundled schema reference to avoid hallucinated config keys or YAML fields.

vostride/agent-qa · 56 tokens

agent-qa-debug-fix

Use after an agent-qa run has failed and you need to debug, patch, and verify the issue using MCP evidence, logs, artifacts, and local code changes instead of generated fix suggestions.

vostride/agent-qa · 46 tokens