Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/kint4/autoframe/flake-checknpx skills add kint4/autoframe --skill flake-checkgit clone --depth 1 https://github.com/kint4/autoframeWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00040 | $0.00445 |
| Opus 5 | $0.00020 | $0.00222 |
| Sonnet 5 | $0.00008 | $0.00089 |
| Haiku 4.5 | $0.00004 | $0.00044 |
Grade A, and why
flake-check scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
/flake-check — Diagnose Flaky Tests
Type: Functional Description: Analyzes failing or unstable tests and diagnoses whether the cause is a flaky selector, a timing/race issue, test data, or a real product bug — then proposes a fix.
Input Format
Any of:
- A path to a spec file or test name
- A pasted failure log / stack trace
- A description of the intermittent behavior
Output Format
A diagnosis report containing:
- Classification: flaky selector / timing issue / test data / environment / real bug
- Evidence: the specific lines or signals that point to the cause
- Recommended fix: concrete code change following Autoframe conventions
- Confidence: high / medium / low
Step-by-Step Instructions
- Read the relevant spec and its Page Object(s).
- Inspect the failure signal (log, trace, or description).
- Check for the common flake causes, in order:
- Selectors: brittle CSS /
page.$()/ nth-based locators → recommend semanticgetByRole/getByLabel. - Timing: hard waits (
waitForTimeout), missing auto-waiting assertions, race conditions → recommend web-firstexpectassertions and removing fixed sleeps. - Test data: shared/mutable state, ordering dependencies → recommend factories and isolation.
- Environment: network, auth token expiry, base URL.
- Real bug: behavior is genuinely wrong → recommend
/bug-from-failure.
- Selectors: brittle CSS /
- State the classification with evidence and a confidence level.
- Propose the fix as a concrete diff that follows POM and spec conventions.
- If it's a real bug, route the user to
/bug-from-failure.
Rules
- Prefer fixing the root cause over adding retries.
- Never recommend
waitForTimeoutas a fix. - Always classify before proposing a change.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 47 lines · 0 tokens per session scan A 2398c0532238
flake-check is a skill published in the GitHub repository kint4/autoframe (6 stars, last pushed 2mo ago), licensed MIT. It adds 40 tokens to every session and 445 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
k6-load-testing
Comprehensive k6 load testing skill for API, browser, and scalability testing. Write realistic load scenarios, analyze results, and integrate with CI/CD.
qawolf-cli
Manage QA Wolf through the qawolf CLI. Use when asked to create, update, or list coverage requests, bug reports, or maintenance reports; start a run of flows or tags on the QA Wolf platform or read a run's results; list, set, or delete environment variables; manage environments, flows, or tags; request automation of…
tokenless
Use when a task can be delegated through the globally installed Tokenless CLI without directly writing to the workspace; route it to a visible AI provider website to save agent tokens.
tokenless-install
Install, upgrade, repair, and verify Tokenless, its agent skills, and local Playwright runtime. Use only when the user explicitly asks for installation, upgrade, repair, browser sign-in handoff, a failed doctor check, or an installation integrity check.
agent-qa-authoring
Use when creating, editing, validating, or running agent-qa tests, suites, or hooks. Prefer agent-qa MCP tools, enforce canonical agent-qa IDs, and use the bundled schema reference to avoid hallucinated config keys or YAML fields.
agent-qa-debug-fix
Use after an agent-qa run has failed and you need to debug, patch, and verify the issue using MCP evidence, logs, artifacts, and local code changes instead of generated fix suggestions.