Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Shogo1222/socratic --skill maieuticgit clone --depth 1 https://github.com/Shogo1222/socraticWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/shogo1222/socratic/maieutic)<a href="https://agentmods.dev/skills/shogo1222/socratic/maieutic"><img src="https://agentmods.dev/badge/skills/shogo1222/socratic/maieutic/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/shogo1222/socratic/maieutic"><img src="https://agentmods.dev/badge/skills/shogo1222/socratic/maieutic.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00131 | $0.03514 |
| Opus 5 | $0.00066 | $0.01757 |
| Sonnet 5 | $0.00026 | $0.00703 |
| Haiku 4.5 | $0.00013 | $0.00351 |
Grade A, and why
maieutic scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 227 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Maieutic
Turn a code diff into a small set of human decisions and a validated, executable unit-test contract. Treat code as evidence of intent, never as the final specification.
Required references
Read references/intent-contract.md before recording decisions. In standalone use only, validate contract artifacts with references/intent-contract.schema.json; under a Socratic-hosted run, never open schema files — the Contract starts from the Runner's scaffold-contract document, its editable_fields and field_guide carry the authoritative structure, and staging performs the validation. Read references/qa-techniques.md when selecting test cases.
When receiving or preparing a proven-test handoff, also read Proven Test Handoff. Elenchus owns its mutation evidence and patch artifact; Maieutic owns whether its mapped expectations are confirmed and still current.
Operating rules
- Ask only when different reasonable answers change an observable expectation or important side effect.
- Do not ask for facts discoverable from repository instructions, issues, authoritative documentation, call sites, history, or reviewed evidence.
- Distinguish observed behavior, inferred intent, confirmed intent, and unresolved intent.
- Prefer a concrete behavior comparison over an abstract specification question.
- Optimize for human review cost as well as risk. Make each question quick to answer.
- Do not change production behavior to make a test pass unless the user separately authorizes a fix.
- Add tests only for confirmed expectations. Never freeze a suspected bug into a regression test.
Human decision interaction
When a decision changes an important observable oracle:
- Prefer the host's structured user-question tool when available —
AskUserQuestionon Claude Code,request_user_inputon Codex. - Ask one to three questions per batch, each with two or three mutually exclusive options and single selection by default.
- Give each option a label and a one-sentence observable consequence, state the oracle the answer changes, and mark a recommended option when one exists.
- Always allow a free-form answer; the specification owner may state a different expectation.
- When no structured tool is available, render the same question as copyable Markdown with lettered options.
- Ask from the main agent only; structured question tools are unavailable in subagents. Subagents may investigate, run tests, and execute mutations, and must return open decisions to the main agent.
- Persist every answer and its provenance in the Intent Contract before acting on it.
What ships with it
4 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 227 lines · 131 tokens per session scan A e6278b8da408
maieutic is a skill published in the GitHub repository Shogo1222/socratic (5 stars, last pushed 1mo ago), licensed MIT. It adds 131 tokens to every session and 3,514 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
test-review
You are an expert DataHub test reviewer. Your role is to evaluate pytest smoke tests against established testing standards, identify issues, and provide actionable feedback.
grade-tests
Grade specified test methods individually and produce a concise PR-ready table with each fully qualified test name, an A-F grade, score band, and one-line note. USE FOR per-test feedback on a curated list such as new or modified tests in a pull request, not a suite-wide audit. Polyglot: .NET, Python, TS/JS, Java, Go…
go-testing
Trigger: Go tests, go test coverage, Bubbletea teatest, golden files. Apply focused Go testing patterns.
quality-checklist
Validate implementation quality through custom checklists, scoring against constitution standards, specification coverage, and producing remediation recommendations.
brooks-test
Test quality review drawing on twelve classic engineering books — with primary focus on xUnit Test Patterns, The Art of Unit Testing, How Google Tests Software, and Working Effectively with Legacy Code — that diagnoses structural problems in an existing test suite: brittleness, mock abuse, coverage illusions, slow…
refactor
Refactors code for quality and maintainability. Triggers: refactor, clean up, restructure, improve code, modernize.