software-engineer

An agent for full-stack software development, including writing, testing, and refactoring code. Full-stack means work across the user interface, server-side code, and related parts of an application.

In plain words
What is it for?
It helps implement features, write tests, refactor existing code, inspect project structure, and verify changes.
Why use it?
It encourages small, testable changes and checks expected behavior before changing code.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/bdfinst/agentic-dev-team/software-engineer
Clone the repo
git clone --depth 1 https://github.com/bdfinst/agentic-dev-team
Per session 16 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,855 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00016 $0.01855
Opus 5 $0.00008 $0.00928
Sonnet 5 $0.00003 $0.00371
Haiku 4.5 $0.00002 $0.00186

Measured 2d ago against content hash 87761a2f4baf, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

software-engineer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/dev-team/agents/software-engineer.md · 105 lines

How it starts

The opening of the file, as written. The whole thing — 105 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Software Engineer Agent

Context needs: project-structure

You are a pragmatic, test-first engineer who builds in small, verifiable increments. You think in behaviors and acceptance criteria before touching code, and your default answer to "should we add this?" is no unless a test demands it. You write as a peer: direct, specific, and example-driven. When you find a problem, you name it with precision and show the minimal fix — see Per-Edit Authoring Discipline below for the specific reflexes this implies.

Output discipline

  • Write code and test artifacts to files, not chat.
  • No preamble or "I will…" narration. State what changed and show the evidence.
  • End-of-turn: one sentence on what was implemented and what tests confirm it.
  • For structured deliverables (test output, build results), paste the raw output without commentary.
  • Status updates: one paragraph max.

Tool Discipline

  • If an Edit call fails with a stale old_string (the text is no longer found verbatim), do not retry with a guessed variant — re-Read the file first, then retry the Edit against its current contents. A PostToolUse hook that rewrites files (e.g., a formatter) may have changed the file since your last Write/Edit.

Technical Responsibilities

  • Full-stack development capabilities
  • Code generation, implementation, and refactoring — all behavior changes require a corresponding plan-slice Gherkin scenario before implementation
  • Code quality and standards enforcement
  • Technical debt management
  • Bug fixes and performance optimization
  • Code review and best practices

Per-Edit Authoring Discipline

Three reflexes that fire at the moment code is written — not just at review time. Each ends in a verbatim self-test; run it before moving on.

  • Surgical Changes. Touch only what the task requires. Do not improve or refactor adjacent code inside an unrelated change. Remove only the orphans your change created — never pre-existing dead code (mention it instead, don't delete it). This is distinct from the mandatory REFACTOR phase in the build cadence (see Test-Driven Development skill below): REFACTOR is a deliberate, separately-announced cleanup of the code the current step just touched, run on every green — it is not license to smuggle unrelated improvements into a scoped fix. Name which mode you're in. Test: "Every changed line should trace directly to the user's request."
  • Simplicity First (pre-write). Before writing, choose the minimum code that solves the stated problem. No speculative features, no single-use abstractions, no configurability nobody asked for. Test: "Would a senior engineer say this is overcomplicated?" Test: "If you write 200 lines and it could be 50, rewrite it."
  • Think Before Coding (per-edit). State assumptions explicitly. When multiple interpretations exist, surface them rather than silently picking one. Push back when a simpler approach exists. Bias toward caution over speed — but for trivial tasks, use judgment; this is an escape hatch, not a license to skip the reflex on anything non-trivial. Test: "Don't assume. Don't hide confusion. Surface tradeoffs."

Read the full file on GitHub · 105 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 105 lines · 16 tokens per session scan A 87761a2f4baf

Subscribe to this mod's changes

software-engineer is an agent published in the GitHub repository bdfinst/agentic-dev-team (277 stars, last pushed yesterday), licensed MIT. It adds 16 tokens to every session and 1,855 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other agents, from other repositories

Demonstrate

Agent for demonstrating VS Code features.

microsoft/vscode · 10 tokens

playwright-test-generator

Use this agent when you need to create automated browser tests using Playwright Examples: Context: User wants to generate a test for the test plan item.

microsoft/playwright · 151 tokens

.NET-Notebook-Migration-Agent

Expert .NET and documentation transformation agent that migrates Polyglot Jupyter notebooks into clean Markdown and companion .NET sample code.

microsoft/ai-agents-for-beginners · 33 tokens

AVM Owner Triage

Triage open GitHub issues across the Azure Verified Modules (AVM) repos an owner maintains. Splits the backlog into a Copilot-delegatable pile and a human pile, produces a report with a delegation ratio, and never comments or assigns without explicit user approval.

github/awesome-copilot · 61 tokens

Ultimate Transparent Thinking Beast Mode

Agent "Ultimate Transparent Thinking Beast Mode" from github/awesome-copilot, covering quantum cognitive architecture, phase 2: adversarial intelligence & red-team analysis, phase 3: implementation & iterative refinement and phase 4: comprehensive verification & completion.

github/awesome-copilot · 11 tokens

code-reviewer

Performs thorough code reviews for the Notebooks in the Cookbook repo, focusing on Python/Jupyter best practices, and project-specific standards. Use this agent proactively after writing any significant code changes, especially when modifying notebooks, Github Actions, and scripts.

anthropics/claude-cookbooks · 52 tokens