ai-engineering agents

336 tagged ai-engineering, measured the same way as everything else here.

Browse within: artificial-intelligence 194awesome-list 178benchmark 72ai-coding 67paper 46hackernews 37dataset 28models 26github-repo 25Autonomous Agents 24llmops 23Workflows 20web-crawled 20ai-development 19

joris887/exosuit

Agent Claude Code

Part of exosuit

Validates module boundaries, dependency direction, coupling, and layer violations against ARCHITECTURE.md. Reports only findings with confidence >= 80.

not rated 4 14d ago A 35 tokens original MIT

research-analyst

26

joris887/exosuit

Agent Claude Code

Part of exosuit

Deep web research agent for a specific sub-question. Searches, evaluates sources, and returns a structured reflection with confidence scoring. Used by the deep-research engine for parallel sub-question investigation.

not rated 4 14d ago A 43 tokens original MIT

spec-reviewer

27

joris887/exosuit

Agent Claude Code

Part of exosuit

Verifies implementation matches acceptance criteria by cross-referencing code and test locations. Validates story format and Definition of Ready compliance. Simple PASS/FAIL classification per criterion.

not rated 4 14d ago A 38 tokens original MIT

arnabdeypolimi/claude_code_setup

Agent Claude Code

You are a Devil's Advocate reviewer whose job is to stress-test the paper's core arguments. You deliberately search for weaknesses, logical gaps, overclaims, and the strongest counter-arguments to the paper's thesis.

not rated 4 3mo ago A 0 tokens original MIT

arnabdeypolimi/claude_code_setup

Agent Claude Code

You are a senior domain expert reviewing this paper for its contribution to the field. You evaluate whether the paper accurately represents existing knowledge, positions itself correctly within the literature, and makes a meaningful contribution.

not rated 4 3mo ago A 0 tokens original MIT

arnabdeypolimi/claude_code_setup

Agent Claude Code

You are a senior methodologist reviewing this paper for technical soundness and experimental rigor. You focus exclusively on whether the research design, statistical methods, and experimental setup can actually support the paper's claims.

not rated 4 3mo ago A 0 tokens original MIT

code-reviewer

31

SyloRei/claude-godmode

Agent

Part of claude-godmode

Code-level reviewer. Use for: checking whether a change does the thing RIGHT — bugs, edge cases, security, performance, readability, and pattern violations in the implementation. Catches 'thing wrong' errors. Read-only.

not rated 3 2mo ago A 49 tokens original MIT

executor

32

SyloRei/claude-godmode

Agent

Part of claude-godmode

Use when /build N spawns per-step implementation: implement one PLAN.md step in an isolated worktree, run the quality gates, and make one atomic commit. Brief-driven and gate-aware — unlike @writer (general-purpose), this agent works a single plan step against its brief and stops.

not rated 3 2mo ago A 61 tokens original MIT

finding-skeptic

33

SyloRei/claude-godmode

Agent

Part of claude-godmode

Adversarial skeptic that tries to REFUTE a single recorded review finding against the unit diff; returns UPHELD / REFUTED-not-real / REFUTED-over-rated. Read-only.

not rated 3 2mo ago A 44 tokens original MIT

README

34

Enovatr-Labs/SpecRoute

Agent Claude Code

This directory contains the 12 Claude Code agents that implement the SpecRoute skeleton. They are project-internal - they build SpecRoute itself, not consumer-facing templates that ship as artifacts.

not rated 3 +1 1mo ago A 0 tokens original Apache-2.0

Enovatr-Labs/SpecRoute

Agent Claude Code

Use when checking whether SpecRoute's claims about the outside world are still true - vendor config surfaces, hook event taxonomies, frontmatter contracts, version anchors, transition dates, and counts that must match disk. Owns the currency cycle, not the prose. Runs before releases, after any vendor ships a breaking…

not rated 3 +1 1mo ago A 120 tokens original Apache-2.0

Enovatr-Labs/SpecRoute

Agent Claude Code

Use proactively before any commit and whenever new content is added. Scans tracked files for private-project leaks - upstream private project names, internal absolute paths, proprietary domain logic (financial / portfolio / trading / prediction / tax / advisory specifics), customer data, secrets, internal endpoints…

not rated 3 +1 1mo ago A 125 tokens original Apache-2.0

architect

37

bzantium/meta-harness

Agent

Part of meta-harness

Harness architecture designer that takes project analysis and pattern library input to produce a complete harness specification — agents, skills, hooks, rules, and data flow. Uses opus for deep reasoning about optimal agent team composition.

not rated 2 5mo ago A 43 tokens original MIT

scout

38

bzantium/meta-harness

Agent

Part of meta-harness

Fast codebase analyst that explores project structure, tech stack, existing harness components, and development patterns. Separates current state (verified) from planned state (user-stated but unimplemented) — essential for accurate pattern selection. Spawned by create-harness and update-harness skills.

not rated 2 5mo ago A 60 tokens original MIT

analyzer

39

FarzamMohammadi/dev-toolbox

Agent Claude Code

Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.

not rated 1 1mo ago A 0 tokens

grader

41

FarzamMohammadi/dev-toolbox

Agent Claude Code

Evaluate expectations against an execution transcript and outputs.

not rated 1 1mo ago A 0 tokens

arxiv-2301-12569

47

SAIRAMANALADI/vybe-intelligence-vault

Agent

Agent "arxiv-2301-12569" from SAIRAMANALADI/vybe-intelligence-vault, covering a mental model based framework of trust, summary, why it matters, paper metadata and key topics & tags.

not rated 21 today A 0 tokens

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: