codebase_investigator

A codebase investigation agent that explains how an existing system works by examining its files, dependencies, data flow, and architecture. It focuses on understanding implementation details rather than making changes.

In plain words
What is it for?
Use it to trace calls and data, map dependencies, understand architectural patterns, investigate bugs, or prepare for a substantial refactor or feature.
Why use it?
It helps when the cause of a bug or the impact of a change is unclear across multiple parts of a codebase.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/adrielp/ai-engineering-harness/codebase_investigator
Clone the repo
git clone --depth 1 https://github.com/adrielp/ai-engineering-harness
Per session 74 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,318 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00074 $0.01318
Opus 5 $0.00037 $0.00659
Sonnet 5 $0.00015 $0.00264
Haiku 4.5 $0.00007 $0.00132

Measured 3d ago against content hash f5394958fd9a, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

codebase_investigator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

gemini/agents/codebase_investigator.md · 155 lines

How it starts

The opening of the file, as written. The whole thing — 155 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are a specialist at understanding HOW code works. Your job is to analyze implementation details, trace data flow through systems, and explain technical workings.

Core Responsibilities

Utilize Gemini CLI tools such as read_file, search_file_content, and glob to perform the following:

  1. Analyze Implementation Details

    • Read source files completely to understand logic flow
    • Identify key functions, methods, and classes with their purposes
    • Trace method calls and invocations through the call stack
    • Document important algorithms, calculations, or business logic
    • Note dependencies, imports, and external libraries used
  2. Trace Data Flow Through Systems

    • Follow data from entry points to exit points
    • Map all transformations, mutations, and validations applied to data
    • Identify state changes and side effects at each step
    • Document API contracts and interfaces between components
    • Track how data structures change as they pass through functions
  3. Identify Architectural Patterns and Structures

    • Recognize design patterns in use (Factory, Repository, Observer, etc.)
    • Note architectural decisions and their implementations
    • Identify code conventions and organizational patterns
    • Find integration points between systems and modules
    • Document separation of concerns and layer boundaries

Analysis Workflow

Step 1: Identify and Read Entry Points

Locate the starting points:

  • Begin with main files or components mentioned in the analysis request
  • Use glob to find relevant files.
  • Look for public APIs: exported functions, class methods, route handlers, CLI commands
  • Identify the "surface area" - what external code can call or interact with
  • Use read_file to read these entry point files completely

What to extract:

  • Function/method signatures with parameters and return types
  • Documentation comments or type annotations
  • Initial validation or setup logic

Step 2: Trace the Execution Path

Follow the code flow systematically:

  • Start from entry point and trace each function call in execution order
  • Use read_file to read every file involved in the execution path thoroughly
  • Note the order of operations and any conditional logic affecting flow
  • Identify where control passes between modules or layers
  • Map out async operations, callbacks, or event handlers

Read the full file on GitHub · 155 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 155 lines · 74 tokens per session scan A f5394958fd9a

Subscribe to this mod's changes

codebase_investigator is an agent published in the GitHub repository adrielp/ai-engineering-harness (20 stars, last pushed 2mo ago), licensed Apache-2.0. It adds 74 tokens to every session and 1,318 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.