mcp-tap: Agent for Claude Code

.claude/agents/llm-integration-architect.md

llm-integration-architect is an agent for Claude Code from felipestenzel/mcp-tap. It costs 433 tokens per session (2,073 once invoked), scanned A, original, MIT.

An engineering guide for adding large-language-model features to an application. Large language models are software systems that generate or interpret text, code, and other data through an API.

In plain words
What is it for?
Use it to build chat and streaming responses, prompts, embeddings and retrieval systems, function calling, structured JSON output, token tracking, caching, batching, and rate-limit handling.
Why use it?
It helps structure integrations so provider changes, retries, limits, costs, and output errors are handled consistently. It also covers how to connect model responses to application functions.

Agent for Claude Code

Written for Claude Code: installed under .claude/. Also seen: model in frontmatter.

This is felipestenzel/mcp-tap's own configuration. It tells Claude Code how to work on mcp-tap itself, so it is not a mod to install elsewhere. Copy it as a starting point and replace the rules that are about this project. Everything mcp-tap configures →

Not installable: its command points at a path on the author’s own machine, so it runs nowhere else. The line is /Users/felipestenzel/Documents/project_cswd/.claude/agent-memory/llm-integration-architect/.

Reuse

Borrowing it

Nothing to install: this file belongs to felipestenzel/mcp-tap. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.

Copy the file
curl -O https://raw.githubusercontent.com/felipestenzel/mcp-tap/main/.claude/agents/llm-integration-architect.md
Clone the repo
git clone --depth 1 https://github.com/felipestenzel/mcp-tap

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for llm-integration-architect

README.md
[![agentmods](https://agentmods.dev/badge/agents/felipestenzel/mcp-tap/llm-integration-architect.svg)](https://agentmods.dev/agents/felipestenzel/mcp-tap/llm-integration-architect)
Your own site
<a href="https://agentmods.dev/agents/felipestenzel/mcp-tap/llm-integration-architect"><img src="https://agentmods.dev/badge/agents/felipestenzel/mcp-tap/llm-integration-architect.svg" alt="Measured on agentmods" height="20"></a>
Per session 433 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 2,073 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00433 $0.02073
Opus 5 $0.00217 $0.01037
Sonnet 5 $0.00087 $0.00415
Haiku 4.5 $0.00043 $0.00207

Measured 7d ago against content hash 010740414328, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-08, from the pricing page.

Security

Grade A, and why

llm-integration-architect scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/agents/llm-integration-architect.md · 136 lines

How it starts

The opening of the file, as written. The whole thing — 136 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are an elite AI Integration Architect with deep expertise in building production-grade LLM-powered features. You have extensive experience with OpenAI, Anthropic, Google, and open-source model APIs. You've shipped LLM features at scale handling millions of requests, and you understand the full stack from prompt engineering to infrastructure.

Core Competencies

  • Chat Completions & Streaming: SSE, WebSocket, and chunked transfer implementations for real-time LLM responses
  • Prompt Engineering: System prompts, few-shot examples, chain-of-thought, structured output extraction
  • Embeddings & RAG: Vector databases, semantic search, chunking strategies, retrieval pipelines
  • Function Calling / Tool Use: Tool definitions, parameter schemas, execution loops, safety guardrails
  • Cost & Performance Optimization: Token counting, caching, model selection, batching, rate limit handling
  • Structured Output: JSON mode, schema validation, Pydantic models, output parsing with retry logic

Implementation Principles

1. API Client Design

  • Always create a thin abstraction layer over LLM provider SDKs to enable provider switching
  • Implement retry logic with exponential backoff for transient failures (429, 500, 503)
  • Use connection pooling and session reuse for HTTP clients
  • Never hardcode API keys; always use environment variables or secret managers
  • Set explicit max_tokens limits on every call to prevent runaway costs
  • Log token usage (input + output) for every call for cost tracking

2. Prompt Architecture

  • Separate system prompts from user content clearly
  • Use structured formats (JSON, XML tags) for complex instructions
  • Include explicit output format specifications with examples
  • Keep prompts as concise as possible without losing clarity — every token costs money
  • Version control prompts alongside code; treat them as first-class artifacts
  • When the task involves classification or extraction, prefer response_format: {"type": "json_object"} or equivalent structured output modes

Read the full file on GitHub · 136 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 7d ago First seen · 136 lines · 433 tokens per session scan A 010740414328

Subscribe to this mod's changes

llm-integration-architect is an agent published in the GitHub repository felipestenzel/mcp-tap (0 stars, last pushed 6mo ago), licensed MIT. It adds 433 tokens to every session and 2,073 once invoked, about $0.0022 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other agents, from other repositories

cortex

Designs and ships production AI features — LLM integration, prompt engineering, RAG pipelines, evals, and MLOps. Use when you need an AI architecture decision, a prompt-first vs RAG vs fine-tune call, or an eval harness for an existing feature. Trigger with "build this AI feature", "design the RAG pipeline".

jeremylongshore/tons-of-skills-marketplace · 75 tokens

ai-engineer

Build LLM applications, RAG systems, and prompt pipelines. Implements vector search, agent orchestration, and AI API integrations. Use PROACTIVELY for LLM features, chatbots, or AI-powered applications.

davepoon/buildwithclaude · 48 tokens

token

Optimizes LLM context windows through token budgeting, chunking strategy, and truncation design. Use when you need to control token spend, design a chunking pipeline, or audit token usage in a production AI system. Trigger with "design my token budget", "fix my context overflow".

jeremylongshore/tons-of-skills-marketplace · 60 tokens

ai-engineer

AI/ML Engineer (Reza Tehrani) - LLM seçimi, prompt engineering, RAG, AI agent mimarisi, fine-tuning.

vibeeval/vibecosystem · 36 tokens

ai-engineer

An AI and machine-learning engineering agent for adding language models and other AI features to software. It covers prompts, document search with generated text, and multi-step agent workflows.

CronusL-1141/AI-company · 41 tokens

ai-engineer

Build LLM applications, RAG systems, and prompt pipelines. Implements vector search, agent orchestration, and AI API integrations. Use PROACTIVELY for LLM features, chatbots, or AI-powered applications.

echoVic/blade-code · 48 tokens