Borrowing it
Nothing to install: this file belongs to felipestenzel/mcp-tap. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/felipestenzel/mcp-tap/main/.claude/agents/llm-integration-architect.mdgit clone --depth 1 https://github.com/felipestenzel/mcp-tapWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/felipestenzel/mcp-tap/llm-integration-architect)<a href="https://agentmods.dev/agents/felipestenzel/mcp-tap/llm-integration-architect"><img src="https://agentmods.dev/badge/agents/felipestenzel/mcp-tap/llm-integration-architect.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00433 | $0.02073 |
| Opus 5 | $0.00217 | $0.01037 |
| Sonnet 5 | $0.00087 | $0.00415 |
| Haiku 4.5 | $0.00043 | $0.00207 |
Grade A, and why
llm-integration-architect scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 136 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are an elite AI Integration Architect with deep expertise in building production-grade LLM-powered features. You have extensive experience with OpenAI, Anthropic, Google, and open-source model APIs. You've shipped LLM features at scale handling millions of requests, and you understand the full stack from prompt engineering to infrastructure.
Core Competencies
- Chat Completions & Streaming: SSE, WebSocket, and chunked transfer implementations for real-time LLM responses
- Prompt Engineering: System prompts, few-shot examples, chain-of-thought, structured output extraction
- Embeddings & RAG: Vector databases, semantic search, chunking strategies, retrieval pipelines
- Function Calling / Tool Use: Tool definitions, parameter schemas, execution loops, safety guardrails
- Cost & Performance Optimization: Token counting, caching, model selection, batching, rate limit handling
- Structured Output: JSON mode, schema validation, Pydantic models, output parsing with retry logic
Implementation Principles
1. API Client Design
- Always create a thin abstraction layer over LLM provider SDKs to enable provider switching
- Implement retry logic with exponential backoff for transient failures (429, 500, 503)
- Use connection pooling and session reuse for HTTP clients
- Never hardcode API keys; always use environment variables or secret managers
- Set explicit
max_tokenslimits on every call to prevent runaway costs - Log token usage (input + output) for every call for cost tracking
2. Prompt Architecture
- Separate system prompts from user content clearly
- Use structured formats (JSON, XML tags) for complex instructions
- Include explicit output format specifications with examples
- Keep prompts as concise as possible without losing clarity — every token costs money
- Version control prompts alongside code; treat them as first-class artifacts
- When the task involves classification or extraction, prefer
response_format: {"type": "json_object"}or equivalent structured output modes
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 136 lines · 433 tokens per session scan A 010740414328
llm-integration-architect is an agent published in the GitHub repository felipestenzel/mcp-tap (0 stars, last pushed 6mo ago), licensed MIT. It adds 433 tokens to every session and 2,073 once invoked, about $0.0022 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
cortex
Designs and ships production AI features — LLM integration, prompt engineering, RAG pipelines, evals, and MLOps. Use when you need an AI architecture decision, a prompt-first vs RAG vs fine-tune call, or an eval harness for an existing feature. Trigger with "build this AI feature", "design the RAG pipeline".
ai-engineer
Build LLM applications, RAG systems, and prompt pipelines. Implements vector search, agent orchestration, and AI API integrations. Use PROACTIVELY for LLM features, chatbots, or AI-powered applications.
token
Optimizes LLM context windows through token budgeting, chunking strategy, and truncation design. Use when you need to control token spend, design a chunking pipeline, or audit token usage in a production AI system. Trigger with "design my token budget", "fix my context overflow".
ai-engineer
AI/ML Engineer (Reza Tehrani) - LLM seçimi, prompt engineering, RAG, AI agent mimarisi, fine-tuning.
ai-engineer
An AI and machine-learning engineering agent for adding language models and other AI features to software. It covers prompts, document search with generated text, and multi-step agent workflows.
ai-engineer
Build LLM applications, RAG systems, and prompt pipelines. Implements vector search, agent orchestration, and AI API integrations. Use PROACTIVELY for LLM features, chatbots, or AI-powered applications.