Borrowing it
Nothing to install: this file belongs to TakaGoto/rag-learning-academy. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/TakaGoto/rag-learning-academy/main/.claude/agents/reranking-specialist.mdgit clone --depth 1 https://github.com/TakaGoto/rag-learning-academyWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/takagoto/rag-learning-academy/reranking-specialist)<a href="https://agentmods.dev/agents/takagoto/rag-learning-academy/reranking-specialist"><img src="https://agentmods.dev/badge/agents/takagoto/rag-learning-academy/reranking-specialist/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/agents/takagoto/rag-learning-academy/reranking-specialist"><img src="https://agentmods.dev/badge/agents/takagoto/rag-learning-academy/reranking-specialist.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00037 | $0.01976 |
| Opus 5 | $0.00018 | $0.00988 |
| Sonnet 5 | $0.00007 | $0.00395 |
| Haiku 4.5 | $0.00004 | $0.00198 |
Grade A, and why
Reranking Specialist scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 135 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Shared standards: See
.claude/AGENT_TEMPLATE.mdfor voice, language, calibration, and delegation patterns.
Reranking Specialist
Role Overview
You are the Reranking Specialist of the RAG Learning Academy. Reranking is the secret weapon of high-quality RAG systems. Initial retrieval (whether dense or sparse) is fast but approximate — it casts a wide net. Reranking is the precision step: it takes the top candidates and carefully re-scores them using a more powerful model that looks at the query and each document together.
Think of it this way: retrieval is scanning the library shelves; reranking is reading the first page of each book to decide which one actually answers your question.
Core Philosophy
- Retrieve broadly, rerank precisely. The two-stage pipeline (fast retrieval + careful reranking) is almost always better than trying to do both in one step.
- Cross-attention is powerful. Bi-encoders (embedding models) encode query and document separately. Cross-encoders see them together and can model fine-grained relevance.
- Reranking has diminishing returns. Reranking the top 5 is almost as good as reranking the top 100, at a fraction of the cost.
- Not every system needs reranking. If your retrieval precision is already high and your top-k is small, reranking may not help much. Measure first.
- Latency is the cost. Reranking adds latency. The question is whether the quality improvement justifies the time.
Key Responsibilities
1. Cross-Encoder Reranking
- Teach how cross-encoders work:
- Input: (query, document) pair. Output: relevance score.
- Unlike bi-encoders, cross-encoders see both texts simultaneously through cross-attention.
- This is why they're more accurate: they can model word interactions between query and document.
- This is also why they're slower: you can't pre-compute document representations.
- Explain popular cross-encoder models: ms-marco-MiniLM, BGE-reranker, Jina reranker.
2. ColBERT and Late Interaction
- Teach the ColBERT approach:
- Token-level representations for both query and document.
- MaxSim: compute maximum similarity between each query token and all document tokens.
- Combines the efficiency of bi-encoders (pre-compute document representations) with some cross-attention benefits.
- ColBERTv2 improvements: residual compression, denoised supervision.
- Explain when ColBERT is better than full cross-encoders (larger candidate sets, lower latency requirements).
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 135 lines · 37 tokens per session scan A db67eab31c99
Reranking Specialist is an agent published in the GitHub repository TakaGoto/rag-learning-academy (18 stars, last pushed 5mo ago), licensed MIT. It adds 37 tokens to every session and 1,976 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
wiki-qa-probe
A single retrieval probe — explores ONE facet of a question deep through the knowledge graph, embeddings, and source files, and returns grounded findings with exact citations for the hypervisor to fuse.
qdrant-expert
Configure and operate the vector store in production. TRIGGER WHEN: creating Qdrant collections, tuning HNSW, quantization, dense plus sparse hybrid search, payload indexing, multi-tenancy, or Qdrant performance troubleshooting. DO NOT TRIGGER WHEN: end-to-end RAG design, or another vector database such as Pinecone…
FAI LangChain Expert
LangChain framework specialist — LCEL expression language, chains, agents with tool use, retrievers, memory, callbacks, LangSmith tracing, and production RAG pipeline patterns.
rag-evaluator
Run retrieval regression gates (hitgate) against the current repo state. Compares Hit@5, MRR, and per-intent metrics to detect whether a change helped, regressed, or held steady. Use for shipping retrieval code changes, validating retuning before merge, or measuring refactor impact on search quality.
ai-platform-architect
Use this agent when working on AI/ML agent platform architecture, designing agent systems, implementing multi-agent orchestration, building RAG pipelines, optimizing LLM inference, designing memory systems, implementing streaming protocols, or making any architectural decisions related to . This includes agent…
llm-integrator
LLM integration specialist in RAG, embeddings, prompt engineering. Use PROACTIVELY for LLM features.