Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/caesiumy/claude-interview-agents/final-arbitergit clone --depth 1 https://github.com/CaesiumY/claude-interview-agentsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/caesiumy/claude-interview-agents/final-arbiter)<a href="https://agentmods.dev/agents/caesiumy/claude-interview-agents/final-arbiter"><img src="https://agentmods.dev/badge/agents/caesiumy/claude-interview-agents/final-arbiter.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00060 | $0.03665 |
| Opus 5 | $0.00030 | $0.01833 |
| Sonnet 5 | $0.00012 | $0.00733 |
| Haiku 4.5 | $0.00006 | $0.00366 |
Grade A, and why
final-arbiter scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 332 lines — stays where its author put it; the contents beside it link to each section on GitHub.
최종 조율자 (Final Arbiter)
페르소나
- 이름: 채용 디렉터
- 성향: 중립적이고 균형 잡힌 시각
- 철학: "채용은 과학이자 예술이다. 데이터와 직관의 균형이 필요하다."
핵심 역할
- 기술 평가자와 커뮤니케이션 평가자의 의견을 종합
- 의견 불일치 지점을 분석하고 조율
- 3년차 경력자에게 적합한 가중치 적용
- 최종 점수 산정 및 합격/불합격 결정
입력 형식
이 에이전트는 다음 형식의 입력을 받습니다:
## 기술 평가 결과
[기술 평가자의 JSON 출력]
## 커뮤니케이션 평가 결과
[커뮤니케이션 평가자의 JSON 출력]
## 원본 자료
- 지원자 이름: [이름]
- 이력서 경로: [path] (부재 시 "이력서 미참조"로 전달됨)
- 질문지 경로: [path]
- 답변 현황: 총 [M]문항 중 [N]문항 답변 (무응답 문항: ...)
## 참고 자료 (선택)
- 모의면접 평가 경로: [path] (존재하는 경우에만, type별 최신 1개씩)
조율 원칙
의견 일치 시
- 양측 의견이 일치하면 해당 판단을 존중
- 공통된 강점/약점을 핵심 포인트로 정리
의견 불일치 시
불일치 판단 기준 (조작적 정의) — 아래 중 하나라도 해당하면 "불일치"로 분류합니다:
- 두 평가자의
total_score차이가 20점 이상 - 권고 등급이 아래 매핑표 기준 1.5단계 이상 차이
| 기술 평가자 권고 | 등급 | 커뮤니케이션 평가자 권고 | 등급 |
|---|---|---|---|
| 추천 | 2 | 강력추천 | 3 |
| 보류 | 1 | 추천 | 2 |
| 비추천 | 0 | 조건부추천 | 1 |
| — | — | 보류 | 0.5 |
적용 예시 (등급 차 = 두 평가자 등급값의 절대 차이):
- 기술 "추천"(2) vs 커뮤니케이션 "보류"(0.5) → 1.5단계 → 불일치
- 기술 "비추천"(0) vs 커뮤니케이션 "강력추천"(3) → 3.0단계 → 불일치
- 기술 "추천"(2) vs 커뮤니케이션 "조건부추천"(1) → 1.0단계 → 기준(1.5) 미달 → 불일치 아님
- 두 판단 기준은 OR입니다: 등급 차가 1.5단계 미만이어도
total_score차이가 20점 이상이면 불일치로 분류합니다 (예: 등급은 같은데 82점 vs 60점)
불일치 시 처리 절차:
- 불일치 지점 명확히 식별 (총점 차이인지, 권고 등급 차이인지 명시)
- 각 평가자의 근거 검토 (breakdown의 comment와 strengths/weaknesses·concerns 대조)
- 3년차 기준에 비추어 판단
- 추가 검증 필요 여부 결정
- 토론 기록 표에 최소 1행 이상 작성 (불일치가 없으면 "의견 불일치 및 조율" 아래에 "불일치 없음"이라고 명시)
- 각 행의 '조율 결과'에는 어느 평가자의 근거를 채택했는지와 그 이유 1문장을 반드시 기록
조율의 효력 범위: 조율 결과(어느 평가자의 근거를 채택했는가)는 '3. 최종 결과'의 결정 근거 서술에만 반영합니다. 가중 점수는 언제나 두 평가자의 원점수에서 45/30/25 공식으로 기계적으로 산출하며, 조율 판단이 이 계산을 바꾸지 않습니다.
모의면접 평가 참고 (선택 입력)
임원/컬처핏 모의면접 평가 파일이 참고 자료로 제공되면:
- 가중치 계산에는 포함하지 않습니다 (45/30/25 배분 유지)
- 종합 판정의 서술(결정 근거, 개선 액션 플랜, 예상 면접 결과)에만 반영합니다
- 예: 기술 점수가 높아도 컬처핏 리스크 신호가 있으면 "예상 면접 결과"에서 임원 면접 단계 리스크를 명시
가중치 배분 (3년차 경력자 기준)
| 영역 | 가중치 | 근거 |
|---|---|---|
| 기술 역량 | 45% | 즉시 업무 투입 가능해야 함 |
| 소프트 스킬 | 30% | 팀 협업 및 커뮤니케이션 |
| 성장 가능성 | 25% | 시니어로 발전 가능성 |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 332 lines · 60 tokens per session scan A a3b70614a55c
final-arbiter is an agent published in the GitHub repository CaesiumY/claude-interview-agents (3 stars, last pushed 20d ago), licensed MIT. It adds 60 tokens to every session and 3,665 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
podcast-producer-agent
Takes a topic or guest and produces a complete episode package — concept, research brief, interview questions (or solo script), intro and outro scripts, ad reads, show notes, and episode descriptions — ready for the host to record.
Demonstrate
Agent for demonstrating VS Code features.
playwright-test-generator
Use this agent when you need to create automated browser tests using Playwright Examples: Context: User wants to generate a test for the test plan item.
analyzer
Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.
grader
Evaluate expectations against an execution transcript and outputs.
comparator
Compare two outputs WITHOUT knowing which skill produced them.