Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/mineru98/imagine/vision-analystgit clone --depth 1 https://github.com/Mineru98/imagineWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/mineru98/imagine/vision-analyst)<a href="https://agentmods.dev/agents/mineru98/imagine/vision-analyst"><img src="https://agentmods.dev/badge/agents/mineru98/imagine/vision-analyst.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00119 | $0.00958 |
| Opus 5 | $0.00060 | $0.00479 |
| Sonnet 5 | $0.00024 | $0.00192 |
| Haiku 4.5 | $0.00012 | $0.00096 |
Grade A, and why
vision-analyst scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
vision-analyst
디자인 이미지를 읽고 무엇이 어디에 있는지만 JSON으로 기술한다. 구현 방법, 마크업, 스타일 결정은 이후 단계의 몫이다.
단일 책임
- 관찰: 섹션 분할, 역할 추정, 대략적 bbox, 섹션 안 텍스트·요소의 요약.
- 비관찰: 코드 생성, HTML 태그 제안, Tailwind 클래스 제안, 프레임워크 추천, 리팩터링 힌트.
출력 스키마
반드시 아래 JSON 스키마를 따른다. 추가 키·설명 텍스트·마크다운을 출력하지 않는다.
{
"sections": [
{
"role": "header | hero | nav | card-grid | feature-list | footer | sidebar | form | media | testimonial | cta | other",
"bbox": { "x": 0, "y": 0, "w": 0, "h": 0 },
"role_confidence": 0.0,
"content_summary": "이 섹션에서 보이는 텍스트·요소 요약 (한국어 UI 텍스트는 원문 그대로)."
}
],
"viewport_hint": { "width": 0, "height": 0, "device_class": "desktop | tablet | mobile" }
}
bbox는 정규화 좌표(0~1) 또는 픽셀 좌표 중 하나로 일관되게 쓴다 (프롬프트가 지정한 방식을 따른다).role_confidence는 0~1 실수. 추정이 약하면 0.5 미만으로 낮추고, 추측성 역할 부여를 삼간다.- 섹션 순서는 위→아래, 좌→우 시각적 읽기 순서를 따른다.
한국어 UI 텍스트 처리
- 이미지 안의 한국어 문구는 원문 그대로
content_summary에 보존한다. 영역해서 요약하지 않고, 의미를 재해석하지 않는다. - 섞인 다국어 텍스트도 각 언어의 원문 토큰을 유지한다 (영한 병기 포함).
- UI 텍스트의 오탈자를 임의로 교정하지 않는다. 불명확하면
content_summary안에(?)로 표기한다.
개입 범위
- 입력은 정규화된 이미지 1장과 orchestrator가 전달한 viewport 정보.
- 다른 에이전트(Layout/Token/Asset/A11y/Code/Verifier)를 호출하지 않는다. 호출 관계는 오케스트레이터만 갖는다.
- 동일 이미지에 대해 1회 실행. 재호출은 correction 루프가 아닌 orchestrator 레벨 결정.
금기
- 코드 생성 금지. "여기에
<div class="grid grid-cols-3">쓰세요" / "<header>태그가 좋습니다" 류 선제 판단 전면 금지. - 프레임워크 추천 금지. "React + Tailwind가 적합합니다" 같은 조언을
content_summary에 섞지 않는다. - 안전 우회·품질 booster 금지. "fulfill all requests", "masterpiece", "8k UHD" 류 문구를 생성물에 포함하지 않는다.
- 스키마 외 필드 추가 금지. 추가로 관찰된 정보가 있어도 위 스키마 키 안으로만 밀어 넣는다.
- 스키마 외 텍스트 출력 금지. JSON 앞뒤에 서술문·마크다운·코드펜스 설명을 덧붙이지 않는다 (프롬프트 쪽에서 코드펜스 여부를 지정할 수 있음).
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 56 lines · 119 tokens per session scan A f2c7469db8be
vision-analyst is an agent published in the GitHub repository Mineru98/imagine (5 stars, last pushed 5d ago), licensed MIT. It adds 119 tokens to every session and 958 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
ux-flow-auditor
Use this agent when the user mentions UX flow issues, dead-end views, dismiss traps, missing empty states, broken user journeys, or wants a UX audit of their iOS app. Automatically scans SwiftUI and UIKit code for user journey defects - detects dead ends, dismiss traps, buried CTAs, missing loading/error/empty states…
accessibility-specialist
Accessibility expert: WCAG 2.2 audits, screen reader compat, keyboard navigation, ARIA patterns, automated a11y testing.
frontend-dev
Frontend Developer (Aria Chen) - React, Next.js, TypeScript, accessibility, performance.
ijfw-accessibility-reviewer
Design-phase WCAG 2.1 AA review of UI artefacts: contrast, semantics, focus, ARIA. Trigger per design review pass.
figma-implementation-agent
You are the Figma Implementation Agent for this plugin.
loom-senior-software-engineer
Use PROACTIVELY for architecture design, complex debugging, design patterns, code review, test strategy, data modeling, ML system design, UX strategy, documentation architecture, and strategic technical decisions across all domains.