scene-vision-analyzer

scene-vision-analyzer is a skill for Claude Code, Codex from IvanYangYangXi/artclaw_bridge. It costs 133 tokens per session (1,058 once invoked), scanned A, original, MIT.

An image-analysis workflow that turns game concept art or screenshots into structured information for rebuilding a 3D scene.

In plain words
What is it for?
Use it to prepare data for 2D-to-3D reconstruction, then refine and check the analysis for missed objects, sizing, shadows, occlusion, and spatial consistency.
Why use it?
It helps convert a flat reference image into details about objects, positions, camera view, lighting, perspective, and depth.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to prepare data for 2D-to-3D reconstruction, then refine and check the analysis for missed objects, sizing, shadows, occlusion, and spatial consistency.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/ivanyangyangxi/artclaw_bridge/scene-vision-analyzer
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add IvanYangYangXi/artclaw_bridge --skill scene-vision-analyzer
Clone the repo
git clone --depth 1 https://github.com/IvanYangYangXi/artclaw_bridge

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for scene-vision-analyzer

README.md
[![agentmods](https://agentmods.dev/badge/skills/ivanyangyangxi/artclaw_bridge/scene-vision-analyzer/github.svg)](https://agentmods.dev/skills/ivanyangyangxi/artclaw_bridge/scene-vision-analyzer)
Your own site
<a href="https://agentmods.dev/skills/ivanyangyangxi/artclaw_bridge/scene-vision-analyzer"><img src="https://agentmods.dev/badge/skills/ivanyangyangxi/artclaw_bridge/scene-vision-analyzer/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for scene-vision-analyzer

Your own site · 80×15
<a href="https://agentmods.dev/skills/ivanyangyangxi/artclaw_bridge/scene-vision-analyzer"><img src="https://agentmods.dev/badge/skills/ivanyangyangxi/artclaw_bridge/scene-vision-analyzer.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 133 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,058 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00133 $0.01058
Opus 5 $0.00067 $0.00529
Sonnet 5 $0.00027 $0.00212
Haiku 4.5 $0.00013 $0.00106

Measured 12d ago against content hash 69402bd301e4, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-11, from the pricing page.

Security

Grade A, and why

scene-vision-analyzer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.

The scan reads SKILL.md. This mod also ships 1 executable file (scripts/scene_analyze.py), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/marketplace/universal/scene-vision-analyzer/SKILL.md · 101 lines

How it starts

The opening of the file, as written. The whole thing — 101 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Scene Vision Analyzer

Analyze game scene concept art into structured JSON — the data foundation for 2D→3D reconstruction pipeline.

Workflow

Step 1: Prepare

  1. Load the image (user provides path or URL)
  2. Read references/output-schema.json for the full JSON schema
  3. Read references/analysis-prompts.md for prompt templates

Step 2: Run Analysis (Round 1 — Required)

Send the image to a multimodal AI with:

  • System prompt: from analysis-prompts.md § System Prompt
  • User prompt: from analysis-prompts.md § User Prompt, with {schema} replaced by output-schema.json content
  • Image: attached as vision input

Parse the AI response as JSON. If the response is wrapped in markdown code blocks, strip them.

Step 3: Refinement (Round 2 — Recommended)

Send the Round 1 result back with the refinement prompt from analysis-prompts.md § 第二轮细化 Prompt.

Focus: missed objects, bbox precision, ground_contact accuracy.

Step 4: Consistency Validation (Round 3 — Optional)

Send the Round 2 result with the validation prompt from analysis-prompts.md § 第三轮一致性校验 Prompt.

Focus: size consistency, depth consistency, occlusion logic, shadow direction, group coherence.

Step 5: Save Results

Use scripts/scene_analyze.py helper functions:

from scene_analyze import save_result, generate_summary
# save_result(output_dir, json_string) → saves analysis_result.json + analysis_summary.txt

Or save manually:

  • analysis_result.json — full structured result
  • analysis_summary.txt — human-readable summary

Default output dir: <image_dir>/scene_analysis_<timestamp>/

Quick Reference: Key Schema Fields

Field Purpose Used by downstream
objects[].ground_contact_pct Where object touches ground Step 3 (perspective → 3D coords)
objects[].estimated_size_m Real-world size in meters Step 3 (scale calibration)
objects[].bbox_pct 2D bounding box (%) Step 2 (annotation drawing)
camera.pitch_angle_deg/yaw_angle_deg View angles Step 3 (camera matrix)
camera.horizon_position_pct Horizon line position Step 3 (vanishing point)
spatial_references.perspective_lines Depth cue lines Step 3 (homography)
spatial_references.vanishing_points_pct Convergence points Step 3 (focal length est.)
ground_plane.polygon_pct Ground area outline Step 3 (ground plane fit)
groups Repeated object patterns Step 5 (batch placement)

Read the full file on GitHub · 101 lines

Files

What ships with it

4 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 12d ago First seen · 101 lines · 133 tokens per session scan A 69402bd301e4

Subscribe to this mod's changes

scene-vision-analyzer is a skill published in the GitHub repository IvanYangYangXi/artclaw_bridge (35 stars, last pushed 4mo ago), licensed MIT. It adds 133 tokens to every session and 1,058 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

usd-tools

Infrastructure skill — low-level OpenUSD scene inspection and validation: read layer stacks, traverse prims, validate USD schemas. Use when working directly with raw USD files (usda, usdc, usdz) or verifying USD compliance. Not for Maya-specific USD export — use maya-pipelineexportusd for that. Not for full DCC…

dcc-mcp/dcc-mcp-core · 83 tokens

spatial-interchange

Plan deterministic coordinate-axis and unit conversions between DCCs, engines, and interchange formats from explicit right/up/forward axes and meters-per-unit values.

dcc-mcp/dcc-mcp-core · 35 tokens

game-level-layout

Game level layout review helpers for editor contexts.

dcc-mcp/dcc-mcp-core · 13 tokens

dcc-mcp-core

Foundation library for the DCC Model Context Protocol (MCP) ecosystem. Provides Rust-powered action management, skills system, IPC transport, MCP Streamable HTTP server (2025-03-26 spec, with 2025-06-18 and 2025-11-25 awareness), sandbox security, shared memory, screen capture, USD scene support, and telemetry for…

dcc-mcp/dcc-mcp-core · 109 tokens

dcc-cua

A routing guide for controlling application interfaces through the project’s DCC-CUA system. DCC applications are digital-content tools such as Maya, Blender, and Houdini.

dcc-mcp/dcc-mcp-core · 131 tokens

marketplace-publish-extension

Infrastructure skill — publish (register/update) an extension package to a marketplace catalog. Reads the extension's SKILL.md frontmatter, constructs a CatalogEntry, and upserts it into the target marketplace.json. Optionally commits and pushes when the catalog source is a git repository. Use after scaffolding an…

dcc-mcp/dcc-mcp-core · 88 tokens