modlens AGENTS.md

Project guidance for modlens, a command-line tool that turns an image from a file or web address into text that language models can use later. It describes its supported image-reading providers and deliberately limited scope.

In plain words
What is it for?
Use it when changing image parsing, provider integrations, structured output, or the boundary between image reading and web-page fetching.
Why use it?
It prevents agents from treating the tool as a camera, browser controller, or pixel-editing utility when its job is only to read images into structured text.

Instructions file for CodexOpenCode

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add instructions/liustack/modlens/agents-md
Clone the repo
git clone --depth 1 https://github.com/liustack/modlens

Made for: Codex, OpenCode.

Per session 1,760 This file is loaded in full into every session.
When invoked 1,760 The same file — it is already loaded in full.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.01760 $0.01760
Opus 5 $0.00880 $0.00880
Sonnet 5 $0.00352 $0.00352
Haiku 4.5 $0.00176 $0.00176

Measured 2d ago against content hash 6ee3be1fda80, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

modlens AGENTS.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

AGENTS.md · 102 lines

How it starts

The opening of the file, as written. The whole thing — 102 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Project Overview (for AI Agent)

Goal

Provide the modlens CLI tool that converts image sources (local path or remote URL) into structured text evidence for non-vision LLM workflows.

Scope

The contract: an image the user hands over is read once, and that read leaves text later turns can quote. The image is in the conversation because the user pasted it, dropped a path, or gave a URL.

Do not add:

  • a camera, screenshot capture, or hotkeys
  • CDP, or holding a browser session
  • computer-use (screen plus keyboard and mouse)
  • a pixel toolbox (grounding, crop, pixel-diff, reconstructing a UI)
  • bounding boxes or confidence scores (dropped from the schema on purpose)

Visual parsing is the only job. Web search and page fetching live in modsearch.

Technical Approach

  • Six vision providers behind one interface (src/providers/index.ts). Subprocess providers implement buildInvocation + parseOutput (antigravity-cli, claude-cli, kimi-cli); in-process API providers implement execute (gemini-api, openai, anthropic). antigravity-cli is the zero-config default, and kimi-cli runs only when named, since it spends a subscription.
  • Schema-enforced JSON output wherever the backend allows: --json-schema on the Claude and Antigravity CLIs (kimi-cli has no such flag and uses the template), responseJsonSchema on gemini-api, a forced tool call on anthropic. The openai route uses a template-instance prompt (weak gateways echo raw schemas back) plus shape validation that fails loudly.
  • Layered config: CLI flags > ~/.modlens/config.json (managed by modlens config init/set/show, 0600, masked rendering) > built-ins. Since 3.17.0 a provider's settings come from one source, whole: the file when it mentions that provider, the bound environment variables (GEMINI_API_KEY, GEMINI_BASE_URL, OPENAI_API_KEY, OPENAI_BASE_URL, ANTHROPIC_API_KEY, ANTHROPIC_BASE_URL) when it does not. They used to merge field by field, and an endpoint and a key are one credential, so drawing halves from two places built a pairing that existed in neither. apiKey (file or env) accepts a comma-separated list and rotates after auth, rate-limit, or quota failures. Quota cooldown state lives in ~/.modlens/state.json. MODLENS_MODEL, MODLENS_HARNESS and the proxy conventions are unaffected.
  • Vendor knobs pass through, they are not modelled: <provider>.extraBody (or --extra-body) deep-merges a JSON object into the request body of the three API providers, which is how thinking gets turned off. No per-vendor table lives in the code, because the spelling differs per gateway and a wrong guess either 400s or is ignored silently. The fields carrying the image, the prompt, and the schema are reserved (src/util/extraBody.ts).
  • Paste recovery across harnesses: modlens recover-paste pulls pasted image bytes out of local session storage (pastes never hit a regular temp file). It supports Claude Code and Pi (JSONL transcripts) and OpenCode (SQLite), detects Codex and defers to its on-disk temp files, and scopes to the harness it runs inside via process ancestry. Exact targeting via --session, else newest-image-timestamp scanning. Storage layouts are each harness's internals, so treat this as best-effort.
  • Single responsibility: visual parsing only. Web search and page fetching live in modsearch.

Read the full file on GitHub · 102 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 102 lines · 1,760 tokens per session scan A 6ee3be1fda80

Subscribe to this mod's changes

modlens AGENTS.md is an instructions file published in the GitHub repository liustack/modlens (3,828 stars, last pushed yesterday), licensed MIT. It adds 1,760 tokens to every session, about $0.0088 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other instructions, from other repositories

dsh-TUI AGENTS.md

AGENTS.md instructions for ccch1mneyyy/dsh-TUI, covering agents.md, 仓库布局, 命令, 上游边界与契约 and 约定与红线.

ccch1mneyyy/dsh-TUI · 2,093 tokens

DSH-better-sidebar AGENTS.md

AGENTS.md instructions for omdsh-dev/DSH-better-sidebar, covering dsh-better-sidebar 仓库规则(agents), 1. 仓库硬约束(必须遵守), 2. ci 挂载冒烟(plugin-mount job / pnpm test:mount), 3. dsh 0.1.2-alpha 适配要点(alpha 通道) and 4. npm 发版(github release → npm publish).

omdsh-dev/DSH-better-sidebar · 3,150 tokens

dsh-worktable AGENTS.md

Instructions for Aisland-SJL/dsh-worktable, covering dsh-worktable 项目规则, 协作方式(用户定案,最高优先级), 边界, 构建与验证 and 领域约定(会话中必须遵守).

Aisland-SJL/dsh-worktable · 4,863 tokens

awesome-deepseek-harness-plugins AGENTS.md

Instructions for imsai-sh/awesome-deepseek-harness-plugins, covering repository instructions, the submission gate is the product, generated files, cross-repo contracts (no ci spans both repositories) and permanent urls.

imsai-sh/awesome-deepseek-harness-plugins · 1,057 tokens

dsh-launcher AGENTS.md

Instructions for Ruler4396/dsh-launcher, covering agents.md — 对本仓库中所有 ai agent 的强制指示, 🚨 第一优先:架构铁律(不读就动代码 = 违规), 🔥 快速合规清单(动代码前自检), 🧭 项目地图(定位代码) and ✅ 完成后必须自查.

Ruler4396/dsh-launcher · 1,505 tokens

seektty AGENTS.md

Instructions for Hilbert-beinghappy/seektty: This repository ships one out-of-tree DeepSeek Harness Bundle. Harness remains the only owner of Agent, Session, model, settings, permissions, Profile, plugin, and persistence state.

Hilbert-beinghappy/seektty · 156 tokens